You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit cfd9dde
Browse filesBrowse the repository at this point in the historyBrowse files
This example demonstrates how to fill and submit a web form using the <ApiLinkto="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLinkto="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLinkto="class/ParselCrawler">`ParselCrawler`</ApiLink>.
15
+
This example demonstrates how to fill and submit a web form using the <ApiLinkto="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLinkto="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLinkto="class/ParselCrawler">`ParselCrawler`</ApiLink>. These two crawlers can also [fill in the form automatically](#fill-in-the-form-automatically).
15
16
16
17
We are going to use the [httpbin.org](https://httpbin.org) website to demonstrate how it works.
17
18
@@ -118,3 +119,21 @@ Finally, run your crawler. Your logs should show something like this:
118
119
```
119
120
120
121
This log output confirms that the crawler successfully submitted the form and processed the response. Congratulations! You have successfully filled and submitted a web form using the <ApiLinkto="class/HttpCrawler">`HttpCrawler`</ApiLink>.
122
+
123
+
## Fill in the form automatically
124
+
125
+
The <ApiLinkto="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLinkto="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> can build the form request for you. Their crawling contexts provide the <ApiLinkto="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper, which reads the form from the page, fills in your values and returns a list with the request that submits it the way a browser does. The action URL, the method and the encoding come from the form itself, so you only need the field names from [Investigate the form fields](#investigate-the-form-fields).
126
+
127
+
The crawler below opens the page with the form. The default handler fills in the form with the `fields` argument and enqueues the submission with a label. A separate handler for that label processes the response.
-`fields` replaces the values of the listed fields and adds the ones the form doesn't have. A list submits the field once per value, as with the `topping` checkboxes.
136
+
- Fields you don't list keep the values from the page, so hidden inputs such as CSRF tokens are submitted as they are. A CSRF token is tied to the session cookie, so pass `session_id=context.session.id` to send the form in the same session. The request also carries the `Referer` and `Origin` headers a browser sends, which some CSRF checks require.
137
+
- On a page with several forms, the helper submits the one sharing the most field names with `fields`, or the first one if none shares any. It skips forms that can't be submitted, for example because their action is JavaScript, but never falls back to a form sharing fewer names, so the list can be empty. To pick a form yourself, pass a CSS selector such as `selector='#order'`. To submit each form, pass `all_forms=True`. Then `fields` only replaces the fields each form has.
138
+
- The first enabled submit button of the form is clicked by default, and a form without one is submitted anyway. Use the `click` argument to pick another button by its attributes, even a disabled one, or to submit without one.
139
+
- The page decides where its form is sent. To enqueue only requests to the same host, call `context.add_requests(requests, strategy='same-hostname')`.
@@ -263,7 +263,7 @@ Scrapy retries failed requests with `RetryMiddleware` and reports terminal failu
263
263
264
264
## Forms and login
265
265
266
-
Scrapy submits forms with `FormRequest`, which encodes `formdata` as `form-urlencoded` and sets the header for you. Crawlee's <ApiLinkto="class/Request#from_url">`payload`</ApiLink> takes the raw request body, so encode the fields yourself with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
266
+
Scrapy submits forms with `FormRequest.from_response`, which reads the form from the page, keeps its hidden fields and encodes the data for you. Crawlee's <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper does the same in the <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink>. Pass your values in `fields` and request options such as `label` or `session_id` as keyword arguments. Like `from_response`, it submits a single form. Scrapy takes the first form by default, while the helper prefers the one sharing the most field names with `fields`. It returns a list, which is empty when no form matches, so enqueue it with `add_requests`. For a plain `FormRequest` that doesn't come from a form on the page, use <ApiLink to="class/Request#from_url">`Request.from_url`</ApiLink>. Its `payload` is the raw request body, so encode the fields with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
0 commit comments