Skip to content
39 changes: 39 additions & 0 deletions docs/examples/code_examples/fill_and_submit_web_form_automated.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
import asyncio

from crawlee.crawlers import ParselCrawler, ParselCrawlingContext


async def main() -> None:
crawler = ParselCrawler()

# Fill in the form on the page and enqueue its submission.
@crawler.router.default_handler
async def request_handler(context: ParselCrawlingContext) -> None:
context.log.info(f'Filling in the form on {context.request.url} ...')
requests = await context.extract_form_requests(
fields={
'custname': 'John Doe',
'custtel': '1234567890',
'custemail': 'johndoe@example.com',
'size': 'large',
'topping': ['bacon', 'cheese', 'mushroom'],
'delivery': '13:00',
'comments': 'Please ring the doorbell upon arrival.',
},
label='form-result',
)
await context.add_requests(requests)

# Process the response to the form submission.
@crawler.router.handler('form-result')
async def form_result_handler(context: ParselCrawlingContext) -> None:
context.log.info(f'Processing {context.request.url} ...')
response = (await context.http_response.read()).decode('utf-8')
context.log.info(f'Response: {response}') # To see the response in the logs.

# Run the crawler with the page containing the form.
await crawler.run(['https://httpbin.org/forms/post'])


if __name__ == '__main__':
asyncio.run(main())
21 changes: 20 additions & 1 deletion docs/examples/fill_and_submit_web_form.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,9 @@ import RunnableCodeBlock from '@site/src/components/RunnableCodeBlock';

import RequestExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_request.py';
import CrawlerExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_crawler.py';
import AutomatedExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_automated.py';

This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>.
This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>. These two crawlers can also [fill in the form automatically](#fill-in-the-form-automatically).

We are going to use the [httpbin.org](https://httpbin.org) website to demonstrate how it works.

Expand Down Expand Up @@ -118,3 +119,21 @@ Finally, run your crawler. Your logs should show something like this:
```

This log output confirms that the crawler successfully submitted the form and processed the response. Congratulations! You have successfully filled and submitted a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink>.

## Fill in the form automatically

The <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> can build the form request for you. Their crawling contexts provide the <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper, which reads the form from the page, fills in your values and returns a list with the request that submits it the way a browser does. The action URL, the method and the encoding come from the form itself, so you only need the field names from [Investigate the form fields](#investigate-the-form-fields).

The crawler below opens the page with the form. The default handler fills in the form with the `fields` argument and enqueues the submission with a label. A separate handler for that label processes the response.

<RunnableCodeBlock className="language-python" language="python">
{AutomatedExample}
</RunnableCodeBlock>

Note that:

- `fields` replaces the values of the listed fields and adds the ones the form doesn't have. A list submits the field once per value, as with the `topping` checkboxes.
- Fields you don't list keep the values from the page, so hidden inputs such as CSRF tokens are submitted as they are. A CSRF token is tied to the session cookie, so pass `session_id=context.session.id` to send the form in the same session. The request also carries the `Referer` and `Origin` headers a browser sends, which some CSRF checks require.
- On a page with several forms, the helper submits the one sharing the most field names with `fields`, or the first one if none shares any. It skips forms that can't be submitted, for example because their action is JavaScript, but never falls back to a form sharing fewer names, so the list can be empty. To pick a form yourself, pass a CSS selector such as `selector='#order'`. To submit each form, pass `all_forms=True`. Then `fields` only replaces the fields each form has.
- The first enabled submit button of the form is clicked by default, and a form without one is submitted anyway. Use the `click` argument to pick another button by its attributes, even a disabled one, or to submit without one.
- The page decides where its form is sent. To enqueue only requests to the same host, call `context.add_requests(requests, strategy='same-hostname')`.
36 changes: 8 additions & 28 deletions docs/guides/code_examples/scrapy_migration/crawlee_post.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,5 @@
import asyncio
from urllib.parse import urlencode

from crawlee import Request
from crawlee.crawlers import ParselCrawler, ParselCrawlingContext


Expand All @@ -14,34 +12,16 @@ async def login_page(context: ParselCrawlingContext) -> None:
if not context.session:
raise RuntimeError('Session not found')

token = context.selector.css('input[name="csrf_token"]::attr(value)').get()

# The CSRF token is required for the POST to succeed. If it's missing,
# the login will fail.
if not token:
raise RuntimeError('CSRF token not found')

form = {'csrf_token': token, 'username': 'user', 'password': 'pass'}

# highlight-start
# Crawlee's `payload` is the raw request body, so encode the fields yourself
# and set the `Content-Type`. Scrapy's `FormRequest` does both for you.
await context.add_requests(
[
Request.from_url(
'https://quotes.toscrape.com/login',
method='POST',
payload=urlencode(form),
headers={'content-type': 'application/x-www-form-urlencoded'},
label='after-login',
# Bind the POST to the same session so its CSRF cookie matches.
session_id=context.session.id,
# The POST shares the GET's URL. Include the method and payload
# in the unique key, or the queue drops it as a duplicate.
use_extended_unique_key=True,
)
]
# Like Scrapy's `FormRequest.from_response`, the helper keeps the hidden
# `csrf_token` field, encodes the data and sets the `Content-Type` header.
requests = await context.extract_form_requests(
fields={'username': 'user', 'password': 'pass'},
label='after-login',
# Bind the POST to the same session so its CSRF cookie matches.
session_id=context.session.id,
)
await context.add_requests(requests)
# highlight-end

@crawler.router.handler('after-login')
Expand Down
4 changes: 2 additions & 2 deletions docs/guides/scrapy_migration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@ Both frameworks give you a request scheduler, filtering of duplicate requests, r
| `response.follow()` / `yield Request(...)` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links`</ApiLink> / <ApiLink to="class/AddRequestsFunction">`add_requests`</ApiLink> |
| `dont_filter=True` | <ApiLink to="class/Request#from_url">`Request.from_url(always_enqueue=True)`</ApiLink> |
| `allowed_domains` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links(strategy=...)`</ApiLink> |
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> |
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> / <ApiLink to="class/BeautifulSoupCrawlingContext#extract_form_requests">`context.extract_form_requests(...)`</ApiLink> |
| Item pipelines | <ApiLink to="class/Dataset">`Dataset`</ApiLink> |
| Downloader / spider middlewares | <ApiLink to="class/Router#use">`router.use()`</ApiLink>, navigation hooks, <ApiLink to="class/HttpClient">HTTP clients</ApiLink> |
| `settings.py` | <ApiLink to="class/Configuration">`Configuration`</ApiLink> + crawler arguments |
Expand Down Expand Up @@ -263,7 +263,7 @@ Scrapy retries failed requests with `RetryMiddleware` and reports terminal failu

## Forms and login

Scrapy submits forms with `FormRequest`, which encodes `formdata` as `form-urlencoded` and sets the header for you. Crawlee's <ApiLink to="class/Request#from_url">`payload`</ApiLink> takes the raw request body, so encode the fields yourself with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
Scrapy submits forms with `FormRequest.from_response`, which reads the form from the page, keeps its hidden fields and encodes the data for you. Crawlee's <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper does the same in the <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink>. Pass your values in `fields` and request options such as `label` or `session_id` as keyword arguments. Like `from_response`, it submits a single form. Scrapy takes the first form by default, while the helper prefers the one sharing the most field names with `fields`. It returns a list, which is empty when no form matches, so enqueue it with `add_requests`. For a plain `FormRequest` that doesn't come from a form on the page, use <ApiLink to="class/Request#from_url">`Request.from_url`</ApiLink>. Its `payload` is the raw request body, so encode the fields with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).

<Tabs groupId="scrapy-migration-login">
<TabItem value="scrapy" label="Scrapy">
Expand Down
Loading
Loading