Skip to content

Commit cfd9dde

Browse files
committed
parse HTML responses starting with an XML declaration as HTML in ParselCrawler
1 parent b2de53a commit cfd9dde

10 files changed

Lines changed: 1418 additions & 40 deletions

File tree

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,39 @@
1+
import asyncio
2+
3+
from crawlee.crawlers import ParselCrawler, ParselCrawlingContext
4+
5+
6+
async def main() -> None:
7+
crawler = ParselCrawler()
8+
9+
# Fill in the form on the page and enqueue its submission.
10+
@crawler.router.default_handler
11+
async def request_handler(context: ParselCrawlingContext) -> None:
12+
context.log.info(f'Filling in the form on {context.request.url} ...')
13+
requests = await context.extract_form_requests(
14+
fields={
15+
'custname': 'John Doe',
16+
'custtel': '1234567890',
17+
'custemail': 'johndoe@example.com',
18+
'size': 'large',
19+
'topping': ['bacon', 'cheese', 'mushroom'],
20+
'delivery': '13:00',
21+
'comments': 'Please ring the doorbell upon arrival.',
22+
},
23+
label='form-result',
24+
)
25+
await context.add_requests(requests)
26+
27+
# Process the response to the form submission.
28+
@crawler.router.handler('form-result')
29+
async def form_result_handler(context: ParselCrawlingContext) -> None:
30+
context.log.info(f'Processing {context.request.url} ...')
31+
response = (await context.http_response.read()).decode('utf-8')
32+
context.log.info(f'Response: {response}') # To see the response in the logs.
33+
34+
# Run the crawler with the page containing the form.
35+
await crawler.run(['https://httpbin.org/forms/post'])
36+
37+
38+
if __name__ == '__main__':
39+
asyncio.run(main())

‎docs/examples/fill_and_submit_web_form.mdx‎

Lines changed: 20 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -10,8 +10,9 @@ import RunnableCodeBlock from '@site/src/components/RunnableCodeBlock';
1010

1111
import RequestExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_request.py';
1212
import CrawlerExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_crawler.py';
13+
import AutomatedExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_automated.py';
1314

14-
This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>.
15+
This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>. These two crawlers can also [fill in the form automatically](#fill-in-the-form-automatically).
1516

1617
We are going to use the [httpbin.org](https://httpbin.org) website to demonstrate how it works.
1718

@@ -118,3 +119,21 @@ Finally, run your crawler. Your logs should show something like this:
118119
```
119120

120121
This log output confirms that the crawler successfully submitted the form and processed the response. Congratulations! You have successfully filled and submitted a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink>.
122+
123+
## Fill in the form automatically
124+
125+
The <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> can build the form request for you. Their crawling contexts provide the <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper, which reads the form from the page, fills in your values and returns a list with the request that submits it the way a browser does. The action URL, the method and the encoding come from the form itself, so you only need the field names from [Investigate the form fields](#investigate-the-form-fields).
126+
127+
The crawler below opens the page with the form. The default handler fills in the form with the `fields` argument and enqueues the submission with a label. A separate handler for that label processes the response.
128+
129+
<RunnableCodeBlock className="language-python" language="python">
130+
{AutomatedExample}
131+
</RunnableCodeBlock>
132+
133+
Note that:
134+
135+
- `fields` replaces the values of the listed fields and adds the ones the form doesn't have. A list submits the field once per value, as with the `topping` checkboxes.
136+
- Fields you don't list keep the values from the page, so hidden inputs such as CSRF tokens are submitted as they are. A CSRF token is tied to the session cookie, so pass `session_id=context.session.id` to send the form in the same session. The request also carries the `Referer` and `Origin` headers a browser sends, which some CSRF checks require.
137+
- On a page with several forms, the helper submits the one sharing the most field names with `fields`, or the first one if none shares any. It skips forms that can't be submitted, for example because their action is JavaScript, but never falls back to a form sharing fewer names, so the list can be empty. To pick a form yourself, pass a CSS selector such as `selector='#order'`. To submit each form, pass `all_forms=True`. Then `fields` only replaces the fields each form has.
138+
- The first enabled submit button of the form is clicked by default, and a form without one is submitted anyway. Use the `click` argument to pick another button by its attributes, even a disabled one, or to submit without one.
139+
- The page decides where its form is sent. To enqueue only requests to the same host, call `context.add_requests(requests, strategy='same-hostname')`.

‎docs/guides/code_examples/scrapy_migration/crawlee_post.py‎

Lines changed: 8 additions & 28 deletions
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,5 @@
11
import asyncio
2-
from urllib.parse import urlencode
32

4-
from crawlee import Request
53
from crawlee.crawlers import ParselCrawler, ParselCrawlingContext
64

75

@@ -14,34 +12,16 @@ async def login_page(context: ParselCrawlingContext) -> None:
1412
if not context.session:
1513
raise RuntimeError('Session not found')
1614

17-
token = context.selector.css('input[name="csrf_token"]::attr(value)').get()
18-
19-
# The CSRF token is required for the POST to succeed. If it's missing,
20-
# the login will fail.
21-
if not token:
22-
raise RuntimeError('CSRF token not found')
23-
24-
form = {'csrf_token': token, 'username': 'user', 'password': 'pass'}
25-
2615
# highlight-start
27-
# Crawlee's `payload` is the raw request body, so encode the fields yourself
28-
# and set the `Content-Type`. Scrapy's `FormRequest` does both for you.
29-
await context.add_requests(
30-
[
31-
Request.from_url(
32-
'https://quotes.toscrape.com/login',
33-
method='POST',
34-
payload=urlencode(form),
35-
headers={'content-type': 'application/x-www-form-urlencoded'},
36-
label='after-login',
37-
# Bind the POST to the same session so its CSRF cookie matches.
38-
session_id=context.session.id,
39-
# The POST shares the GET's URL. Include the method and payload
40-
# in the unique key, or the queue drops it as a duplicate.
41-
use_extended_unique_key=True,
42-
)
43-
]
16+
# Like Scrapy's `FormRequest.from_response`, the helper keeps the hidden
17+
# `csrf_token` field, encodes the data and sets the `Content-Type` header.
18+
requests = await context.extract_form_requests(
19+
fields={'username': 'user', 'password': 'pass'},
20+
label='after-login',
21+
# Bind the POST to the same session so its CSRF cookie matches.
22+
session_id=context.session.id,
4423
)
24+
await context.add_requests(requests)
4525
# highlight-end
4626

4727
@crawler.router.handler('after-login')

‎docs/guides/scrapy_migration.mdx‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,7 @@ Both frameworks give you a request scheduler, filtering of duplicate requests, r
8080
| `response.follow()` / `yield Request(...)` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links`</ApiLink> / <ApiLink to="class/AddRequestsFunction">`add_requests`</ApiLink> |
8181
| `dont_filter=True` | <ApiLink to="class/Request#from_url">`Request.from_url(always_enqueue=True)`</ApiLink> |
8282
| `allowed_domains` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links(strategy=...)`</ApiLink> |
83-
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> |
83+
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> / <ApiLink to="class/BeautifulSoupCrawlingContext#extract_form_requests">`context.extract_form_requests(...)`</ApiLink> |
8484
| Item pipelines | <ApiLink to="class/Dataset">`Dataset`</ApiLink> |
8585
| Downloader / spider middlewares | <ApiLink to="class/Router#use">`router.use()`</ApiLink>, navigation hooks, <ApiLink to="class/HttpClient">HTTP clients</ApiLink> |
8686
| `settings.py` | <ApiLink to="class/Configuration">`Configuration`</ApiLink> + crawler arguments |
@@ -263,7 +263,7 @@ Scrapy retries failed requests with `RetryMiddleware` and reports terminal failu
263263

264264
## Forms and login
265265

266-
Scrapy submits forms with `FormRequest`, which encodes `formdata` as `form-urlencoded` and sets the header for you. Crawlee's <ApiLink to="class/Request#from_url">`payload`</ApiLink> takes the raw request body, so encode the fields yourself with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
266+
Scrapy submits forms with `FormRequest.from_response`, which reads the form from the page, keeps its hidden fields and encodes the data for you. Crawlee's <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper does the same in the <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink>. Pass your values in `fields` and request options such as `label` or `session_id` as keyword arguments. Like `from_response`, it submits a single form. Scrapy takes the first form by default, while the helper prefers the one sharing the most field names with `fields`. It returns a list, which is empty when no form matches, so enqueue it with `add_requests`. For a plain `FormRequest` that doesn't come from a form on the page, use <ApiLink to="class/Request#from_url">`Request.from_url`</ApiLink>. Its `payload` is the raw request body, so encode the fields with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
267267

268268
<Tabs groupId="scrapy-migration-login">
269269
<TabItem value="scrapy" label="Scrapy">

0 commit comments

Comments
 (0)