Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions docs/examples/code_examples/fill_and_submit_web_form_automated.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
import asyncio

from crawlee.crawlers import ParselCrawler, ParselCrawlingContext


async def main() -> None:
crawler = ParselCrawler()

# Fill in the form on the page and enqueue its submission.
@crawler.router.default_handler
async def request_handler(context: ParselCrawlingContext) -> None:
context.log.info(f'Filling in the form on {context.request.url} ...')
requests = await context.extract_form_requests(
fields={
'custname': 'John Doe',
'custtel': '1234567890',
'custemail': 'johndoe@example.com',
'size': 'large',
'topping': ['bacon', 'cheese', 'mushroom'],
'delivery': '13:00',
'comments': 'Please ring the doorbell upon arrival.',
},
label='form-result',
)
await context.add_requests(requests)

# Process the response to the form submission.
@crawler.router.handler('form-result')
async def form_result_handler(context: ParselCrawlingContext) -> None:
context.log.info(f'Processing {context.request.url} ...')
response = (await context.http_response.read()).decode('utf-8')
context.log.info(f'Response: {response}') # To see the response in the logs.

# Run the crawler with the page containing the form.
await crawler.run(['https://httpbin.org/forms/post'])


if __name__ == '__main__':
asyncio.run(main())
21 changes: 20 additions & 1 deletion docs/examples/fill_and_submit_web_form.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,9 @@ import RunnableCodeBlock from '@site/src/components/RunnableCodeBlock';

import RequestExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_request.py';
import CrawlerExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_crawler.py';
import AutomatedExample from '!!raw-loader!roa-loader!./code_examples/fill_and_submit_web_form_automated.py';

This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>.
This example demonstrates how to fill and submit a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> crawler. The same approach applies to any crawler that inherits from it, such as the <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> or <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink>. These two crawlers can also [fill in the form automatically](#fill-in-the-form-automatically).

We are going to use the [httpbin.org](https://httpbin.org) website to demonstrate how it works.

Expand Down Expand Up @@ -118,3 +119,21 @@ Finally, run your crawler. Your logs should show something like this:
```

This log output confirms that the crawler successfully submitted the form and processed the response. Congratulations! You have successfully filled and submitted a web form using the <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink>.

## Fill in the form automatically

The <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> can build the form request for you. Their crawling contexts provide the <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper, which reads the form from the page, fills in your values and returns a list with the request that submits it the way a browser does. The action URL, the method and the encoding come from the form itself, so you only need the field names from [Investigate the form fields](#investigate-the-form-fields).

The crawler below opens the page with the form. The default handler fills in the form with the `fields` argument and enqueues the submission with a label. A separate handler for that label processes the response.

<RunnableCodeBlock className="language-python" language="python">
{AutomatedExample}
</RunnableCodeBlock>

Note that:

- `fields` replaces the values of the listed fields and adds the ones the form doesn't have. A list submits the field once per value, as with the `topping` checkboxes.
- Fields you don't list keep the values from the page, so hidden inputs such as CSRF tokens are submitted as they are. A CSRF token is tied to the session cookie, so pass `session_id=context.session.id` to send the form in the same session. The request also carries the `Referer` and `Origin` headers a browser sends, which some CSRF checks require.
- On a page with several forms, the helper submits the one sharing the most field names with `fields`, or the first one if none shares any. It skips forms that can't be submitted, for example because their action is JavaScript, but never falls back to a form sharing fewer names, so the list can be empty. To pick a form yourself, pass a CSS selector such as `selector='#order'`. To submit each form, pass `all_forms=True`. Then `fields` only replaces the fields each form has.
- The first enabled submit button of the form is clicked by default, and a form without one is submitted anyway. Use the `click` argument to pick another button by its attributes, even a disabled one, or to submit without one.
- The page decides where its form is sent. To enqueue only requests to the same host, call `context.add_requests(requests, strategy='same-hostname')`.
36 changes: 8 additions & 28 deletions docs/guides/code_examples/scrapy_migration/crawlee_post.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,5 @@
import asyncio
from urllib.parse import urlencode

from crawlee import Request
from crawlee.crawlers import ParselCrawler, ParselCrawlingContext


Expand All @@ -14,34 +12,16 @@ async def login_page(context: ParselCrawlingContext) -> None:
if not context.session:
raise RuntimeError('Session not found')

token = context.selector.css('input[name="csrf_token"]::attr(value)').get()

# The CSRF token is required for the POST to succeed. If it's missing,
# the login will fail.
if not token:
raise RuntimeError('CSRF token not found')

form = {'csrf_token': token, 'username': 'user', 'password': 'pass'}

# highlight-start
# Crawlee's `payload` is the raw request body, so encode the fields yourself
# and set the `Content-Type`. Scrapy's `FormRequest` does both for you.
await context.add_requests(
[
Request.from_url(
'https://quotes.toscrape.com/login',
method='POST',
payload=urlencode(form),
headers={'content-type': 'application/x-www-form-urlencoded'},
label='after-login',
# Bind the POST to the same session so its CSRF cookie matches.
session_id=context.session.id,
# The POST shares the GET's URL. Include the method and payload
# in the unique key, or the queue drops it as a duplicate.
use_extended_unique_key=True,
)
]
# Like Scrapy's `FormRequest.from_response`, the helper keeps the hidden
# `csrf_token` field, encodes the data and sets the `Content-Type` header.
requests = await context.extract_form_requests(
fields={'username': 'user', 'password': 'pass'},
label='after-login',
# Bind the POST to the same session so its CSRF cookie matches.
session_id=context.session.id,
)
await context.add_requests(requests)
# highlight-end

@crawler.router.handler('after-login')
Expand Down
4 changes: 2 additions & 2 deletions docs/guides/scrapy_migration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@ Both frameworks give you a request scheduler, filtering of duplicate requests, r
| `response.follow()` / `yield Request(...)` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links`</ApiLink> / <ApiLink to="class/AddRequestsFunction">`add_requests`</ApiLink> |
| `dont_filter=True` | <ApiLink to="class/Request#from_url">`Request.from_url(always_enqueue=True)`</ApiLink> |
| `allowed_domains` | <ApiLink to="class/EnqueueLinksFunction">`enqueue_links(strategy=...)`</ApiLink> |
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> |
| `scrapy.FormRequest` | <ApiLink to="class/Request#from_url">`Request.from_url(method='POST', payload=...)`</ApiLink> / <ApiLink to="class/BeautifulSoupCrawlingContext#extract_form_requests">`context.extract_form_requests(...)`</ApiLink> |
| Item pipelines | <ApiLink to="class/Dataset">`Dataset`</ApiLink> |
| Downloader / spider middlewares | <ApiLink to="class/Router#use">`router.use()`</ApiLink>, navigation hooks, <ApiLink to="class/HttpClient">HTTP clients</ApiLink> |
| `settings.py` | <ApiLink to="class/Configuration">`Configuration`</ApiLink> + crawler arguments |
Expand Down Expand Up @@ -263,7 +263,7 @@ Scrapy retries failed requests with `RetryMiddleware` and reports terminal failu

## Forms and login

Scrapy submits forms with `FormRequest`, which encodes `formdata` as `form-urlencoded` and sets the header for you. Crawlee's <ApiLink to="class/Request#from_url">`payload`</ApiLink> takes the raw request body, so encode the fields yourself with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).
Scrapy submits forms with `FormRequest.from_response`, which reads the form from the page, keeps its hidden fields and encodes the data for you. Crawlee's <ApiLink to="class/ParselCrawlingContext#extract_form_requests">`extract_form_requests`</ApiLink> helper does the same in the <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> and <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink>. Pass your values in `fields` and request options such as `label` or `session_id` as keyword arguments. Like `from_response`, it submits a single form. Scrapy takes the first form by default, while the helper prefers the one sharing the most field names with `fields`. It returns a list, which is empty when no form matches, so enqueue it with `add_requests`. For a plain `FormRequest` that doesn't come from a form on the page, use <ApiLink to="class/Request#from_url">`Request.from_url`</ApiLink>. Its `payload` is the raw request body, so encode the fields with `urllib.parse.urlencode` and set the `Content-Type` through `headers=`. For a full login flow with session reuse, see the [Logging in with a crawler guide](./logging-in-with-a-crawler).

<Tabs groupId="scrapy-migration-login">
<TabItem value="scrapy" label="Scrapy">
Expand Down
Loading
Loading