Skip to content

fix(crawlers): Resolve relative and invalid base URLs when extracting links - #2263

Open
Mantisus wants to merge 1 commit into
apify:masterfrom
Mantisus:fix-relative-base
Open

Mantisus wants to merge 1 commit into
apify:masterfrom
Mantisus:fix-relative-base

Conversation

@Mantisus

Copy link
Copy Markdown
Collaborator

Description

extract_links and enqueue_links in the HTTP-based crawlers now resolve a relative <base href> against the page URL and fall back to the page URL when the base is invalid. Previously, a relative base dropped every relative link on the page, and an invalid one made link extraction fail.

PlaywrightCrawler also falls back to the page URL when Chromium reports an invalid base as about:blank, and ignores the whitespace around link URLs.

Testing

  • Added new regression tests.

✍️ Drafted by Claude Code

@Mantisus Mantisus self-assigned this Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant