🔥 Supercharge your AI agents with data from the web and beyond. A web data API to search, scrape, and access more sources.
-
Updated
Sep 29, 2026 - TypeScript
🔥 Supercharge your AI agents with data from the web and beyond. A web data API to search, scrape, and access more sources.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Python scraper based on AI
The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.
Turn any website into a structured API. Extract, automate, search and monitor the web.
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.
Declarative data automation language and Go runtime for structured extraction workflows.
Extract Keywords from sentence or Replace keywords in sentences.
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
The undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.
Python package for scraping recipes data
ContextGem: Effortless LLM extraction from documents
To extract article from given URL
Converts a pdf file into a text file while keeping the layout of the original pdf. Useful to extract the content from a table in a pdf file for instance. This is a subclass of PDFTextStripper class (from the Apache PDFBox library).
🚚 Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone
Lightweight library for scraping web-sites with LLMs
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
To associate your repository with the data-extraction topic, visit your repo's landing page and select "manage topics."