Data and scraping for AI

Tools that turn web pages, PDFs and documents into clean text or structured data that models can use.

Crawl4AI leads with 72, ahead of Firecrawl (68) and Jina Reader (44). 3 projects ranked by score.

The ranking

Data and scraping for AI: full ranking
RankProjectAdoptionFreshnessMaintenanceEasy to runAgent-readyScore
1Crawl4AIPython crawler that turns pages into LLM-ready markdown, with a Docker API85.2k stars, Apache-2.0, last commit Oct 2026701008267072 out of 100
2FirecrawlWeb scraping and crawling API that returns LLM-ready markdown Live demo ↗ (opens in a new tab)190.2k stars, AGPL-3.0, last commit Oct 20269910045337068 out of 100
3Jina ReaderConverts any URL or search query into LLM-friendly markdown Live demo ↗ (opens in a new tab)12.1k stars, Apache-2.0, last commit May 2026208115504044 out of 100

Momentum, Verified build, Docs and Privacy are not measured yet; their weight goes to the signals shown. A dash means the signal is not scored for that kind of project. Hover a number for its rating in words.

Reviews

172 out of 100

Crawl4AI

Python crawler that turns pages into LLM-ready markdown, with a Docker API

85.2k stars, Apache-2.0, last commit Oct 2026

Async Playwright crawler (pip install crawl4ai) that renders pages in Chromium, Firefox or WebKit and emits clean or filtered markdown, with CSS, XPath and regex extraction needing no LLM, or LLM extraction via any LiteLLM provider. Deep crawling (BFS, DFS, priority-scored) and adaptive crawling are built in. A Docker server on port 11235 exposes /md, /html, /crawl, /screenshot, /pdf and MCP behind an API token.

Strengths

  • Structured extraction with CSS, XPath or regex schemas needs no LLM or API key
  • Docker server with REST, streaming crawl, MCP, dashboard and playground; amd64 and arm64
  • Deep crawl strategies with crash recovery via resume_state
  • Persistent browser profiles, CDP remote browsers and an undetected-browser adapter

Weaknesses

  • Apache-2.0 but requires attribution (badge or text) in your project
  • Docker server answers only inside the container until CRAWL4AI_API_TOKEN is set
  • Web search and answer endpoints exist only in the paid cloud
  • Runs full browsers; the docker run example allocates 1 GB shared memory
  • no GPU
  • Docker + Compose
  • Needs Playwright Chromium (installed by crawl4ai-setup)
  • Models: any LiteLLM provider for LLM extraction (OpenAI, Ollama and others)
  • port 11235
268 out of 100

Firecrawl

Web scraping and crawling API that returns LLM-ready markdown

Live demo ↗ (opens in a new tab)190.2k stars, AGPL-3.0, last commit Oct 2026

API that turns URLs into markdown, HTML, screenshots or schema-based JSON, with endpoints for search, scrape, crawl, map, batch scrape, page interaction and a prompt-driven agent. Handles JS-rendered pages and parses hosted PDFs and DOCX. SDKs for Python, Node, Go, Java, Elixir, Rust and Ruby plus an MCP server and CLI, for teams feeding web content to RAG pipelines and agents.

Strengths

  • Seven SDKs plus CLI and MCP server; SDKs poll async crawl jobs automatically
  • Crawl, map and batch-scrape endpoints return job IDs for large sites
  • Scrape supports actions (click, scroll, write, wait) before extraction
  • Compose file at the repo root for self-hosting

Weaknesses

  • README is written around the hosted API and keys; self-hosting lives in separate docs
  • AGPL-3.0 license; network use of a modified version triggers source obligations
  • Agent endpoint runs the hosted spark-2 model, not a local LLM
  • Proxy rotation and anti-bot handling are hosted-service features
  • no GPU
  • Compose
344 out of 100

Jina Reader

Converts any URL or search query into LLM-friendly markdown

Live demo ↗ (opens in a new tab)12.1k stars, Apache-2.0, last commit May 2026

Open-source branch of the service behind r.jina.ai and s.jina.ai: fetches a page with headless Chrome or curl-impersonate, parses PDFs and Office files, and returns markdown, text, HTML, screenshots or JSON controlled by request headers (engine, timeout, token limits). The ghcr.io image bundles Chrome, LibreOffice and CJK fonts, serves HTTP/1.1 on 8081 and h2c on 8080, and runs stateless or with S3-compatible caching.

Strengths

  • Prebuilt image with Chrome, LibreOffice and CJK fonts; stateless by default
  • Fine-grained headers: x-respond-timing, x-max-tokens, x-token-budget, x-target-selector
  • Optional VLM captions for images without alt text
  • Semantic markdown chunking by heading or block level

Weaknesses

  • Hosted proxy pool, rate limiting and MongoDB storage layer are not in the OSS branch
  • Needs non-redistributable assets (MaxMind GeoLite2, Source Han Sans) fetched at build
  • Last commit May 2026; the SaaS resync was April 2026
  • Default h2c port 8080 needs --http2-prior-knowledge from curl; use 8081 otherwise
  • no GPU
  • Docker + Compose
  • Needs Headless Chrome and LibreOffice (bundled in image), S3-compatible bucket (optional cache), VLM endpoint for image captions (optional)
  • port 8081

Written from each project's README and checked facts. Spot something wrong? Report it on GitHub (opens in a new tab).