11st of 3 in Data and scraping for AI

Crawl4AI

Python crawler that turns pages into LLM-ready markdown, with a Docker API

Stars
85.2k
License
Apache-2.0
Last commit
Oct 2026
Last release
Sep 2026

Overview

Async Playwright crawler (pip install crawl4ai) that renders pages in Chromium, Firefox or WebKit and emits clean or filtered markdown, with CSS, XPath and regex extraction needing no LLM, or LLM extraction via any LiteLLM provider. Deep crawling (BFS, DFS, priority-scored) and adaptive crawling are built in. A Docker server on port 11235 exposes /md, /html, /crawl, /screenshot, /pdf and MCP behind an API token.

Who it is for: Developers building scrapers and RAG ingestion pipelines

Strengths

  • Structured extraction with CSS, XPath or regex schemas needs no LLM or API key
  • Docker server with REST, streaming crawl, MCP, dashboard and playground; amd64 and arm64
  • Deep crawl strategies with crash recovery via resume_state
  • Persistent browser profiles, CDP remote browsers and an undetected-browser adapter

Weaknesses

  • Apache-2.0 but requires attribution (badge or text) in your project
  • Docker server answers only inside the container until CRAWL4AI_API_TOKEN is set
  • Web search and answer endpoints exist only in the paid cloud
  • Runs full browsers; the docker run example allocates 1 GB shared memory

What it needs

  • no GPU
  • Docker + Compose
  • Needs Playwright Chromium (installed by crawl4ai-setup)
  • Models: any LiteLLM provider for LLM extraction (OpenAI, Ollama and others)
  • port 11235

Also in Data and scraping for AI

See all 3
Also in Data and scraping for AI
RankProjectScore
2FirecrawlWeb scraping and crawling API that returns LLM-ready markdown Live demo ↗ (opens in a new tab)190.2k stars, AGPL-3.068 out of 100
3Jina ReaderConverts any URL or search query into LLM-friendly markdown Live demo ↗ (opens in a new tab)12.1k stars, Apache-2.044 out of 100