22nd of 3 in Data and scraping for AI
Firecrawl
Web scraping and crawling API that returns LLM-ready markdown
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 190.2k
- License
- AGPL-3.0
- Last commit
- Oct 2026
- Last release
- Jun 2026
Overview
API that turns URLs into markdown, HTML, screenshots or schema-based JSON, with endpoints for search, scrape, crawl, map, batch scrape, page interaction and a prompt-driven agent. Handles JS-rendered pages and parses hosted PDFs and DOCX. SDKs for Python, Node, Go, Java, Elixir, Rust and Ruby plus an MCP server and CLI, for teams feeding web content to RAG pipelines and agents.
Who it is for: Teams feeding web content to RAG pipelines and agents
Strengths
- Seven SDKs plus CLI and MCP server; SDKs poll async crawl jobs automatically
- Crawl, map and batch-scrape endpoints return job IDs for large sites
- Scrape supports actions (click, scroll, write, wait) before extraction
- Compose file at the repo root for self-hosting
Weaknesses
- README is written around the hosted API and keys; self-hosting lives in separate docs
- AGPL-3.0 license; network use of a modified version triggers source obligations
- Agent endpoint runs the hosted spark-2 model, not a local LLM
- Proxy rotation and anti-bot handling are hosted-service features
What it needs
- no GPU
- Compose
Also in Data and scraping for AI
See all 3| Rank | Project | Score |
|---|---|---|
| 1 | Crawl4AIPython crawler that turns pages into LLM-ready markdown, with a Docker API | 72 out of 100 |
| 3 | Jina ReaderConverts any URL or search query into LLM-friendly markdown | 44 out of 100 |