22nd of 3 in Data and scraping for AI

Firecrawl

Web scraping and crawling API that returns LLM-ready markdown

Stars
190.2k
License
AGPL-3.0
Last commit
Oct 2026
Last release
Jun 2026

Overview

API that turns URLs into markdown, HTML, screenshots or schema-based JSON, with endpoints for search, scrape, crawl, map, batch scrape, page interaction and a prompt-driven agent. Handles JS-rendered pages and parses hosted PDFs and DOCX. SDKs for Python, Node, Go, Java, Elixir, Rust and Ruby plus an MCP server and CLI, for teams feeding web content to RAG pipelines and agents.

Who it is for: Teams feeding web content to RAG pipelines and agents

Strengths

  • Seven SDKs plus CLI and MCP server; SDKs poll async crawl jobs automatically
  • Crawl, map and batch-scrape endpoints return job IDs for large sites
  • Scrape supports actions (click, scroll, write, wait) before extraction
  • Compose file at the repo root for self-hosting

Weaknesses

  • README is written around the hosted API and keys; self-hosting lives in separate docs
  • AGPL-3.0 license; network use of a modified version triggers source obligations
  • Agent endpoint runs the hosted spark-2 model, not a local LLM
  • Proxy rotation and anti-bot handling are hosted-service features

What it needs

  • no GPU
  • Compose

Also in Data and scraping for AI

See all 3
Also in Data and scraping for AI
RankProjectScore
1Crawl4AIPython crawler that turns pages into LLM-ready markdown, with a Docker API85.1k stars, Apache-2.072 out of 100
3Jina ReaderConverts any URL or search query into LLM-friendly markdown Live demo ↗ (opens in a new tab)12.1k stars, Apache-2.044 out of 100