33rd of 3 in Data and scraping for AI
Jina Reader
Converts any URL or search query into LLM-friendly markdown
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 12.1k
- License
- Apache-2.0
- Last commit
- May 2026
Overview
Open-source branch of the service behind r.jina.ai and s.jina.ai: fetches a page with headless Chrome or curl-impersonate, parses PDFs and Office files, and returns markdown, text, HTML, screenshots or JSON controlled by request headers (engine, timeout, token limits). The ghcr.io image bundles Chrome, LibreOffice and CJK fonts, serves HTTP/1.1 on 8081 and h2c on 8080, and runs stateless or with S3-compatible caching.
Who it is for: Developers feeding web pages and documents to LLMs
Strengths
- Prebuilt image with Chrome, LibreOffice and CJK fonts; stateless by default
- Fine-grained headers: x-respond-timing, x-max-tokens, x-token-budget, x-target-selector
- Optional VLM captions for images without alt text
- Semantic markdown chunking by heading or block level
Weaknesses
- Hosted proxy pool, rate limiting and MongoDB storage layer are not in the OSS branch
- Needs non-redistributable assets (MaxMind GeoLite2, Source Han Sans) fetched at build
- Last commit May 2026; the SaaS resync was April 2026
- Default h2c port 8080 needs --http2-prior-knowledge from curl; use 8081 otherwise
What it needs
- no GPU
- Docker + Compose
- Needs Headless Chrome and LibreOffice (bundled in image), S3-compatible bucket (optional cache), VLM endpoint for image captions (optional)
- port 8081