11st of 13 in Observability
Langfuse
Tracing, prompt management and evals for LLM apps on ClickHouse
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 35.6k
- License
- custom license
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
Ingests traces of LLM calls, retrieval and agent steps via Python and JS/TS SDKs or drop-in OpenAI, LangChain, LlamaIndex, LiteLLM and Vercel AI SDK integrations, then adds prompt versioning with caching, LLM-as-a-judge and code evaluators, datasets and a playground. Stores data in ClickHouse; deploys with docker compose, Helm on Kubernetes, or Terraform for AWS, Azure and GCP. For teams debugging and evaluating LLM apps.
Who it is for: Teams debugging and evaluating LLM apps
Strengths
- Public OpenAPI spec, Postman collection and typed Python and JS/TS SDKs
- Prompt management with server and client caching adds no request latency
- Deployment paths from docker compose to Helm and Terraform templates
- Integrations with Dify, Flowise, Langflow, OpenWebUI, LobeChat, CrewAI, smolagents
Weaknesses
- MIT except the ee folders; enterprise features need a commercial license
- Runs on ClickHouse plus other services; heavier than single-binary tools
- Default compose inherits Docker json-file logging with no rotation; disk can fill
- No Dockerfile at the repo root; images come from Docker Hub
What it needs
- no GPU
- Compose
- Needs ClickHouse
Also in Observability
See all 13| Rank | Project | Score |
|---|---|---|
| 2 | PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry | 74 out of 100 |
| 3 | promptfooCLI for evaluating and red-teaming prompts, agents and RAG | 71 out of 100 |
| #4 | MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking | 70 out of 100 |
| #5 | OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end | 64 out of 100 |
| #6 | LatitudeAgent observability that groups failures and dispatches coding agents to fix them | 62 out of 100 |