Langfuse vs Phoenix
Two of the top observability, side by side: score, setup, license, activity and what each review found.
Langfuse
Tracing, prompt management and evals for LLM apps on ClickHouse
Phoenix
LLM tracing, evals, datasets and prompt playground built on OpenTelemetry
| What we compare | Langfuse | Phoenix |
|---|---|---|
| Score parts, out of 100 | ||
| Adoption | 85, widely used | 52, popular |
| Freshness | 100, active | 100, active |
| Maintenance | 87, healthy | 79, fair |
| Easy to run | 50, easy | 67, easy |
| Agent-ready | 30, minimal | 85, ready |
| Facts from GitHub and the README | ||
| Stars | 35.6k | 11.8k |
| License | custom license (read the license) | custom license (read the license) |
| Last commit | Oct 2026 | Oct 2026 |
| Last release | Oct 2026 | Oct 2026 |
| Language | Not stated | Python |
| Docker | Yes | Yes |
| GPU | Not needed | Not needed |
| arm64 or Apple Silicon | Not stated | Not stated |
Langfuse
Ingests traces of LLM calls, retrieval and agent steps via Python and JS/TS SDKs or drop-in OpenAI, LangChain, LlamaIndex, LiteLLM and Vercel AI SDK integrations, then adds prompt versioning with caching, LLM-as-a-judge and code evaluators, datasets and a playground. Stores data in ClickHouse; deploys with docker compose, Helm on Kubernetes, or Terraform for AWS, Azure and GCP. For teams debugging and evaluating LLM apps.
Who it is for: Teams debugging and evaluating LLM apps
Strengths
- Public OpenAPI spec, Postman collection and typed Python and JS/TS SDKs
- Prompt management with server and client caching adds no request latency
- Deployment paths from docker compose to Helm and Terraform templates
- Integrations with Dify, Flowise, Langflow, OpenWebUI, LobeChat, CrewAI, smolagents
Weaknesses
- MIT except the ee folders; enterprise features need a commercial license
- Runs on ClickHouse plus other services; heavier than single-binary tools
- Default compose inherits Docker json-file logging with no rotation; disk can fill
- No Dockerfile at the repo root; images come from Docker Hub
- no GPU
- Compose
- Needs ClickHouse
Phoenix
Phoenix collects traces from LLM applications through OpenTelemetry/OpenInference instrumentation and shows them in a web UI. It also covers LLM-based evals, versioned datasets, experiments, prompt management and a prompt playground that can replay traced calls. It runs via pip, uvx, Docker or a Helm chart, and exposes a remote MCP endpoint at /mcp for coding agents.
Who it is for: Teams debugging and evaluating LLM apps who want self-hosted tracing
Strengths
- Install with pip or uvx and run `phoenix serve`; no separate setup shown
- Auto-instrumentation for LangGraph, LlamaIndex, CrewAI, DSPy, Vercel AI SDK and more
- Built-in MCP server at /mcp lets Claude Code and Cursor query traces
- Python and TypeScript packages for OTel, client and evals
Weaknesses
- License reported as NOASSERTION; terms need checking before commercial use
- Managed production workflows are pushed to the paid Arize AX product
- Azure template serves plain HTTP and needs a TLS proxy in front
- RAM, storage backend and default port not stated in the README excerpt
- no GPU
- Docker + Compose
- Compose runs PostgreSQL
- Models: OpenAI, Anthropic, Google GenAI, AWS Bedrock, OpenRouter