Langfuse vs Phoenix

Two of the top observability, side by side: score, setup, license, activity and what each review found.

11st of 13 in Observability

Langfuse

Tracing, prompt management and evals for LLM apps on ClickHouse

22nd of 13 in Observability

Phoenix

LLM tracing, evals, datasets and prompt playground built on OpenTelemetry

74 out of 100
Langfuse vs Phoenix: score parts and facts
What we compareLangfusePhoenix
Score parts, out of 100
Adoption85, widely used52, popular
Freshness100, active100, active
Maintenance87, healthy79, fair
Easy to run50, easy67, easy
Agent-ready30, minimal85, ready
Facts from GitHub and the README
Stars35.6k11.8k
Licensecustom license (read the license)custom license (read the license)
Last commitOct 2026Oct 2026
Last releaseOct 2026Oct 2026
LanguageNot statedPython
DockerYesYes
GPUNot neededNot needed
arm64 or Apple SiliconNot statedNot stated

Langfuse

Ingests traces of LLM calls, retrieval and agent steps via Python and JS/TS SDKs or drop-in OpenAI, LangChain, LlamaIndex, LiteLLM and Vercel AI SDK integrations, then adds prompt versioning with caching, LLM-as-a-judge and code evaluators, datasets and a playground. Stores data in ClickHouse; deploys with docker compose, Helm on Kubernetes, or Terraform for AWS, Azure and GCP. For teams debugging and evaluating LLM apps.

Who it is for: Teams debugging and evaluating LLM apps

Strengths

  • Public OpenAPI spec, Postman collection and typed Python and JS/TS SDKs
  • Prompt management with server and client caching adds no request latency
  • Deployment paths from docker compose to Helm and Terraform templates
  • Integrations with Dify, Flowise, Langflow, OpenWebUI, LobeChat, CrewAI, smolagents

Weaknesses

  • MIT except the ee folders; enterprise features need a commercial license
  • Runs on ClickHouse plus other services; heavier than single-binary tools
  • Default compose inherits Docker json-file logging with no rotation; disk can fill
  • No Dockerfile at the repo root; images come from Docker Hub
  • no GPU
  • Compose
  • Needs ClickHouse

Phoenix

Phoenix collects traces from LLM applications through OpenTelemetry/OpenInference instrumentation and shows them in a web UI. It also covers LLM-based evals, versioned datasets, experiments, prompt management and a prompt playground that can replay traced calls. It runs via pip, uvx, Docker or a Helm chart, and exposes a remote MCP endpoint at /mcp for coding agents.

Who it is for: Teams debugging and evaluating LLM apps who want self-hosted tracing

Strengths

  • Install with pip or uvx and run `phoenix serve`; no separate setup shown
  • Auto-instrumentation for LangGraph, LlamaIndex, CrewAI, DSPy, Vercel AI SDK and more
  • Built-in MCP server at /mcp lets Claude Code and Cursor query traces
  • Python and TypeScript packages for OTel, client and evals

Weaknesses

  • License reported as NOASSERTION; terms need checking before commercial use
  • Managed production workflows are pushed to the paid Arize AX product
  • Azure template serves plain HTTP and needs a TLS proxy in front
  • RAM, storage backend and default port not stated in the README excerpt
  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL
  • Models: OpenAI, Anthropic, Google GenAI, AWS Bedrock, OpenRouter

More in Observability