22nd of 13 in Observability
Phoenix
LLM tracing, evals, datasets and prompt playground built on OpenTelemetry
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 11.8k
- License
- custom license
- Last commit
- Oct 2026
- Last release
- Oct 2026
- Language
- Python
Overview
Phoenix collects traces from LLM applications through OpenTelemetry/OpenInference instrumentation and shows them in a web UI. It also covers LLM-based evals, versioned datasets, experiments, prompt management and a prompt playground that can replay traced calls. It runs via pip, uvx, Docker or a Helm chart, and exposes a remote MCP endpoint at /mcp for coding agents.
Who it is for: Teams debugging and evaluating LLM apps who want self-hosted tracing
Strengths
- Install with pip or uvx and run `phoenix serve`; no separate setup shown
- Auto-instrumentation for LangGraph, LlamaIndex, CrewAI, DSPy, Vercel AI SDK and more
- Built-in MCP server at /mcp lets Claude Code and Cursor query traces
- Python and TypeScript packages for OTel, client and evals
Weaknesses
- License reported as NOASSERTION; terms need checking before commercial use
- Managed production workflows are pushed to the paid Arize AX product
- Azure template serves plain HTTP and needs a TLS proxy in front
- RAM, storage backend and default port not stated in the README excerpt
What it needs
- no GPU
- Docker + Compose
- Compose runs PostgreSQL
- Models: OpenAI, Anthropic, Google GenAI, AWS Bedrock, OpenRouter
Also in Observability
See all 13| Rank | Project | Score |
|---|---|---|
| 1 | LangfuseTracing, prompt management and evals for LLM apps on ClickHouse | 74 out of 100 |
| 3 | promptfooCLI for evaluating and red-teaming prompts, agents and RAG | 71 out of 100 |
| #4 | MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking | 70 out of 100 |
| #5 | OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end | 64 out of 100 |
| #6 | LatitudeAgent observability that groups failures and dispatches coding agents to fix them | 62 out of 100 |