22nd of 13 in Observability

Phoenix

LLM tracing, evals, datasets and prompt playground built on OpenTelemetry

Stars
11.8k
License
custom license
Last commit
Oct 2026
Last release
Oct 2026
Language
Python

Overview

Phoenix collects traces from LLM applications through OpenTelemetry/OpenInference instrumentation and shows them in a web UI. It also covers LLM-based evals, versioned datasets, experiments, prompt management and a prompt playground that can replay traced calls. It runs via pip, uvx, Docker or a Helm chart, and exposes a remote MCP endpoint at /mcp for coding agents.

Who it is for: Teams debugging and evaluating LLM apps who want self-hosted tracing

Strengths

  • Install with pip or uvx and run `phoenix serve`; no separate setup shown
  • Auto-instrumentation for LangGraph, LlamaIndex, CrewAI, DSPy, Vercel AI SDK and more
  • Built-in MCP server at /mcp lets Claude Code and Cursor query traces
  • Python and TypeScript packages for OTel, client and evals

Weaknesses

  • License reported as NOASSERTION; terms need checking before commercial use
  • Managed production workflows are pushed to the paid Arize AX product
  • Azure template serves plain HTTP and needs a TLS proxy in front
  • RAM, storage backend and default port not stated in the README excerpt

What it needs

  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL
  • Models: OpenAI, Anthropic, Google GenAI, AWS Bedrock, OpenRouter

Also in Observability

See all 13
Also in Observability
RankProjectScore
1LangfuseTracing, prompt management and evals for LLM apps on ClickHouse Live demo ↗ (opens in a new tab)35.6k stars, custom license74 out of 100
3promptfooCLI for evaluating and red-teaming prompts, agents and RAG25.9k stars, MIT71 out of 100
#4MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking Live demo ↗ (opens in a new tab)28.3k stars, Apache-2.070 out of 100
#5OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end22.5k stars, Apache-2.064 out of 100
#6LatitudeAgent observability that groups failures and dispatches coding agents to fix them4.7k stars, MIT62 out of 100