Observability

Tracing, evaluation and prompt management for LLM applications.

Langfuse leads with 74, ahead of Phoenix (74) and promptfoo (71). 13 projects ranked by score.

The ranking

Observability: full ranking
RankProjectAdoptionFreshnessMaintenanceEasy to runAgent-readyScore
1LangfuseTracing, prompt management and evals for LLM apps on ClickHouse Live demo ↗ (opens in a new tab)35.6k stars, custom license, last commit Oct 20268510087503074 out of 100
2PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry11.8k stars, custom license, last commit Oct 20265210079678574 out of 100
3promptfooCLI for evaluating and red-teaming prompts, agents and RAG25.9k stars, MIT, last commit Oct 20267010083504571 out of 100
#4MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking Live demo ↗ (opens in a new tab)28.3k stars, Apache-2.0, last commit Oct 20267510087338570 out of 100
#5OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end22.5k stars, Apache-2.0, last commit Oct 20266310089333064 out of 100
#6LatitudeAgent observability that groups failures and dispatches coding agents to fix them4.7k stars, MIT, last commit Oct 20262510099503062 out of 100
#7Future AGITracing, evals, simulation, guardrails and an LLM gateway for AI agents Live demo ↗ (opens in a new tab)2.1k stars, Apache-2.0, last commit Oct 2026210072831561 out of 100
#8LangWatchAgent observability, simulation testing, AI gateway and governance in one4.9k stars, Apache-2.0, last commit Oct 20263510096335560 out of 100
#9LaminarRust-based agent tracing with SQL queries, signals and evals3.4k stars, Apache-2.0, last commit Sep 20261810076504057 out of 100
#10OpenLITOpenTelemetry-based tracing, evals and guardrails for LLM apps and coding agents2.8k stars, Apache-2.0, last commit Oct 2026810077507057 out of 100
#11AgentaTeam workspace for building chat-driven agents that run in Slack and WhatsApp4.8k stars, custom license, last commit Oct 2026301008808550 out of 100
#12HeliconeLLM proxy gateway with request logging, cost tracking and sessions Live demo ↗ (opens in a new tab)6.2k stars, Apache-2.0, last commit Sep 2026428033337049 out of 100
#13PezzoPrompt management, observability and caching for LLM apps3.3k stars, Apache-2.0, last commit Aug 202613702433033 out of 100

Momentum, Verified build, Docs and Privacy are not measured yet; their weight goes to the signals shown. A dash means the signal is not scored for that kind of project. Hover a number for its rating in words.

Reviews

174 out of 100

Langfuse

Tracing, prompt management and evals for LLM apps on ClickHouse

Live demo ↗ (opens in a new tab)35.6k stars, custom license, last commit Oct 2026

Ingests traces of LLM calls, retrieval and agent steps via Python and JS/TS SDKs or drop-in OpenAI, LangChain, LlamaIndex, LiteLLM and Vercel AI SDK integrations, then adds prompt versioning with caching, LLM-as-a-judge and code evaluators, datasets and a playground. Stores data in ClickHouse; deploys with docker compose, Helm on Kubernetes, or Terraform for AWS, Azure and GCP. For teams debugging and evaluating LLM apps.

Strengths

  • Public OpenAPI spec, Postman collection and typed Python and JS/TS SDKs
  • Prompt management with server and client caching adds no request latency
  • Deployment paths from docker compose to Helm and Terraform templates
  • Integrations with Dify, Flowise, Langflow, OpenWebUI, LobeChat, CrewAI, smolagents

Weaknesses

  • MIT except the ee folders; enterprise features need a commercial license
  • Runs on ClickHouse plus other services; heavier than single-binary tools
  • Default compose inherits Docker json-file logging with no rotation; disk can fill
  • No Dockerfile at the repo root; images come from Docker Hub
  • no GPU
  • Compose
  • Needs ClickHouse
274 out of 100

Phoenix

LLM tracing, evals, datasets and prompt playground built on OpenTelemetry

11.8k stars, custom license, last commit Oct 2026

Phoenix collects traces from LLM applications through OpenTelemetry/OpenInference instrumentation and shows them in a web UI. It also covers LLM-based evals, versioned datasets, experiments, prompt management and a prompt playground that can replay traced calls. It runs via pip, uvx, Docker or a Helm chart, and exposes a remote MCP endpoint at /mcp for coding agents.

Strengths

  • Install with pip or uvx and run `phoenix serve`; no separate setup shown
  • Auto-instrumentation for LangGraph, LlamaIndex, CrewAI, DSPy, Vercel AI SDK and more
  • Built-in MCP server at /mcp lets Claude Code and Cursor query traces
  • Python and TypeScript packages for OTel, client and evals

Weaknesses

  • License reported as NOASSERTION; terms need checking before commercial use
  • Managed production workflows are pushed to the paid Arize AX product
  • Azure template serves plain HTTP and needs a TLS proxy in front
  • RAM, storage backend and default port not stated in the README excerpt
  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL
  • Models: OpenAI, Anthropic, Google GenAI, AWS Bedrock, OpenRouter
371 out of 100

promptfoo

CLI for evaluating and red-teaming prompts, agents and RAG

25.9k stars, MIT, last commit Oct 2026

Runs prompt and model evaluations from a YAML config via promptfoo eval, compares providers side by side, and generates red-team vulnerability reports; promptfoo view opens a local web viewer. Installs with npm, Homebrew or pip, runs in CI/CD, and can scan pull requests for LLM security issues, for developers testing prompts and agents before release.

Strengths

  • Evals run locally; prompts stay on your machine
  • Red-team scans produce vulnerability reports alongside quality evals
  • Live reload and caching for fast iteration; npx usage needs no install
  • MIT licensed and still open source after joining OpenAI

Weaknesses

  • Primarily a CLI; the web viewer is a local results UI, not a multi-user server
  • Most providers require an API key; local use needs Ollama or similar
  • README is short; config syntax, assertions and providers are only in the docs
  • Dockerfile exists at the root but the README gives no Docker instructions
  • no GPU
  • Docker
  • Needs Node.js (npm) or Python (pip), LLM provider API key or Ollama
  • Models: OpenAI, Anthropic, Azure, Bedrock, Ollama
#470 out of 100

MLflow

Tracing, evals, prompt registry and AI gateway plus classic ML tracking

Live demo ↗ (opens in a new tab)28.3k stars, Apache-2.0, last commit Oct 2026

Single mlflow server (port 5000) that records OpenTelemetry traces from 60+ frameworks via one-line autolog, runs evaluations with 50+ metrics and LLM judges, versions and optimizes prompts, and fronts providers through an OpenAI-compatible AI Gateway with rate limits, fallbacks and traffic splitting. Keeps the original experiment tracking, model registry and deployment tooling. For teams wanting one platform for GenAI and ML.

Strengths

  • One-line autolog for 60+ frameworks in Python, TypeScript and Java; MCP and OTel native
  • Starts with uvx mlflow server; no separate database needed to begin
  • AI Gateway adds credential management, guardrails and A/B traffic splitting
  • Setup wizard lets Claude Code, Codex or OpenCode add tracing to a project

Weaknesses

  • README covers the quickstart; production backend store and auth setup live in docs
  • Broad scope (ML tracking plus GenAI) means a large install and UI surface
  • No Dockerfile or compose file at the repo root
  • TypeScript and Java coverage is smaller than Python (5 TS and 2 Java frameworks listed)
  • no GPU
  • Docker + Compose
  • Models: any LLM provider via autolog or the AI Gateway
  • port 5000
#564 out of 100

Opik

Trace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end

22.5k stars, Apache-2.0, last commit Oct 2026

Logs trace trees for LLM calls, tool executions and agent steps via Python and TypeScript SDKs, OpenTelemetry or framework integrations, then runs datasets, experiments and LLM-as-a-judge metrics for hallucination, moderation and RAG quality, with online evaluation rules in production. Self-hosts with ./opik.sh (Docker Compose, UI on port 5173) or a Helm chart. For ML engineers moving agents to production.

Strengths

  • Full platform (backend, web app, evals, prompt management) under Apache-2.0
  • Designed for 40M+ traces per day; online evaluation rules on production traffic
  • PyTest integration gates LLM pipelines in CI
  • MCP server lets Claude Code, Cursor, Codex or opencode query traces and run evals

Weaknesses

  • No Dockerfile or compose file at the repo root; install goes through opik.sh
  • Multi-service stack (databases, caches, backend, frontend); not a single binary
  • Guardrails and the optimizer are separate profiles and SDKs to enable
  • README is heavy with Comet Cloud links and UTM tracking
  • no GPU
  • Compose
  • Models: any LLM via SDK, OpenTelemetry or framework integrations (Google ADK, AG2, Autogen, Flowise)
  • port 5173
#662 out of 100

Latitude

Agent observability that groups failures and dispatches coding agents to fix them

4.7k stars, MIT, last commit Oct 2026

Captures traces, sessions and tool calls via a one-line SDK (TypeScript, Python) or OpenTelemetry, groups failing traces into tracked signals, then dispatches Claude Code or Cursor with those traces to open a fix PR and replays fixes against regression datasets. The UI is also reachable from an MCP server and CLI; self-hosts from Docker Hub images via Compose or Helm. For teams operating agents in production.

Strengths

  • Signals auto-group failing traces with status, size and trend
  • Agent Dispatch sends sample traces to Claude Code or Cursor via Linear or webhooks
  • Regression datasets replay fixes against the real failing traces
  • MIT license; Compose and Helm paths plus Railway one-click

Weaknesses

  • README quickstart targets the cloud; self-host steps are in external docs
  • Automatic fixing depends on third-party coding agents and their subscriptions
  • Claude Code session capture is a separate telemetry package
  • Storage and service requirements are not stated in the README
  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL, ClickHouse, Redis
  • Models: OpenAI, Anthropic, Bedrock, Vercel AI SDK and LangChain apps, any OpenTelemetry source
#761 out of 100

Future AGI

Tracing, evals, simulation, guardrails and an LLM gateway for AI agents

Live demo ↗ (opens in a new tab)2.1k stars, Apache-2.0, last commit Oct 2026

Future AGI is a Django and Go platform that traces agents over OpenTelemetry, scores outputs with 50+ evaluators, simulates multi-turn conversations, and applies guardrail scanners. It also ships an OpenAI-compatible gateway with routing, caching and virtual keys, plus prompt-optimization algorithms. The default install runs one app container with Postgres and ClickHouse; a distributed Compose setup and a Helm chart cover larger deployments.

Strengths

  • OTel tracing with instrumentors for 50+ frameworks in Python, TypeScript, Java and C#
  • Gateway is OpenAI-compatible with 100+ providers, semantic caching and virtual keys
  • Standalone install is one command and needs 2 vCPUs and 4 GB for Docker
  • Apache-2.0 core; Compose, production overlay, Helm and air-gapped modes documented

Weaknesses

  • No supported path to move Standalone data to Distributed or Helm later
  • Distributed setup needs 4+ vCPUs and 12-16 GB, plus Kafka and PeerDB
  • Many components (Postgres, ClickHouse, Redis, Temporal) make it heavy to operate
  • Benchmark figures come from the README; independent verification is unknown
  • RAM ≥ 4 GB
  • no GPU
  • Docker + Compose
  • Needs PostgreSQL, ClickHouse, Redis, Temporal, Docker Compose v2.24+
  • Models: OpenAI-compatible providers (100+ via gateway)
  • port 3000
#860 out of 100

LangWatch

Agent observability, simulation testing, AI gateway and governance in one

4.9k stars, Apache-2.0, last commit Oct 2026

Traces LLM and agent calls through OpenTelemetry and SDK integrations, runs simulation-based agent tests and evaluations, manages prompts, and adds an OpenAI- and Anthropic-compatible gateway with virtual keys and budgets. Also tracks coding-agent sessions (Claude Code, Codex, Copilot) with cost per pull request, and starts locally with npx @langwatch/server. For platform teams governing AI use across a company.

Strengths

  • npx @langwatch/server starts a local instance with only Node.js installed
  • Coding-agent tracking: sessions and cost per PR for Claude Code, Codex, Copilot
  • Gateway virtual keys with budgets for customers or employees
  • Governance ingests Copilot Studio, Claude and OpenAI compliance APIs, Workato, S3 audit feeds

Weaknesses

  • Open-core: modules under platform/app/ee need a commercial license in production
  • No Dockerfile or compose file at the repo root; production setup is in external docs
  • README is a feature index; architecture and storage needs are not described
  • Cloud signup is the first call to action; self-host gets one line
  • no GPU
  • Needs Node.js
  • Models: OpenAI, Anthropic, Azure OpenAI, Vertex AI, Bedrock
#957 out of 100

Laminar

Rust-based agent tracing with SQL queries, signals and evals

3.4k stars, Apache-2.0, last commit Sep 2026

OpenTelemetry-native tracing for Vercel AI SDK, LangChain, OpenAI, Anthropic, Gemini and more with one line of SDK code, stored in ClickHouse and queried with SQL from the UI, MCP server or CLI. Signals watch every run for behaviors described in plain English and ping Slack; evals run from an SDK and CLI. docker compose up serves the UI on port 5667, for teams debugging browser and tool-using agents.

Strengths

  • Signals: describe a failure in plain English and get a Slack ping when it occurs
  • SQL over traces, spans, metrics and events, also from your coding agent via MCP
  • Rust backend with 20x trace compression and a realtime trace viewer
  • Custom Postgres schema support for shared database deployments

Weaknesses

  • Anonymous usage telemetry is on by default; LAMINAR_TELEMETRY_DISABLED=true opts out
  • Production is steered to the managed platform or the heavier docker-compose-full stack
  • AI features (chat-with-trace, SQL-with-AI) need a configured LLM provider
  • ClickHouse upgrades need manual container recreation and log-table truncation
  • no GPU
  • Compose
  • Needs ClickHouse, PostgreSQL, LLM provider (optional, for AI features)
  • Models: Gemini, OpenAI and OpenAI-compatible gateways (LiteLLM, OpenRouter, vLLM), AWS Bedrock, Azure AI Foundry
  • port 5667
#1057 out of 100

OpenLIT

OpenTelemetry-based tracing, evals and guardrails for LLM apps and coding agents

2.8k stars, Apache-2.0, last commit Oct 2026

OpenLIT collects OpenTelemetry traces and metrics from LLM apps and agents through Python and TypeScript SDKs, and stores them in ClickHouse behind a web dashboard on port 3000. It adds cost tracking, LLM-as-a-judge evals, SDK guardrails, a versioned Prompt Hub, a Vault for API keys, and GPU monitoring. A CLI installs tracing for Claude Code, Cursor and Codex sessions.

Strengths

  • Follows OpenTelemetry GenAI conventions; accepts OTLP on :4317 (gRPC) and :4318 (HTTP)
  • One-line auto-instrumentation via openlit.init(); README claims 70+ integrations
  • Covers cost tracking, evals, guardrails, prompt versioning and secrets in one tool
  • Docker Compose quickstart; Apache-2.0 and free to self-host

Weaknesses

  • Needs ClickHouse as the telemetry store; RAM and disk requirements are not stated
  • Broad scope (Vault, Rule Engine, OpenGround) means a larger surface than a tracing-only tool
  • Guardrails run in the SDK, so they only protect instrumented code
  • No Dockerfile detected by our tools; the README only shows Compose
  • GPU optional
  • Compose
  • Needs ClickHouse
  • Models: OpenAI, Anthropic, Ollama, vLLM, Amazon Bedrock
  • port 3000
#1150 out of 100

Agenta

Team workspace for building chat-driven agents that run in Slack and WhatsApp

4.8k stars, custom license, last commit Oct 2026

Lets teams create agents by describing work in chat, connect tools through MCP or Composio, set per-agent read or write permissions, and talk to them from the web app, Slack, Telegram or WhatsApp. Agents keep memory and skills, run on schedules or events, and each session gets a sandbox with a browser and filesystem; every run is traced and costed. Runs Claude Code, Pi or Codex harnesses on API models, Ollama or a Claude or ChatGPT subscription.

Strengths

  • Runs on an existing Claude or ChatGPT subscription instead of metered API billing
  • Per-agent tool permissions with read or write scopes and human-in-the-loop gates
  • Every run traced and cost-tracked; configurations and versions are visible
  • Agents reachable from Slack, Telegram and WhatsApp Business

Weaknesses

  • README no longer covers the earlier prompt-management and evaluation product
  • Self-host instructions are delegated to an agent skill, not written out
  • Harness support limited to Claude Code, Pi and Codex today
  • No Dockerfile or compose file at the repo root
  • no GPU
  • Needs Claude Code, Pi or Codex harness, LLM API, Ollama, or a Claude or ChatGPT subscription, Composio (optional, 1,000+ app integrations)
  • Models: hosted models via API, Ollama, Claude and ChatGPT subscriptions
#1249 out of 100

Helicone

LLM proxy gateway with request logging, cost tracking and sessions

Live demo ↗ (opens in a new tab)6.2k stars, Apache-2.0, last commit Sep 2026

Sits as an OpenAI-compatible gateway in front of 100+ models with routing and automatic fallbacks, logging every request with cost, latency and session traces, plus a playground and prompt versioning. Self-hosts via a compose script that runs six services: web app, Jawn log server, Workers proxy, Supabase, ClickHouse and MinIO. For engineers who want observability by swapping an endpoint.

Strengths

  • One-line integration: point the OpenAI SDK baseURL at the gateway
  • Async logging path via OpenLLMetry for apps that cannot proxy
  • Open LLM cost database covering 300+ models; MCP server for data export
  • Apache-2.0; self-host compose script included

Weaknesses

  • Six-service stack including Supabase, ClickHouse, MinIO and a Cloudflare Workers proxy
  • Production Helm chart is enterprise only, by contacting sales
  • Manual deployment is explicitly not recommended
  • README quickstart is cloud-first; self-hosting details are in external docs
  • no GPU
  • Docker + Compose
  • Needs Supabase (database and auth), ClickHouse, MinIO, Cloudflare Workers runtime (proxy)
  • Models: 100+ providers via gateway (OpenAI, Azure, Anthropic, Bedrock, Gemini, Groq, Together, Fireworks, Ollama)
#1333 out of 100

Pezzo

Prompt management, observability and caching for LLM apps

3.3k stars, Apache-2.0, last commit Aug 2026

Stores and versions prompts, logs requests with cost and latency, and caches LLM responses, exposed through Node.js and Python clients and a LangChain integration. Runs on PostgreSQL, ClickHouse, Redis and SuperTokens via Docker Compose, with a GraphQL API server and a console UI. For small teams that want prompt delivery without code changes.

Strengths

  • Prompts delivered from the console without redeploying application code
  • Built-in response caching to cut repeated-call cost and latency
  • Node.js and Python clients plus LangChain support
  • Apache-2.0; infra is all open source (PostgreSQL, ClickHouse, Redis, SuperTokens)

Weaknesses

  • Last commit August 2026 with no release notes in the README
  • Four backing services for a modest feature set
  • README is thin; features are shown as screenshots, details only in docs
  • No evaluation or dataset features mentioned
  • no GPU
  • Compose
  • Needs PostgreSQL, ClickHouse, Redis, SuperTokens, Node.js 18+
  • port 4200

Written from each project's README and checked facts. Spot something wrong? Report it on GitHub (opens in a new tab).