#55th of 13 in Observability

Opik

Trace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end

Stars
22.5k
License
Apache-2.0
Last commit
Oct 2026
Last release
Oct 2026

Overview

Logs trace trees for LLM calls, tool executions and agent steps via Python and TypeScript SDKs, OpenTelemetry or framework integrations, then runs datasets, experiments and LLM-as-a-judge metrics for hallucination, moderation and RAG quality, with online evaluation rules in production. Self-hosts with ./opik.sh (Docker Compose, UI on port 5173) or a Helm chart. For ML engineers moving agents to production.

Who it is for: ML engineers moving LLM agents to production

Strengths

  • Full platform (backend, web app, evals, prompt management) under Apache-2.0
  • Designed for 40M+ traces per day; online evaluation rules on production traffic
  • PyTest integration gates LLM pipelines in CI
  • MCP server lets Claude Code, Cursor, Codex or opencode query traces and run evals

Weaknesses

  • No Dockerfile or compose file at the repo root; install goes through opik.sh
  • Multi-service stack (databases, caches, backend, frontend); not a single binary
  • Guardrails and the optimizer are separate profiles and SDKs to enable
  • README is heavy with Comet Cloud links and UTM tracking

What it needs

  • no GPU
  • Compose
  • Models: any LLM via SDK, OpenTelemetry or framework integrations (Google ADK, AG2, Autogen, Flowise)
  • port 5173

Also in Observability

See all 13
Also in Observability
RankProjectScore
1LangfuseTracing, prompt management and evals for LLM apps on ClickHouse Live demo ↗ (opens in a new tab)35.6k stars, custom license74 out of 100
2PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry11.8k stars, custom license74 out of 100
3promptfooCLI for evaluating and red-teaming prompts, agents and RAG25.9k stars, MIT71 out of 100
#4MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking Live demo ↗ (opens in a new tab)28.3k stars, Apache-2.070 out of 100
#6LatitudeAgent observability that groups failures and dispatches coding agents to fix them4.7k stars, MIT62 out of 100