#44th of 13 in Observability
MLflow
Tracing, evals, prompt registry and AI gateway plus classic ML tracking
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 28.3k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
Single mlflow server (port 5000) that records OpenTelemetry traces from 60+ frameworks via one-line autolog, runs evaluations with 50+ metrics and LLM judges, versions and optimizes prompts, and fronts providers through an OpenAI-compatible AI Gateway with rate limits, fallbacks and traffic splitting. Keeps the original experiment tracking, model registry and deployment tooling. For teams wanting one platform for GenAI and ML.
Who it is for: Teams wanting one platform for GenAI tracing and ML tracking
Strengths
- One-line autolog for 60+ frameworks in Python, TypeScript and Java; MCP and OTel native
- Starts with uvx mlflow server; no separate database needed to begin
- AI Gateway adds credential management, guardrails and A/B traffic splitting
- Setup wizard lets Claude Code, Codex or OpenCode add tracing to a project
Weaknesses
- README covers the quickstart; production backend store and auth setup live in docs
- Broad scope (ML tracking plus GenAI) means a large install and UI surface
- No Dockerfile or compose file at the repo root
- TypeScript and Java coverage is smaller than Python (5 TS and 2 Java frameworks listed)
What it needs
- no GPU
- Docker + Compose
- Models: any LLM provider via autolog or the AI Gateway
- port 5000
Also in Observability
See all 13| Rank | Project | Score |
|---|---|---|
| 1 | LangfuseTracing, prompt management and evals for LLM apps on ClickHouse | 74 out of 100 |
| 2 | PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry | 74 out of 100 |
| 3 | promptfooCLI for evaluating and red-teaming prompts, agents and RAG | 71 out of 100 |
| #5 | OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end | 64 out of 100 |
| #6 | LatitudeAgent observability that groups failures and dispatches coding agents to fix them | 62 out of 100 |