Langfuse vs MLflow

Two of the top observability, side by side: score, setup, license, activity and what each review found.

11st of 13 in Observability

Langfuse

Tracing, prompt management and evals for LLM apps on ClickHouse

#44th of 13 in Observability

MLflow

Tracing, evals, prompt registry and AI gateway plus classic ML tracking

Langfuse vs MLflow: score parts and facts
What we compareLangfuseMLflow
Score parts, out of 100
Adoption85, widely used75, popular
Freshness100, active100, active
Maintenance87, healthy87, healthy
Easy to run50, easy33, some setup
Agent-ready30, minimal85, ready
Facts from GitHub and the README
Stars35.6k28.3k
Licensecustom license (read the license)Apache-2.0 (permissive)
Last commitOct 2026Oct 2026
Last releaseOct 2026Oct 2026
LanguageNot statedNot stated
DockerYesYes
GPUNot neededNot needed
arm64 or Apple SiliconNot statedNot stated

Langfuse

Ingests traces of LLM calls, retrieval and agent steps via Python and JS/TS SDKs or drop-in OpenAI, LangChain, LlamaIndex, LiteLLM and Vercel AI SDK integrations, then adds prompt versioning with caching, LLM-as-a-judge and code evaluators, datasets and a playground. Stores data in ClickHouse; deploys with docker compose, Helm on Kubernetes, or Terraform for AWS, Azure and GCP. For teams debugging and evaluating LLM apps.

Who it is for: Teams debugging and evaluating LLM apps

Strengths

  • Public OpenAPI spec, Postman collection and typed Python and JS/TS SDKs
  • Prompt management with server and client caching adds no request latency
  • Deployment paths from docker compose to Helm and Terraform templates
  • Integrations with Dify, Flowise, Langflow, OpenWebUI, LobeChat, CrewAI, smolagents

Weaknesses

  • MIT except the ee folders; enterprise features need a commercial license
  • Runs on ClickHouse plus other services; heavier than single-binary tools
  • Default compose inherits Docker json-file logging with no rotation; disk can fill
  • No Dockerfile at the repo root; images come from Docker Hub
  • no GPU
  • Compose
  • Needs ClickHouse

MLflow

Single mlflow server (port 5000) that records OpenTelemetry traces from 60+ frameworks via one-line autolog, runs evaluations with 50+ metrics and LLM judges, versions and optimizes prompts, and fronts providers through an OpenAI-compatible AI Gateway with rate limits, fallbacks and traffic splitting. Keeps the original experiment tracking, model registry and deployment tooling. For teams wanting one platform for GenAI and ML.

Who it is for: Teams wanting one platform for GenAI tracing and ML tracking

Strengths

  • One-line autolog for 60+ frameworks in Python, TypeScript and Java; MCP and OTel native
  • Starts with uvx mlflow server; no separate database needed to begin
  • AI Gateway adds credential management, guardrails and A/B traffic splitting
  • Setup wizard lets Claude Code, Codex or OpenCode add tracing to a project

Weaknesses

  • README covers the quickstart; production backend store and auth setup live in docs
  • Broad scope (ML tracking plus GenAI) means a large install and UI surface
  • No Dockerfile or compose file at the repo root
  • TypeScript and Java coverage is smaller than Python (5 TS and 2 Java frameworks listed)
  • no GPU
  • Docker + Compose
  • Models: any LLM provider via autolog or the AI Gateway
  • port 5000

More in Observability