promptfoo vs MLflow

Two of the top observability, side by side: score, setup, license, activity and what each review found.

33rd of 13 in Observability

promptfoo

CLI for evaluating and red-teaming prompts, agents and RAG

71 out of 100
#44th of 13 in Observability

MLflow

Tracing, evals, prompt registry and AI gateway plus classic ML tracking

promptfoo vs MLflow: score parts and facts
What we comparepromptfooMLflow
Score parts, out of 100
Adoption70, popular75, popular
Freshness100, active100, active
Maintenance82, healthy87, healthy
Easy to run50, easy33, some setup
Agent-ready45, minimal85, ready
Facts from GitHub and the README
Stars25.9k28.3k
LicenseMIT (permissive)Apache-2.0 (permissive)
Last commitOct 2026Oct 2026
Last releaseOct 2026Oct 2026
LanguageNot statedNot stated
DockerYesYes
GPUNot neededNot needed
arm64 or Apple SiliconNot statedNot stated

promptfoo

Runs prompt and model evaluations from a YAML config via promptfoo eval, compares providers side by side, and generates red-team vulnerability reports; promptfoo view opens a local web viewer. Installs with npm, Homebrew or pip, runs in CI/CD, and can scan pull requests for LLM security issues, for developers testing prompts and agents before release.

Who it is for: Developers testing prompts and agents before release

Strengths

  • Evals run locally; prompts stay on your machine
  • Red-team scans produce vulnerability reports alongside quality evals
  • Live reload and caching for fast iteration; npx usage needs no install
  • MIT licensed and still open source after joining OpenAI

Weaknesses

  • Primarily a CLI; the web viewer is a local results UI, not a multi-user server
  • Most providers require an API key; local use needs Ollama or similar
  • README is short; config syntax, assertions and providers are only in the docs
  • Dockerfile exists at the root but the README gives no Docker instructions
  • no GPU
  • Docker
  • Needs Node.js (npm) or Python (pip), LLM provider API key or Ollama
  • Models: OpenAI, Anthropic, Azure, Bedrock, Ollama

MLflow

Single mlflow server (port 5000) that records OpenTelemetry traces from 60+ frameworks via one-line autolog, runs evaluations with 50+ metrics and LLM judges, versions and optimizes prompts, and fronts providers through an OpenAI-compatible AI Gateway with rate limits, fallbacks and traffic splitting. Keeps the original experiment tracking, model registry and deployment tooling. For teams wanting one platform for GenAI and ML.

Who it is for: Teams wanting one platform for GenAI tracing and ML tracking

Strengths

  • One-line autolog for 60+ frameworks in Python, TypeScript and Java; MCP and OTel native
  • Starts with uvx mlflow server; no separate database needed to begin
  • AI Gateway adds credential management, guardrails and A/B traffic splitting
  • Setup wizard lets Claude Code, Codex or OpenCode add tracing to a project

Weaknesses

  • README covers the quickstart; production backend store and auth setup live in docs
  • Broad scope (ML tracking plus GenAI) means a large install and UI surface
  • No Dockerfile or compose file at the repo root
  • TypeScript and Java coverage is smaller than Python (5 TS and 2 Java frameworks listed)
  • no GPU
  • Docker + Compose
  • Models: any LLM provider via autolog or the AI Gateway
  • port 5000

More in Observability