33rd of 13 in Observability

promptfoo

CLI for evaluating and red-teaming prompts, agents and RAG

Stars
25.9k
License
MIT
Last commit
Oct 2026
Last release
Oct 2026

Overview

Runs prompt and model evaluations from a YAML config via promptfoo eval, compares providers side by side, and generates red-team vulnerability reports; promptfoo view opens a local web viewer. Installs with npm, Homebrew or pip, runs in CI/CD, and can scan pull requests for LLM security issues, for developers testing prompts and agents before release.

Who it is for: Developers testing prompts and agents before release

Strengths

  • Evals run locally; prompts stay on your machine
  • Red-team scans produce vulnerability reports alongside quality evals
  • Live reload and caching for fast iteration; npx usage needs no install
  • MIT licensed and still open source after joining OpenAI

Weaknesses

  • Primarily a CLI; the web viewer is a local results UI, not a multi-user server
  • Most providers require an API key; local use needs Ollama or similar
  • README is short; config syntax, assertions and providers are only in the docs
  • Dockerfile exists at the root but the README gives no Docker instructions

What it needs

  • no GPU
  • Docker
  • Needs Node.js (npm) or Python (pip), LLM provider API key or Ollama
  • Models: OpenAI, Anthropic, Azure, Bedrock, Ollama

Also in Observability

See all 13
Also in Observability
RankProjectScore
1LangfuseTracing, prompt management and evals for LLM apps on ClickHouse Live demo ↗ (opens in a new tab)35.6k stars, custom license74 out of 100
2PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry11.8k stars, custom license74 out of 100
#4MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking Live demo ↗ (opens in a new tab)28.3k stars, Apache-2.070 out of 100
#5OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end22.5k stars, Apache-2.064 out of 100
#6LatitudeAgent observability that groups failures and dispatches coding agents to fix them4.7k stars, MIT62 out of 100