33rd of 13 in Observability
promptfoo
CLI for evaluating and red-teaming prompts, agents and RAG
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 25.9k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
Runs prompt and model evaluations from a YAML config via promptfoo eval, compares providers side by side, and generates red-team vulnerability reports; promptfoo view opens a local web viewer. Installs with npm, Homebrew or pip, runs in CI/CD, and can scan pull requests for LLM security issues, for developers testing prompts and agents before release.
Who it is for: Developers testing prompts and agents before release
Strengths
- Evals run locally; prompts stay on your machine
- Red-team scans produce vulnerability reports alongside quality evals
- Live reload and caching for fast iteration; npx usage needs no install
- MIT licensed and still open source after joining OpenAI
Weaknesses
- Primarily a CLI; the web viewer is a local results UI, not a multi-user server
- Most providers require an API key; local use needs Ollama or similar
- README is short; config syntax, assertions and providers are only in the docs
- Dockerfile exists at the root but the README gives no Docker instructions
What it needs
- no GPU
- Docker
- Needs Node.js (npm) or Python (pip), LLM provider API key or Ollama
- Models: OpenAI, Anthropic, Azure, Bedrock, Ollama
Also in Observability
See all 13| Rank | Project | Score |
|---|---|---|
| 1 | LangfuseTracing, prompt management and evals for LLM apps on ClickHouse | 74 out of 100 |
| 2 | PhoenixLLM tracing, evals, datasets and prompt playground built on OpenTelemetry | 74 out of 100 |
| #4 | MLflowTracing, evals, prompt registry and AI gateway plus classic ML tracking | 70 out of 100 |
| #5 | OpikTrace, evaluate and monitor LLM apps and agents, Apache-2.0 end to end | 64 out of 100 |
| #6 | LatitudeAgent observability that groups failures and dispatches coding agents to fix them | 62 out of 100 |