promptfoo
promptfoo is an open-source tool for testing and evaluating LLM applications. It helps you systematically test prompts, models, and RAG pipelines to catch regressions, compare outputs, and ensure quality before deploying to production.
Key Features
- Prompt Testing: Define test cases in YAML and run them against any LLM provider
- Model Comparison: Side-by-side evaluation of different models and prompt variants
- Assertions: Built-in assertion types for factuality, relevance, toxicity, and custom metrics
- RAG Evaluation: Test retrieval quality, context relevance, and end-to-end RAG pipelines
- Red Teaming: Automated adversarial testing to find prompt injection vulnerabilities
- CI/CD Integration: Run evaluations in your CI pipeline to catch regressions
- Web UI: Interactive results viewer for exploring and comparing outputs
- Provider Agnostic: Works with OpenAI, Anthropic, Google, Azure, local models, and custom endpoints
Installation
npm install -g promptfoo
# or
npx promptfoo@latest init
Quick Start
1. Create a test configuration
# promptfooconfig.yaml
prompts:
- "Summarize the following text: {{text}}"
- "Write a concise summary of: {{text}}"
providers:
- openai:gpt-4
- anthropic:messages:claude-3-sonnet-20240229
tests:
- vars:
text: "The quick brown fox jumps over the lazy dog."
assert:
- type: contains
value: "fox"
- type: llm-rubric
value: "The summary captures the main action"
2. Run the evaluation
promptfoo eval
3. View results
promptfoo view
Use Cases
- Prompt Engineering: Compare prompt variants systematically instead of manual testing
- Model Migration: Evaluate a new model against your existing test suite before switching
- Regression Testing: Catch output quality regressions in CI/CD pipelines
- Red Teaming: Automatically test for prompt injection, jailbreaks, and harmful outputs
- RAG Quality: Measure retrieval precision, context relevance, and answer accuracy
Documentation
License
MIT License