Skip to main content
EVOKORE// SKILLS / README
promptfoo/promptfoomarkdownsha 0EAE10444FF0

README

project overview

promptfoo

promptfoo is an open-source tool for testing and evaluating LLM applications. It helps you systematically test prompts, models, and RAG pipelines to catch regressions, compare outputs, and ensure quality before deploying to production.

Key Features

  • Prompt Testing: Define test cases in YAML and run them against any LLM provider
  • Model Comparison: Side-by-side evaluation of different models and prompt variants
  • Assertions: Built-in assertion types for factuality, relevance, toxicity, and custom metrics
  • RAG Evaluation: Test retrieval quality, context relevance, and end-to-end RAG pipelines
  • Red Teaming: Automated adversarial testing to find prompt injection vulnerabilities
  • CI/CD Integration: Run evaluations in your CI pipeline to catch regressions
  • Web UI: Interactive results viewer for exploring and comparing outputs
  • Provider Agnostic: Works with OpenAI, Anthropic, Google, Azure, local models, and custom endpoints

Installation

npm install -g promptfoo
# or
npx promptfoo@latest init

Quick Start

1. Create a test configuration

# promptfooconfig.yaml
prompts:
  - "Summarize the following text: {{text}}"
  - "Write a concise summary of: {{text}}"

providers:
  - openai:gpt-4
  - anthropic:messages:claude-3-sonnet-20240229

tests:
  - vars:
      text: "The quick brown fox jumps over the lazy dog."
    assert:
      - type: contains
        value: "fox"
      - type: llm-rubric
        value: "The summary captures the main action"

2. Run the evaluation

promptfoo eval

3. View results

promptfoo view

Use Cases

  • Prompt Engineering: Compare prompt variants systematically instead of manual testing
  • Model Migration: Evaluate a new model against your existing test suite before switching
  • Regression Testing: Catch output quality regressions in CI/CD pipelines
  • Red Teaming: Automatically test for prompt injection, jailbreaks, and harmful outputs
  • RAG Quality: Measure retrieval precision, context relevance, and answer accuracy

Documentation

License

MIT License