promptfoo
View on GitHubTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Promptfoo is a CLI and library for evaluating, benchmarking, and red-teaming LLM apps. Declarative YAML configs define prompts, providers, and assertions, then run locally or in CI/CD to compare models like GPT, Claude, and Gemini and catch jailbreaks or regressions.
Use Cases
Evaluate prompt quality with automated assertionsCompare LLM providers side by side on the same test setRed team and pentest LLM apps for jailbreaks and vulnerabilitiesRun LLM regression tests in CI/CD pipelinesScore RAG pipelines for retrieval and answer qualityGenerate security vulnerability reports for GenAI appsScan pull requests for LLM security and compliance issuesRun evals fully locally to keep prompts privateBenchmark model cost, latency, and accuracyTest agent and multi-step workflows
Built With
- Language
- TypeScript
- Frameworks
- Node.js · TypeScript · React · Drizzle ORM · GitHub Actions · Biome
Tags
llm-eval · red-teaming · prompt-testing · llmops · vulnerability-scanning · ci-cd · benchmarking · model-comparison · security · cli · rag-eval · assertions · jailbreak-testing · typescript