awesome-evals
View on GitHubA curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
An annotated collection of papers, articles, talks, tools, and benchmarks for building and evaluating AI agents. Includes a practical playbook with runnable examples for graders, trajectory evaluation, error analysis, and CI gating.
Use Cases
Find papers and practical guides on evaluating AI agentsDesign LLM-as-judge evaluationsBuild benchmarks and evaluation datasetsAssess agent tool use and multi-turn behaviorAdd evaluation checks to CIAudit benchmark quality and safety
Tags
AI agent evaluation · LLM evaluation · benchmarks · evals · research resources · evaluation playbook