awesome-evals
★ 923An annotated collection of papers, articles, talks, tools, and benchmarks for building and evaluating AI agents. Includes a practical playbook with runnable examples for graders, trajectory evaluation, error analysis, and CI gating.
AI Tools | AI agent evaluation · LLM evaluation
View Project →