evalscope
View on GitHubA streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
EvalScope is a Python framework for evaluating LLM, vision-language, embedding, reranking, and generative models. It includes benchmark integrations, RAG and agent evaluation, inference load testing, and a dashboard for comparing results.
Use Cases
Evaluate LLM capabilities on benchmark datasetsCompare model outputs and scoresBenchmark inference latency and throughputEvaluate vision-language and multimodal modelsMeasure RAG pipeline qualityRun agent benchmarks and inspect execution tracesGenerate interactive evaluation reports
Built With
- Language
- Python
- Frameworks
- OpenCompass · VLMEvalKit · RAGAS · MTEB
Tags
LLM evaluation · benchmarking · performance testing · VLM evaluation · RAG evaluation · agent evaluation · multimodal · model comparison · evaluation reports · load testing