evalscope
★ 3.5KEvalScope is a Python framework for evaluating LLM, vision-language, embedding, reranking, and generative models. It includes benchmark integrations, RAG and agent evaluation, inference load testing, and a dashboard for comparing results.
AI Tools | Python · LLM evaluation · benchmarking
View Project →