ragas
View on GitHubSupercharge Your LLM Application Evaluations 🚀
Python toolkit for evaluating RAG and other LLM applications with built-in and custom metrics, test-data generation, and integrations for common LLM frameworks and observability tools.
Use Cases
Evaluate RAG application qualityScore LLM responses with custom metricsGenerate evaluation datasetsTrack and improve production LLM applications
Built With
- Language
- Python
- Frameworks
- LangChain · LlamaIndex · Haystack · DSPy
Tags
LLM evaluation · RAG evaluation · metrics · test data generation · LLMOps · production feedback