judgeval
View on GitHubThe Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python SDK for tracing and evaluating LLM applications and agents. Instrument agent runs, score live or historical traces with prompt-based judges, and monitor production behavior for failures and regressions.
Use Cases
Trace LLM application and agent behaviorScore production traffic with prompt-based judgesFind and triage agent failuresReplay historical traces to validate fixesMonitor for regressions and alert on detected behaviors
Built With
- Language
- Python
- Frameworks
- OpenTelemetry · LangGraph · OpenLit · Claude Agent SDK
Tags
LLM observability · agent evaluation · tracing · online monitoring · evals · agent judges · failure analysis · OpenTelemetry