agent-learning-kit
View on GitHubGeneral Purpose Evaluation and Simulation Environment for all your AI related Workflows
A Python SDK and CLI for evaluating, simulating, red-teaming, and optimizing AI agents. Run local tests and framework adapters, inspect replayable artifacts, and add agent quality gates to CI.
Use Cases
Evaluate agent behavior and runtime tracesSimulate agent tasks and environmentsRed-team agents and workflowsOptimize prompts and agent configurationsTest multi-agent coordination and handoffsCheck retrieval and memory qualityRun regression gates in CIProduce replayable release evidence
Built With
- Language
- Python
- Frameworks
- LangChain · LangGraph · LlamaIndex · AutoGen · CrewAI · LiveKit · Pipecat · PydanticAI · OpenAI Agents · Google ADK
Tags
agent-evaluation · agent-simulation · red-teaming · optimization · CI/CD · regression-testing · multi-agent · local-first