safelabs-eval
View on GitHubOWASP ASI-aligned red-teaming and evaluation framework for AI agents
safelabs-eval is a framework-agnostic Python red-teaming toolkit that fires 131 OWASP ASI-aligned adversarial prompts at any agent HTTP endpoint or Python callable, scoring responses with pure-Python detectors (no LLM cost) and emitting structured JSON reports for CI.
Use Cases
Red-team a live AI agent HTTP endpointWrap any Python callable as an eval targetRun the 131-prompt OWASP ASI Top 10 suite in CIDetect prompt injection complianceDetect data leakage and confidentiality breachesDetect jailbreak and behavioral driftDetect excessive agency and scope violationsDetect hallucination and misinformation outputsTest agents built with LangChain, CrewAI, AutoGen or LlamaIndexTest OpenAI Agents SDK, Google ADK or Semantic Kernel agentsProduce JSON security reports for compliance auditsBrowse and filter the adversarial prompt library by severity
Built With
- Language
- Python
- Frameworks
- LangChain · CrewAI · AutoGen/ag2 · LlamaIndex · OpenAI Agents SDK · Google ADK · Semantic Kernel · Pydantic · Click · httpx · pytest · hatchling
Tags
red-teaming · agent-security · owasp-asi · prompt-injection · jailbreak-detection · llm-security · evaluation · ai-safety · adversarial-prompts · security-testing · benchmark · cli · ci-cd · detectors · python · guardrails