rrsi
View on GitHubA Python system that iteratively proposes, screens, evaluates, and selects changes to AI agent harnesses, including prompts, tools, memory, and control flow. It supports coding, document-work, and engineering-design benchmarks, with auditable Git-based candidate histories.
Use Cases
Evolve and evaluate coding agent harnessesOptimize document-work agent harnessesImprove engineering-design agentsCompare harness changes on in-distribution and held-out benchmarksTrack agent edits, measured outcomes, and evaluation costs
Built With
- Language
- Python
- Frameworks
- Anthropic Vertex AI · LiteLLM · Harbor · MCP · FastMCP
Tags
agent harnesses · self-improvement · agent optimization · benchmarking · evaluation · harness evolution · coding agents · multi-domain agents