SWE-agent
View on GitHubSWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
SWE-agent lets a chosen LM autonomously use tools to fix GitHub issues in real repos, hitting state-of-the-art open-source SWE-bench results. It also supports offensive-security (EnIGMA) and competitive coding tasks, configured via YAML. Note: development has moved to mini-SWE-agent.
Use Cases
Automatically fix GitHub issues in real repositoriesBenchmark LLMs on SWE-bench and SWE-bench VerifiedSolve offensive cybersecurity / CTF challenges via EnIGMA modeAutonomous multi-step code editing and test executionResearch platform for agent-computer interface designRun custom LM agent tasks from a single YAML configBatch-mode evaluation of coding agentsReproduce agent trajectories and replay runs
Built With
- Language
- Python
- Frameworks
- litellm · pydantic · swe-rex · flask · textual · pytest · GitPython · docker
Tags
swe-bench · software-engineering · autonomous-agent · github-issues · bug-fixing · computer-use · cybersecurity · ctf · agent-computer-interface · llm · python · research · code-repair · benchmarking · tool-use · neurips