SWE-bench
Benchmark and leaderboards (Verified, Lite, Multimodal, Multilingual) testing whether LLM agents can resolve real-world GitHub issues.
Tags
Related
- Terminal-Bench: Benchmark and leaderboard measuring how well AI agents complete complicated tasks in the terminal.
Sources
Machine-readable
- JSON
- Markdown
- Source file on GitHub (edit via pull request)