Terminal-Bench
Benchmark and leaderboard measuring how well AI agents complete complicated tasks in the terminal.
Tags
Related
- SWE-bench: Benchmark and leaderboards (Verified, Lite, Multimodal, Multilingual) testing whether LLM agents can resolve real-world GitHub issues.
Sources
Machine-readable
- JSON
- Markdown
- Source file on GitHub (edit via pull request)