Information (21)
Knowledge sources, documentation hubs, datasets, benchmarks and research.
Machine-readable: /api/information.json.
- Aider LLM Leaderboards
Quantitative benchmarks of LLM code-editing skill maintained by the Aider project.
- Arena (formerly LMArena)
Public, crowd-voted leaderboard comparing AI models on text, image and code through real-world head-to-head evaluation.
- Artificial Analysis
Independent benchmarks comparing AI models and API providers on quality, price, output speed and latency.
- arXiv cs.AI (Artificial Intelligence)
arXiv's listing of the latest preprints in Artificial Intelligence, a primary source for new agent research.
- Berkeley Function Calling Leaderboard (BFCL)
UC Berkeley leaderboard evaluating how accurately LLMs call functions and tools; part of the Gorilla project.
- Building Effective AI Agents (Anthropic)
Anthropic engineering essay (Dec 19, 2024) on agent design patterns: workflows vs. agents, and when to use simple composable patterns.
- Common Crawl
Open repository of web crawl data that anyone can access and analyze.
- Epoch AI Benchmarking Hub
Epoch AI's hub of benchmark results for leading AI models, with trends over time by benchmark and by model.
- GAIA
Benchmark for general AI assistants ("Benchmarking General AI Agents"), published as a Hugging Face dataset under the gaia-benchmark org.
- Hugging Face Daily Papers
Daily curated feed of new AI research papers on Hugging Face.
- Hugging Face Datasets Hub
Hugging Face hub of ready-to-use datasets for AI models, plus the open-source `datasets` library for loading and manipulating them.
- Humanity's Last Exam
Multi-modal benchmark of 2,500 expert-written questions at the frontier of human knowledge, from the Center for AI Safety and Scale AI.
- METR Task-Completion Time Horizons
METR's up-to-date measurements of how long tasks frontier AI models can complete autonomously (time horizons).
- MLE-bench
OpenAI benchmark measuring how well AI agents perform machine learning engineering tasks.
- Models.dev
Open-source database of AI model specifications, pricing and features, also served as JSON.
- OpenRouter LLM Rankings
LLM leaderboard by real-world usage, ranked by tokens processed through the OpenRouter API.
- OSWorld
NeurIPS 2024 benchmark of multimodal agents on open-ended tasks in real computer environments.
- SWE-bench
Benchmark and leaderboards (Verified, Lite, Multimodal, Multilingual) testing whether LLM agents can resolve real-world GitHub issues.
- Terminal-Bench
Benchmark and leaderboard measuring how well AI agents complete complicated tasks in the terminal.
- WebArena
A realistic web environment for building autonomous agents, with benchmark tasks used to evaluate web agents.
- τ-bench (tau2-bench)
Sierra's benchmark for tool-agent-user interaction in real-world domains.