Benchmark catalogue
My personal collection of AI related benchmarks and benchmark comparison pages.
Benchmark round-ups and comparisons
Result posts and comparison reports that point to multiple benchmark sources.
- North micro vision benchmark results
Cohere North Micro Vision result report covering document understanding, visual question answering, and other visual-understanding benchmarks.
- LFM2.5-VL visual benchmark results
Vision-language model result post reporting ScreenSpot-v2, RealWorldQA, TextVQA, RefCOCO, and ToolSandbox scores.
- Agentic coding effort and cost comparison
Model-configuration guidance comparing coding performance and cost across reasoning-effort levels.
- TDD impact on coding agents
Empirical evaluation reporting that test-driven development harmed coding-agent performance.
- Atomic Chat HTML5 physics evaluation
Model comparison on self-contained HTML5 physics demos, including token use, cost, and simulation quality.
- Databricks coding-agent evaluation
Harness and model comparison on Databricks' multi-million-line internal codebase.
- Coding benchmark round-up
References ProgramBench, CodeClash, AlgoTune, SWEfficiency, CritPt, and VideoGameBench.
- Fable benchmark results
Results across Vals Index, ProgramBench, ProofBench, VibeCodeBench, and Public Benefits Bench.
- Design Arena, WebDev, and SWE-Bench leaderboard round-up
Comparison of three current model leaderboards.
- APEX-SWE result update
Model results broken down by production-task category.
- Vals model result update
Gemini 3 Pro results across Vals benchmarks.
- SWEfficiency, SciCode, and AlgoTune comparison
Discussion of tougher coding benchmarks beyond SWE-bench.
- Box enterprise extraction evaluation
Internal evaluation of document and image data-extraction accuracy across 40,000 fields.
- Surge internal coding evaluation
Real-world agentic coding comparison across codebases.
Code generation and repository tasks
General code generation, repository-scale work, and agentic coding evaluations.
- ProgramBench
Program-level coding tasks and evaluation.
- MirrorCode
Long-horizon coding benchmark where agents reimplement complete programs from behavior without their source code.
- SlopCodeBench
Coding-quality benchmark with published frontier-model results.
- RepoPrompt Bench
Repository-scale coding benchmark.
- LiveCodeBench
Contamination-resistant evaluation of code generation.
- Interactive coding language evaluation
Result analysis of coding-agent performance across novel interactive environments and programming languages.
- Vals-Smith
Custom coding evaluation that turns merged pull requests into tasks and measures an agent's resolution rate.
- IBench
Coding-agent leaderboard with model result updates.
- GBENCH model result
Coding-agent result comparison covering task performance, cost efficiency, code density, and review quality.
- GertLabs one-shot coding rankings
One-shot coding-agent ranking and model comparison.
- KingBench
Agentic coding benchmark and leaderboard.
- CORE-Bench
Scientific-reproducibility tasks that require agents to run and interpret research repositories.
Frontend, mobile and product coding
Framework-specific, mobile, game, and product-oriented coding work.
- ReactBench
Coding-agent evaluation on realistic React work, including production quality, performance, and accessibility criteria.
- Android Bench
Google's Android-specific coding benchmark and leaderboard covering 100 real-world Android development tasks, with quality, latency, and cost reporting.
- GameDevBench result
Game-development agent benchmark result comparing model task-success rates.
- BridgeBench
Vibe-coding benchmark covering debugging, refactoring, UI, security, cost, and speed.
- ViBench
End-to-end web-app benchmark that assesses whether agents meet product requirements through behavioral test plans.
- Next.js Evals
Next.js's evaluation suite for AI coding agents.
Terminals, harnesses and cost
Terminal agents, harness comparisons, sandbox infrastructure, cost efficiency, and enterprise constraints.
- Terminal-Bench
Terminal-agent benchmark and leaderboard for real command-line tasks.
- Frontier-Bench
Community benchmark for frontier agent work, built by the team behind Terminal-Bench and Harbor.
- OpenBench
Compares coding-agent harnesses by speed, cost per solve, and task performance.
- Databricks Coding Bench
Internal coding-agent evaluation on Databricks' multi-million-line codebase, comparing quality and cost.
- Token-saver benchmark notes
Independent real-task evaluations caution that context-compression tools can increase token cost when they force rereads or break cache reuse.
- Copilot evaluation: neutral-to-higher token use after compression.
- Hermes evaluation: most Headroom changes increased token cost, with one file-search exception.
- Claude Code real-task benchmark: 17% higher token cost from cache disruption.
- RTK evaluation and analysis of command-output token reduction for coding agents.
- High-Performance Sandbox Benchmarks
Compares sandbox providers on repeatable developer and CI/CD workloads, including end-to-end workflow time.
- Boundary-Bench
Enterprise-agent benchmark that measures task performance under realistic EDR, SASE, and DLP policy constraints.
Code review
Code-review agents evaluated on real pull requests and substantive findings.
- ReviewBench
Code-review agent evaluation based on curated, substantive findings from real merged pull requests.
- Martian Code Review Bench
Open-source comparison of code-review agents, measuring bug detection, code improvements, and developer adoption.
- Factory Code Review Benchmark
Evaluation of AI code-review performance.
- Qodo real-PR benchmark
Code-review comparison on 400 real pull requests.
Security and systems coding
Security agents, vulnerability rediscovery, reverse engineering, and low-level systems work.
- ExploitGym
Security-agent benchmark spanning 869 real vulnerabilities across userspace software, V8, and the Linux kernel.
- Aikido cybersecurity CVE benchmark
Fresh-CVE security-agent benchmark testing vulnerability rediscovery, recall, consistency, and cost across 32 recent vulnerabilities and three runs per model.
- Warden security benchmark result
Security-agent benchmark comparison covering accuracy, price-normalized precision, and model cost.
- SRE-Bench
Contamination-resistant software reverse-engineering benchmark with 262 binary instances and 1,572 deterministically graded tasks.
- KernelBench
Agentic GPU-kernel benchmark and leaderboard for writing correct, efficient CUDA kernels.
Model leaderboards and indexes
Cross-model comparisons, benchmark hubs, and aggregate performance indexes.
- Artificial Analysis
Cross-provider model comparisons, including coding and speed.
- BenchmarkList
Registry and capability tracker for AI benchmarks, models, and their release history.
- Roboflow Vision Evals
Ground-truth vision benchmark across object detection, counting, identification, OCR, data extraction, and image reasoning, with cost and speed reporting.
- SpeechMap Lexical Fingerprints
Controlled lexical-fingerprint experiment comparing models' use of signature words across identical questions; the linked ox-alpha page analyzes 1,469 complete responses against 374 comparison models.
- LiveBench
Frequently refreshed benchmark designed to reduce contamination.
- Epoch AI Benchmarks
Collection of AI benchmark analyses and datasets.
- BenchLM
Cross-benchmark model leaderboard combining quality, cost, speed, context, and confidence data.
- Scale Labs Leaderboard
Scale AI's general model-evaluation leaderboard.
- Kaggle Benchmarks
Kaggle's benchmark hub and leaderboards.
- VibeBench
Live, community-voted LLM leaderboard with real-time Vibe Score and Curse Trend metrics.
- VibeBench (Standard Agents)
Practitioner-panel benchmark that compares model releases through engineers' real-world, substantive use.
- Vals Index
Aggregate evaluation of model performance across tasks.
General reasoning and learning
Reasoning, knowledge, abstraction, long-context, and adaptive-learning evaluations.
- EBR-bench
Measures whether agents learn from feedback during repeated play.
- SealedBench
Continuously rotated, sealed evaluations designed to reduce benchmark gaming and training contamination.
- ZeroBench
Difficult multimodal reasoning benchmark covering spatial reasoning, visual understanding, and multi-step reasoning across images.
- PostTrainBench
Post-training-agent benchmark with anti-cheat checks for data contamination, model substitution, API use, and trace lookup.
- EdgeBench
Long-horizon agent learning across scientific, professional, coding, math, and game tasks.
- SimpleBench
General model-capability benchmark with a public high-score tracker.
- AA-Omniscience
Knowledge and hallucination evaluation from Artificial Analysis.
- BullshitBench
Measures whether models detect and clearly challenge flawed or misleading premises across domains and techniques.
- Humanity's Last Exam and ARC-AGI
Reasoning and visual abstraction benchmark comparison.
- ARC-AGI-2 efficiency result
Comparison of top ARC-AGI-2 performance and cost efficiency for GPT-5.6 Sol.
- ARC-AGI-3
Game-like evaluation of how well models orient in unfamiliar situations; includes verified frontier-model results.
- GraphWalks
OpenAI long-context reasoning benchmark for breadth-first graph traversal over very large inputs.
- ClockBench model result
Visual-reasoning benchmark result tracking a new high score and performance by reasoning level.
- WeirdML model result
WeirdML comparison of model performance, result consistency, and cost at a near-saturated frontier score.
Professional workflows and domain agents
Business, language, document, spatial, and customer-workflow evaluations.
- CanvasBench
Framer's 236-task evaluation of AI agents building and maintaining responsive websites, spanning layout, interactions, and design workflows.
- Framer translation-model evaluation
Framer's continuous evaluation of translation models for quality, structural correctness, latency, and cost on real website content.
- AutomationBench
Multi-step business-automation workflows with cost per task.
- ShortcutBench
Internal production-harness evaluation of agent performance on everyday and long-horizon spreadsheet tasks.
- SalesEvals
Sales-call coaching benchmark scored against hidden ground truth across synthetic calls.
- Notion Knowledge Board (live eval)
Ongoing production evaluation of models on anonymized Notion knowledge-work traffic, scored for resolution rate, cost per task, and time; Notion explicitly describes it as a live eval rather than a traditional fixed benchmark.
- DataBench
Hex's frontier benchmark for agentic analytics: 100 realistic Q&A and open-ended data-work tasks in a synthetic warehouse, scored for analytical judgment rather than code generation alone.
- HealthBench Professional
Medical-response leaderboard reporting overall and length-adjusted evaluation scores.
- MedAgentBench
Agentic clinical-EHR benchmark in which models call simulated FHIR APIs to complete multi-step tasks across ten clinical task types.
- DiligenceBench
Equity-research agent evaluation.
- Box Complex Work Eval
End-to-end enterprise document-work evaluation across life sciences, technology, due diligence, and legal analysis.
- Writing and tone benchmark result
Internal evaluation of model writing quality, tone adherence, reasoning effort, and per-task cost.
- τ-voice
Real-time voice-agent evaluation on grounded customer-service tasks.
- Braintrust speech-to-text evaluation
Six-provider, 240-case speech-to-text evaluation covering transcription quality, critical entities, downstream answer quality, and latency.
- LiveKit voice benchmarks
Live voice-agent benchmark reporting real-scenario pass rate, p99 latency, and model comparisons over time.
- Fluid
Benchmark for evaluating AI voice-cleaning systems that remove noise and improve spoken-audio clarity.
- Cura 1T medical benchmark results
Healthcare-model comparison across HealthBench Hard, HealthBench Professional, AgentClinic, and MedAgentBench-v2.
- MedGemma medical benchmarks
Google's evaluation suite for MedGemma across clinical knowledge, radiology, dermatology, pathology, ophthalmology, and multimodal reasoning.
- Harvey Legal Agent Benchmark (LAB)
Open benchmark for long-horizon legal-agent work, evaluating whether agents can produce reliable legal work product across realistic workflows.
- Harvey LAB: Contracts
LAB extension with 500 contract-drafting, review, and negotiation tasks for in-house legal work.
- APEX-Agents: Corporate Law
Long-horizon corporate-law agent benchmark built from expert-authored work in simulated professional environments.
- RedlineBench
Multi-turn contract-redlining benchmark that preserves native tracked changes and comments and grades legal correctness, negotiation quality, and deal-closing judgment.
- Professional Reasoning Bench (PRBench)
Open-ended, expert-authored benchmark for high-stakes professional reasoning in finance and law.
- Context-Bench
Agentic context engineering: file operations, entity tracing, and long-horizon tool calls.
- Finance Agent Benchmark
Leaderboard for finance-agent performance.
- EQ-Bench
Writing and emotional-intelligence evaluation.
- Toloka Arena
Arena-style evaluation for AI agents.
- ParseBench
Document-understanding leaderboard across frontier models, open weights, and OCR solutions.
- Vending-Bench 2
Long-horizon vending-machine agent evaluation, including deceptive-behavior analysis.
- Blueprint-Bench 2
Spatial-reasoning benchmark: agents convert apartment photographs into accurate 2D floor plans.
Web, UI, and interactive agents
Agents that operate user interfaces or build interface-heavy frontend experiences.
- WebUIBench
Agents operating and reasoning over web user interfaces.
- Online-Mind2Web result
Computer-use and browser-use benchmark result for web-navigation agents.
- Composite-Bench
Long-horizon browser computer-use benchmark with certified-optimal answers and verified compute.
- App Control Bench
Mobile-app agent benchmark covering 60 scenarios across two apps, with task completion and speed comparisons.
- Browser-use evaluation result
Browser-agent result comparison covering task performance, token use, and model cost.
- Code Arena: Frontend
Frontend coding leaderboard covering product, content-creation, simulation, gaming, and reference-based design tasks.
Domain reasoning and tool use
Customer-service, finance, mathematics, agent skills, and MCP-based evaluations.
- τ²-Bench-Verified
Verified customer-service agent tasks with tool use.
- Big Finance Benchmark
Finance-domain benchmark from Rogo.
- FrontierFinance
Open, long-horizon investment-workflow benchmark spanning research, events, and coverage monitoring.
- FrontierFinance model result
FrontierFinance's GPT-5.6 Sol comparison across financial-analysis quality and cost per query.
- MathArena
Competitive mathematics evaluation and leaderboard.
- SkillsBench
Benchmarking agent skills and tool-use workflows.
- MCP Atlas
Leaderboard for agent performance with MCP tools.
- Firecrawl Research Index (arXivQA)
Research-agent retrieval evaluation reporting arXivQA recall and comparable-cost results.
- Exa Instant search evaluation
Search-provider comparison using SealQA queries to evaluate retrieval quality and latency for agent workflows.
- Vals Web Search Index
Compares native provider search with independent web-search tools on real-world tasks.
- Exa research-paper benchmarks
Two benchmarks for evaluating scientific-paper search and retrieval across a corpus of more than 300 million papers.
- BrowseComp
Web-research benchmark with 1,266 hard-to-locate factual questions, isolating the effect of search providers, tool configuration, and search budget.
- Bench to the Future 3 (BTF-3)
FutureSearch pastcasting benchmark with 1,907 resolved binary and numeric forecasting questions researched against a frozen web corpus.
- Bench to the Future 2 (BTF-2)
Original FutureSearch pastcasting benchmark with hard forecasting questions, frozen web corpora, and strategic-reasoning evaluation.
- Deep Research Bench (DRB)
Web-research-agent benchmark with 169 diverse real-world tasks, offline search corpora, and curated reference answers.
- Metaculus forecasting evaluations
FutureSearch's live bot evaluations across Metaculus tournaments, including rolling MiniBench contests and bot-versus-human comparisons.
- ForecastBench
Dynamic, contamination-resistant forecasting benchmark scored with a Brier Index on unresolved real-world questions.
- WANDR
Open benchmark for research agents that must search broadly and deeply.