Benchmark catalogue

My personal collection of AI related benchmarks and benchmark comparison pages.

143 benchmarks across 12 categories

Benchmark round-ups and comparisons

Result posts and comparison reports that point to multiple benchmark sources.

14 benchmarks
Back to categories

Code generation and repository tasks

General code generation, repository-scale work, and agentic coding evaluations.

12 benchmarks
Back to categories

Frontend, mobile and product coding

Framework-specific, mobile, game, and product-oriented coding work.

6 benchmarks
Back to categories

Terminals, harnesses and cost

Terminal agents, harness comparisons, sandbox infrastructure, cost efficiency, and enterprise constraints.

7 benchmarks
Back to categories

Code review

Code-review agents evaluated on real pull requests and substantive findings.

4 benchmarks
Back to categories

Security and systems coding

Security agents, vulnerability rediscovery, reverse engineering, and low-level systems work.

5 benchmarks
Back to categories

Model leaderboards and indexes

Cross-model comparisons, benchmark hubs, and aggregate performance indexes.

12 benchmarks
Back to categories

General reasoning and learning

Reasoning, knowledge, abstraction, long-context, and adaptive-learning evaluations.

14 benchmarks
  • EBR-bench

    Measures whether agents learn from feedback during repeated play.

  • SealedBench

    Continuously rotated, sealed evaluations designed to reduce benchmark gaming and training contamination.

  • ZeroBench

    Difficult multimodal reasoning benchmark covering spatial reasoning, visual understanding, and multi-step reasoning across images.

  • PostTrainBench

    Post-training-agent benchmark with anti-cheat checks for data contamination, model substitution, API use, and trace lookup.

  • EdgeBench

    Long-horizon agent learning across scientific, professional, coding, math, and game tasks.

  • SimpleBench

    General model-capability benchmark with a public high-score tracker.

  • AA-Omniscience

    Knowledge and hallucination evaluation from Artificial Analysis.

  • BullshitBench

    Measures whether models detect and clearly challenge flawed or misleading premises across domains and techniques.

  • Humanity's Last Exam and ARC-AGI

    Reasoning and visual abstraction benchmark comparison.

  • ARC-AGI-2 efficiency result

    Comparison of top ARC-AGI-2 performance and cost efficiency for GPT-5.6 Sol.

  • ARC-AGI-3

    Game-like evaluation of how well models orient in unfamiliar situations; includes verified frontier-model results.

  • GraphWalks

    OpenAI long-context reasoning benchmark for breadth-first graph traversal over very large inputs.

  • ClockBench model result

    Visual-reasoning benchmark result tracking a new high score and performance by reasoning level.

  • WeirdML model result

    WeirdML comparison of model performance, result consistency, and cost at a near-saturated frontier score.

Back to categories

Professional workflows and domain agents

Business, language, document, spatial, and customer-workflow evaluations.

30 benchmarks
Back to categories

Web, UI, and interactive agents

Agents that operate user interfaces or build interface-heavy frontend experiences.

6 benchmarks
  • WebUIBench

    Agents operating and reasoning over web user interfaces.

  • Online-Mind2Web result

    Computer-use and browser-use benchmark result for web-navigation agents.

  • Composite-Bench

    Long-horizon browser computer-use benchmark with certified-optimal answers and verified compute.

  • App Control Bench

    Mobile-app agent benchmark covering 60 scenarios across two apps, with task completion and speed comparisons.

  • Browser-use evaluation result

    Browser-agent result comparison covering task performance, token use, and model cost.

  • Code Arena: Frontend

    Frontend coding leaderboard covering product, content-creation, simulation, gaming, and reference-based design tasks.

Back to categories

Domain reasoning and tool use

Customer-service, finance, mathematics, agent skills, and MCP-based evaluations.

18 benchmarks
Back to categories