Benchmark catalogue

My personal collection of AI related benchmarks and benchmark comparison pages.

151 benchmarks across 12 categories

Benchmark round-ups and comparisons

Result posts and comparison reports that point to multiple benchmark sources.

14 benchmarks
Back to categories

Code generation and repository tasks

General code generation, repository-scale work, and agentic coding evaluations.

12 benchmarks
Back to categories

Frontend, mobile and product coding

Framework-specific, mobile, game, and product-oriented coding work.

6 benchmarks
Back to categories

Terminals, harnesses and cost

Terminal agents, harness comparisons, sandbox infrastructure, cost efficiency, and enterprise constraints.

7 benchmarks
Back to categories

Code review

Code-review agents evaluated on real pull requests and substantive findings.

4 benchmarks
Back to categories

Security and systems coding

Security agents, vulnerability rediscovery, reverse engineering, and low-level systems work.

5 benchmarks
Back to categories

Model leaderboards and indexes

Cross-model comparisons, benchmark hubs, and aggregate performance indexes.

12 benchmarks
Back to categories

General reasoning and learning

Reasoning, knowledge, abstraction, long-context, and adaptive-learning evaluations.

17 benchmarks
Back to categories

Professional workflows and domain agents

Business, language, document, spatial, and customer-workflow evaluations.

34 benchmarks
Back to categories

Web, UI, and interactive agents

Agents that operate user interfaces or build interface-heavy frontend experiences.

6 benchmarks
  • WebUIBench

    Agents operating and reasoning over web user interfaces.

  • Online-Mind2Web result

    Computer-use and browser-use benchmark result for web-navigation agents.

  • Composite-Bench

    Long-horizon browser computer-use benchmark with certified-optimal answers and verified compute.

  • App Control Bench

    Mobile-app agent benchmark covering 60 scenarios across two apps, with task completion and speed comparisons.

  • Browser-use evaluation result

    Browser-agent result comparison covering task performance, token use, and model cost.

  • Code Arena: Frontend

    Frontend coding leaderboard covering product, content-creation, simulation, gaming, and reference-based design tasks.

Back to categories

Domain reasoning and tool use

Customer-service, finance, mathematics, agent skills, and MCP-based evaluations.

19 benchmarks
Back to categories