Frontier — evals, composites & provenance
Frontier plots release date against one benchmark or one versioned percentile composite — never a fabricated cross-benchmark index. Models lacking the selected axis are shown on the unscored rail, not hidden. This page defines every rule and the change log. Open Frontier →
The y-axis is always one basis
An axis is a single benchmark on its native scale (e.g. GPQA Diamond %), or a single versioned category composite. Comparing a 2023 model's MMLU with a 2026 model's GPQA on one axis is a claim the data cannot support and is structurally impossible here. Arena Score (Bradley-Terry) is kept on its own rating scale and is never blended into a percentage composite.
Category taxonomy
Nine categories; the canonical benchmark set per composite category:
- reasoning:
gpqa_diamond, mmlu_pro, hle - math:
aime_2025, math_500, frontier_math - coding:
swe_bench_verified, swe_bench_pro, livecodebench, aider_polyglot - agentic:
terminal_bench, bfcl, tau_bench
Composite method (versioned)
A category composite is the unweighted mean of per-benchmark percentile ranks within a fixed reference cohort (all scored models), displayed only when coverage ≥ 50% of the category's canonical benchmarks and ≥ 2 are present. Below that, the model drops to the unscored rail rather than showing a misleadingly high partial composite. Percentile ranks use the fixed cohort, so a model's composite does not move when you filter. Every composite carries its version (v1) and coverage. Unweighted by design — opaque weighting would manufacture an invented quality number, which we refuse. We reject cohort min–max (unstable under filters), z-score (assumes normality), and raw-scores-on-one-axis (mixes bases).
Tools & scaffold handling
Tools-enabled vs not is a first-class stored flag (HLE-with-tools ≠ closed-book HLE); the two are never merged. Scaffold-sensitive benchmarks (SWE-bench, Aider, Terminal-Bench) always store the harness string — identical weights in different harnesses can swing 10–20 points, so we label and never cross-compare across harnesses.
Benchmark registry & attribution
Scores are provenance-labeled (measured / vendor_claimed / community_reported / estimated, plus demo in mock mode). The primary source is the CC-BY Epoch AI Benchmarking Hub, with CC-BY-4.0 LMArena and Apache-2.0 Aider / Terminal-Bench where they add coverage. Artificial Analysis and Scale SEAL are cited as design precedent only and never ingested; the HF Open LLM archive is deferred pending license clarification.
| Benchmark | Category | Version | License |
|---|---|---|---|
| GPQA Diamond | reasoning | 2024 | CC-BY 4.0 |
| MMLU-Pro | reasoning | v1 | CC-BY 4.0 |
| Humanity's Last Exam | reasoning | 2025 | CC-BY 4.0 |
| AIME 2025 | math | 2025 | CC-BY 4.0 |
| MATH-500 | math | v1 | CC-BY 4.0 |
| FrontierMath | math | 2024 | CC-BY 4.0 |
| SWE-bench Verified | coding | 2024-08-13 | CC-BY 4.0 |
| SWE-bench Pro | coding | 2025 | leaderboard (display-only) |
| LiveCodeBench | coding | v6 | open (verified at ingest) |
| Aider Polyglot | coding | 2024 | Apache-2.0 |
| Terminal-Bench | agentic | 1.0 | Apache-2.0 |
| BFCL v3 | agentic | v3 | open |
| τ²-bench | agentic | v2 | open |
| IFEval | instruction | v1 | open |
| MMMU | multimodal | v1 | open (verified at ingest) |
| LMArena Score | preference | 2026-09 | CC-BY 4.0 |
Change log
v1— initial category composites (reasoning, math, coding, agentic); percentile-rank against the fixed reference cohort; coverage gate 50% / 2 benchmarks.