inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Methodology

Frontier — evals, composites & provenance

Frontier plots release date against one benchmark or one versioned percentile composite — never a fabricated cross-benchmark index. Models lacking the selected axis are shown on the unscored rail, not hidden. This page defines every rule and the change log. Open Frontier →

The y-axis is always one basis

An axis is a single benchmark on its native scale (e.g. GPQA Diamond %), or a single versioned category composite. Comparing a 2023 model's MMLU with a 2026 model's GPQA on one axis is a claim the data cannot support and is structurally impossible here. Arena Score (Bradley-Terry) is kept on its own rating scale and is never blended into a percentage composite.

Category taxonomy

Nine categories; the canonical benchmark set per composite category:

  • reasoning: gpqa_diamond, mmlu_pro, hle
  • math: aime_2025, math_500, frontier_math
  • coding: swe_bench_verified, swe_bench_pro, livecodebench, aider_polyglot
  • agentic: terminal_bench, bfcl, tau_bench

Composite method (versioned)

A category composite is the unweighted mean of per-benchmark percentile ranks within a fixed reference cohort (all scored models), displayed only when coverage ≥ 50% of the category's canonical benchmarks and ≥ 2 are present. Below that, the model drops to the unscored rail rather than showing a misleadingly high partial composite. Percentile ranks use the fixed cohort, so a model's composite does not move when you filter. Every composite carries its version (v1) and coverage. Unweighted by design — opaque weighting would manufacture an invented quality number, which we refuse. We reject cohort min–max (unstable under filters), z-score (assumes normality), and raw-scores-on-one-axis (mixes bases).

Tools & scaffold handling

Tools-enabled vs not is a first-class stored flag (HLE-with-tools ≠ closed-book HLE); the two are never merged. Scaffold-sensitive benchmarks (SWE-bench, Aider, Terminal-Bench) always store the harness string — identical weights in different harnesses can swing 10–20 points, so we label and never cross-compare across harnesses.

Benchmark registry & attribution

Scores are provenance-labeled (measured / vendor_claimed / community_reported / estimated, plus demo in mock mode). The primary source is the CC-BY Epoch AI Benchmarking Hub, with CC-BY-4.0 LMArena and Apache-2.0 Aider / Terminal-Bench where they add coverage. Artificial Analysis and Scale SEAL are cited as design precedent only and never ingested; the HF Open LLM archive is deferred pending license clarification.

BenchmarkCategoryVersionLicense
GPQA Diamondreasoning2024CC-BY 4.0
MMLU-Proreasoningv1CC-BY 4.0
Humanity's Last Examreasoning2025CC-BY 4.0
AIME 2025math2025CC-BY 4.0
MATH-500mathv1CC-BY 4.0
FrontierMathmath2024CC-BY 4.0
SWE-bench Verifiedcoding2024-08-13CC-BY 4.0
SWE-bench Procoding2025leaderboard (display-only)
LiveCodeBenchcodingv6open (verified at ingest)
Aider Polyglotcoding2024Apache-2.0
Terminal-Benchagentic1.0Apache-2.0
BFCL v3agenticv3open
τ²-benchagenticv2open
IFEvalinstructionv1open
MMMUmultimodalv1open (verified at ingest)
LMArena Scorepreference2026-09CC-BY 4.0

Change log

  • v1 — initial category composites (reasoning, math, coding, agentic); percentile-rank against the fixed reference cohort; coverage gate 50% / 2 benchmarks.

← Pricing & spread methodology