inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Methodology

Pricing, spread & provenance

Every price we show is provider-published and carries its caveats. This page defines the numbers behind Price Spread, Blended cost, and Find & Fit — and the things we refuse to claim.

Capacity-planner methodology →

Source & provenance

Model capabilities, limits, and prices come from models.dev (MIT), a community-maintained catalog used by OpenCode. Prices are vendor-published list prices, not measured invoices — provenance vendor_claimed. We snapshot the whole catalog daily; every price row carries the day it was observed. Zero-cost aggregate rows are tagged placeholder and excluded from every ranking.

Price spread

For each canonical model we take the latest non-placeholder snapshot from every provider and compute dispersion on published output $/M, base tier only: min, median, max, and the spread ratio = max ÷ min. A provider whose price is above Q3 + 1.5·IQR is flagged an outlier. Only providers matched to the model with ≥ 0.9 confidence enter the stats; weaker (fuzzy) matches are excluded by default so two different models are never averaged together.

Blended cost

Headline input+output pricing is wrong for agentic workloads, which run ~10:1 input:output with heavy cache reads. Blended cost is computed per 1M “workload tokens” for a chosen mix:

cost = uncached_in × input
     + cached_in   × (cache_read ?? input)
     + cache_write × uncached_in × cacheWriteRate
     + output_tokens × output

Presets (all ASSUMPTIONS pending measured agent traces):

  • Agentic coding (10:1) — input share 91%, cache hit 70%.
  • Chat (1:1) — input share 50%, cache hit 20%.
  • Batch summarize (20:1) — input share 95%, cache hit 0%.

When a provider publishes no cache_read, cached reads are charged at full input price and the result is flagged (hatched bar, *) rather than silently guessed. Tiered / context-band prices are computed on the base tier and badged.

Quant is unknown

models.dev records no serving precision. A provider listed cheaper may be serving an fp8/int4 quantization that trades quality for price — so every cross-provider view carries a quant: unknown caveat. Where a provider's own model id leaks a quant (e.g. …-fp8) we surface it as a hint, but we never claim two providers serve identical precision.

Fit

“Fits on” runs the VRAM formula for open-weights models with known parameters (Q4_K_M, 8K context, 90% usable VRAM, searching 1–8 GPU rigs). Closed models and models without an HF weights bridge show “fit unknown” — never “doesn't fit”.

Provider class

Each provider is classed (first-party lab, open-weights host, hyperscaler, relay/router, local runtime) from connection metadata and its catalog mix — not authored. Rules are editorial and versioned: 2026.09.0.