Token prices across providers, GPU rental across clouds, and the buy-vs-rent math — captured daily, each figure marked measured, estimated, or vendor-claimed.
GLM-5.2 · 101 providers
output $/M · min → max across providers
Open boardA100 80GB · on-demand · rate-card
Rate limits and quotas not tracked yet.
Open boardNo open-weights model fits at Q4
Nothing in the catalog fits an RTX 4090 (24 GB) at Q4 right now.
Open the fit calculator25 people · coding · standard latency
Speed and reliability: not measured yet — we don't run probes, so we don't show a number.
68,915 observations since 2025-07-01
How we label numbersSix ways people come to the inference market. Pick the one you're holding — each links to the board that answers it.
The same model is priced very differently across providers. We line them up so the spread is obvious.
GLM-5.2 · 101 providers today
output $/M · non-placeholder providers only
Open the spreadA100 80GB · on-demand · rate-card
B200 · on-demand · rate-card
H100 SXM 80GB · on-demand · rate-card
H200 SXM 141GB · on-demand · rate-card
L40S · on-demand · rate-card
Provider coverage today, by population: 178 token/API · 0 GPU-rental (rental capture paused).
A rental number only means something with its tier and price basis attached. We never mix them.
One deterministic engine costs every path for a workload and ranks them — on cost, never on marketing.
Costs carry a ±50% accuracy band from the engine's back-tests. Archetypes are ranked on cost only — never on reliability or latency, which we don't measure yet.
Every model by blended $/M against coding score. Toggle the efficiency frontier; pin models to compare.
Prices are provider-published (vendor-claimed). The coding axis is a seeded Aider pass-rate — a size/family heuristic, not a live-ingested benchmark. Free and interactive; share the URL to share the exact view.
Arithmetic from known quantities — VRAM fit is weights + KV cache, not a guess.
A real benchmarked observation. Prefill tok/s is only ever shown when measured.
Computed from a formula (bandwidth × derate). Decode tok/s is always estimated.
Published by the vendor, not independently verified — most posted prices live here.
Speed and reliability: not measured yet — we don't run probes, so we don't show a number. 68,915 observations since 2025-07-01. How we label numbers
7 days — the public preview of any price series.
30 days — signed in, free, for the whole archive window.
Full history & dynamics — the API surface, for anyone pricing against the market.