inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards

The market read on AI inference:what it costs to borrow.

Token prices across providers, GPU rental across clouds, and the buy-vs-rent math — captured daily, each figure marked measured, estimated, or vendor-claimed.

› cheapest way to run GLM-5.2 today

GLM-5.2 · 101 providers

cheapest — CrofAI$1.05in $0.300·
median$4.40in $1.40·
dearest — Wafer$10.25in $3.00·

output $/M · min → max across providers

Open board

A100 80GB · on-demand · rate-card

runpod$1.19community·datacrunch$1.74secure·coreweave$2.70secure·
median $/GPU-hr$1.74·
Open board
bb1975/mimo-v2.5-proxiaomi-token-plan-cnfree in+out·
deepseek-ai/deepseek-r1iflowcnfree in+out·
deepseek-ai/deepseek-v3iflowcnfree in+out·
deepseek-ai/deepseek-v3.2iflowcnfree in+out·
deepseek-ai/deepseek-v4-proalibaba-token-planfree in+out·

Rate limits and quotas not tracked yet.

Open board

No open-weights model fits at Q4

Nothing in the catalog fits an RTX 4090 (24 GB) at Q4 right now.

Open the fit calculator

25 people · coding · standard latency

API (per-token)cheapest$44/mo±50%COST-03
Rent GPUs$402/mo±50%COST-02
Build (self-host)$1.6k/mo±50%COST-01
Open board
exactarithmetic from known quantities
measureda real benchmarked observation
estimatedcomputed from a formula
vendor-claimedpublished by the vendor, unverified

Speed and reliability: not measured yet — we don't run probes, so we don't show a number.

68,915 observations since 2025-07-01

How we label numbers

Start with your question.

Six ways people come to the inference market. Pick the one you're holding — each links to the board that answers it.

One model, every provider, one view.

The same model is priced very differently across providers. We line them up so the spread is obvious.

  • Cross-provider spread and outliers — ranked on a robust ratio, so one bad row can't headline.
  • Blended cost at your in:out ratio — not just the sticker input price.
  • Cache-read pricing where published — flagged when a provider doesn't.
  • A config snippet to point your client at the cheapest.

GLM-5.2 · 101 providers today

cheapest
$1.05·
CrofAI
median
$4.40·
—
dearest
$10.25·
Wafer

output $/M · non-placeholder providers only

Open the spread

A100 80GB · on-demand · rate-card

runpod$1.19community·datacrunch$1.74secure·coreweave$2.70secure·
median $/GPU-hr$1.74·
Open board

B200 · on-demand · rate-card

runpod$5.98community·datacrunch$6.49secure·coreweave$8.60secure·
median $/GPU-hr$6.49·
Open board

H100 SXM 80GB · on-demand · rate-card

runpod$2.69community·datacrunch$3.35secure·coreweave$6.16secure·
median $/GPU-hr$3.35·
Open board

H200 SXM 141GB · on-demand · rate-card

runpod$3.59community·datacrunch$4.37secure·coreweave$6.31secure·
median $/GPU-hr$4.37·
Open board

L40S · on-demand · rate-card

runpod$0.79community·datacrunch$1.45secure·coreweave$2.25secure·
median $/GPU-hr$1.45·
Open board
1/5 GPUs

Provider coverage today, by population: 178 token/API · 0 GPU-rental (rental capture paused).

What GPUs rent for, by tier, today.

A rental number only means something with its tier and price basis attached. We never mix them.

  • Same tier and price basis only — on-demand rate-card is never blended with an interruptible bid.
  • Marketplace liquidity signals — per-marketplace, with the caveat that they aren't comparable across sources.
  • 7-day preview, 30 days signed in — the full panel is the paid surface.
  • A daily archive since 2025-07-01 — appended, never overwritten.

Borrow, rent, API, PTU, or build.

One deterministic engine costs every path for a workload and ranks them — on cost, never on marketing.

  • Deterministic engine — ruleset 2026.09.6, same inputs, same numbers.
  • Every figure carries a rule id and a ± band — nothing is a bare guess.
  • Break-even utilization is shown, never hidden inside a total.
  • Same engine behind the full Capacity Planner.
API (per-token)cheapestrecommended
$44/mo ±50%
Pay per token to a hosted provider
comparableCOST-03
Rent GPUs
$402/mo ±50%
Marketplace + neocloud, cheapest live rate
comparableCOST-02
Build (self-host)
$1.6k/mo ±50%
Own the GPUs — capex + power + colo + staff
estimatedCOST-01
Provisioned throughput
not offered at this scale
Reserved capacity (Azure PTU-style)
comparableCOST-04

Costs carry a ±50% accuracy band from the engine's back-tests. Archetypes are ranked on cost only — never on reliability or latency, which we don't measure yet.

Cost vs capability, plotted.

Every model by blended $/M against coding score. Toggle the efficiency frontier; pin models to compare.

Prices are provider-published (vendor-claimed). The coding axis is a seeded Aider pass-rate — a size/family heuristic, not a live-ingested benchmark. Free and interactive; share the URL to share the exact view.

Every number is labeled.

exact

Arithmetic from known quantities — VRAM fit is weights + KV cache, not a guess.

measured

A real benchmarked observation. Prefill tok/s is only ever shown when measured.

estimated

Computed from a formula (bandwidth × derate). Decode tok/s is always estimated.

vendor-claimed

Published by the vendor, not independently verified — most posted prices live here.

Speed and reliability: not measured yet — we don't run probes, so we don't show a number. 68,915 observations since 2025-07-01. How we label numbers

History is the product.

7 days — the public preview of any price series.

30 days — signed in, free, for the whole archive window.

Full history & dynamics — the API surface, for anyone pricing against the market.

Sign in for 30 days of history — free.

Get started