inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
About

Provenance & labels

Every derived figure on this site carries a provenance label so you always know how much to trust it. Labeling is the product's differentiator, not a disclaimer — the developer audience rewards “tells you when it doesn't know” over confident fake precision. Here is exactly what each label means.

exact

Arithmetic from known quantities. No formula fudge factor, no estimate.

e.g. VRAM fit = model weights + KV cache (2·layers·kv_heads·head_dim·bytes) + overhead. It either fits or it doesn't.

measured

A real benchmarked observation from a corpus, not a formula.

e.g. Prefill tok/s from the llmfit community corpus. Prefill is only ever shown when measured — never estimated.

estimated

Computed from a formula. Treat it as a ballpark, not a benchmark.

e.g. Decode tok/s = memory bandwidth × a per-backend derate (llama.cpp 0.65, vLLM 0.70, MLX 0.75). Always badged estimated.

vendor-claimed

Published by the vendor and not independently verified. Most posted prices live here.

e.g. Token $/M from models.dev / OpenRouter and GPU rental $/hr from provider rate cards — list prices, not measured invoices.

community-reported

Reported by the community rather than a first-party or by us.

e.g. Measured tok/s contributed to the llmfit corpus, and marketplace liquidity signals from Vast listings.

What we refuse to claim (yet)

Full methodology →Roadmap →