inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Explore

Compare on a level field.

Percentile-normalized so raw tok/s and dollars across hardware classes become comparable — 80th percentile means better than 80% of the set.

Model profiles

deepseek-v4-prodeepseek-v4.1-flashglm-5.3

GPU price vs performance

Perf = estimated decode tok/s on Llama-3.1-8B Q4 (bandwidth-driven).