inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Find & Fit

Filter by capability, price, and what fits.

160 of 160 models · cross-provider price from models.dev · smallest-GPU-that-fits computed live.

CacheSmallest rigRent $/hrWeights
deepseek-ai/deepseek-v4-pro1.6T
0.80
$0.696routing-run766.9×89%3× mi455x (1296GB)—open
zai-org/glm-5.3753B
0.75
$0——–2× mi355x (576GB)—open
deepseek-ai/deepseek-v4.1-flashnew763B
0.75
$0.280amd498.6×96%2× mi355x (576GB)—open
deepseek-ai/deepseek-v3.2-speciale685B
0.74
$1.68——–2× mi325x (512GB)—open
deepseek-ai/deepseek-v3.2685B
0.74
$0.295orcarouter496.3×63%2× mi325x (512GB)—open
deepseek-ai/deepseek-v3.2-exp685B
0.74
$0.320——–2× mi325x (512GB)—open
inclusionai/ring-2.6-1t1.0T
0.74
$2.50requesty31.0×67%2× mi455x (864GB)—closed
bb1975/mimo-v2.5-pro1.0T
0.74
$0——–2× mi455x (864GB)—open
inclusionai/ring-2.5-1t1.0T
0.74
———–2× mi455x (864GB)—open
inclusionai/ring-1t1000B
0.74
$2.24——–2× mi455x (864GB)—closed
avlp12/inkling-975b-alis-mlx-dynamic-3.7bpw947B
0.73
———–2× mi455x (864GB)—open
thinkingmachines/inkling952B
0.73
$4.05edenai242.3×88%2× mi455x (864GB)—open
redhatai/inkling-fp8-block952B
0.73
———–2× mi455x (864GB)—open
willfalco/glm-5.2-exl3-tr3-3.36bpw175B
0.73
———–b100 (180GB)—open
blackfrost-ai/glm-5.2-abliterated-reap-nu176-nvfp4273B
0.73
———–mi325x (256GB)—open
1–15 of 160 models
1 / 11

Fit status applies to open-weights models with known params (Q4_K_M, 8K context, 90% usable VRAM, 1–8 GPUs); closed models show “fit unknown”, never “doesn't fit”. Cross-provider price, providers, spread, and cache from models.dev. Live data.