inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Fit calculator

What runs on my hardware?

Pick your rig — every model that fits, with estimated speed and max context. Fit is exact; decode is estimated from bandwidth; prefill is shown only when measured.

24GB total
8K · capped per model
On rtx-4090 at q4_k_m, 0 of 160 models fit.

Models on rtx-4090 · 0 match · click a row for detail

Nothing fits this rig at these settings. Try a smaller quant, more GPUs, or turn off “Fits only”.
0–0 of 0 models

Fit at q4_k_m, up to 8,192 ctx, 90% usable VRAM across 1 GPU. Decode estimated on llama.cpp. Live data.