inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Methodology

How the capacity planner sizes

Every number the planner emits carries a rule id. Each rule below lists its formula, the inputs it consumes, its defaults, a confidence tag, and the source. This page is generated from the rules registry the engine runs on, so it can't drift from the code. Ruleset 2026.09.6.

← Back to the planner

Concurrency

CONC-01Arrival-rate concurrency (Little's Law)COMPARABLE
L = λ_peak × W; λ_peak = seats × activeShare × reqPerActiveUserHr/3600 × busyHourFactor; W = ttft + out/perUserTps + uncachedPrefill/prefill_rate; peak = Poisson-p95(L). Agentic: L = seats × activeShare × dutyCycle (continuously in-flight).
Inputs
people, workType, latency, busyHourFactor, peakConcurrencyPct
Defaults
busyHourFactor 0.17×24≈4.08; agent duty 0.6; peakConcurrencyPct is a direct override
Source
Little's Law; NVIDIA '2–3 active per 1,000'; iternal 1:10–1:20; Cisco 17% busy-hour; VMware 'requests actively processed by the engine'.
QUEUE-01External arrival-rate queue (Erlang-C)COMPARABLE
offered = arrivals/hr ÷ 3600 × handle_sec; servers = min c where Erlang-C mean wait ≤ target
Inputs
arrivalsPerHour, targetP95WaitSec, latency
Defaults
target wait 20s; handle = avg_output ÷ tok/s + 8s think
Source
Erlang-C; contact-center per-agent 2–6 concurrent chats.
CONC-03Agentic in-flight steps (pinned KV)COMPARABLE
L_request = Poisson-p95(λ_step × W_step); W_step = ttft + step_output/perUserTps + uncached_prefill/prefill_rate; step_output ≈ 250 (per-step, not per-task)
Inputs
people, workType(agentic), latency
Defaults
step_output 250 (PRD §1.2); only in-flight steps pin KV, not all active sessions
Source
TraceLab per-step output (Claude 252/Codex 184); agents keep few steps genuinely mid-decode at once.
MIX-01Mixed-workload envelopeESTIMATED
per-workload mean L=λW by weight; combine (sum by default, or largest if peaksCoincide:false) → Poisson-p95
Inputs
mixedWeights, peaksCoincide
Defaults
sum of per-workload means (they share the cluster) then p95
Source
PRD §2.7 mixed workloads — concurrent workloads add on one cluster.

Token budget

TOK-01Monthly input tokensCOMPARABLE
req/s × avg_input × seconds/month
Inputs
people, workType, prefixCacheHitRate, imagesPerRequest
Defaults
work-type req/hr, avg input; 730h month
Source
Azure LLM inference traces (Splitwise/DynamoLLM medians).
TOK-02Monthly output tokensCOMPARABLE
req/s × avg_output × seconds/month
Inputs
people, workType
Defaults
work-type avg output
Source
Azure trace output medians (chat 129, code-completion 13).
TOK-03Sustained output tok/sCOMPARABLE
active_users × req/s × avg_output
Inputs
people, workType
Defaults
active-user share per work type
Source
Derived from the token profile.
TOK-04Cached (prefix) tokensESTIMATED
monthly_input × prefix_cache_hit_rate
Inputs
prefixCacheHitRate
Defaults
agentic 60%, else 20%
Source
PRD §1.2 — prefix reuse dominant for agents/RAG.

Latency & throughput

LAT-01Interactivity + TTFT presetCOMPARABLE
per-user tok/s + TTFT from the latency tolerance preset
Inputs
latency, targetPerUserTps
Defaults
relaxed 15/2s · standard 25/1s · snappy 50/0.5s
Source
PRD §1.4 latency targets (chat >20 tok/s, TTFT <1s; code 100 tok/s; voice <500ms).
TPUT-01Per-GPU aggregate — measured corpusVERIFIED
look up (model, GPU, quant) in the benchmark corpus
Inputs
model, GPU, quant
Defaults
corpus rows (InferenceMAX/MLPerf/community)
Source
NVIDIA InferenceMAX, MLPerf v5.0, vLLM community.
TPUT-02Per-GPU aggregate — bandwidth-scaledESTIMATED
class_reference × (gpu_bandwidth ÷ reference_bandwidth)
Inputs
GPU memory bandwidth
Defaults
class reference per model size
Source
Decode is bandwidth-bound; scale a class reference by memory bandwidth.
TPUT-03Per-GPU aggregate — formula fallbackESTIMATED
single-stream bandwidth estimate × batch-8
Inputs
GPU bandwidth, active params, quant
Defaults
batch factor 8
Source
Roofline decode estimate when no corpus/class row exists.
TPUT-04Pareto interactivity adjustCOMPARABLE
at target > reference tok/s: aggregate × (ref_per_user ÷ target), floored at single-stream
Inputs
targetPerUserTps
Defaults
reference interactivity per corpus row
Source
PRD §1.4 throughput-vs-interactivity Pareto (up to 100×+ apart).
TPUT-05Required decode + binding constraintESTIMATED
required = max(concurrent × op, sustained); binding = vram | throughput | latency | prefill
Inputs
all sizing inputs
Defaults
—
Source
Whichever of VRAM / throughput / interactivity / prefill dominated the count.
TPUT-06Operating-point search (latency = floor)ESTIMATED
the latency target is a FLOOR: for each GPU search operating points perUserTps ∈ [target, measured-reference], recompute W (hence concurrency) and aggregate at each, pick the min-GPU point (tie → lower op)
Inputs
latency, targetPerUserTps
Defaults
candidates {target, corpus-reference}; never extrapolate above the measured reference
Source
Running faster than the target can cut GPUs (shorter KV occupancy); guarantees gpu_count(relaxed) ≤ standard ≤ snappy.

VRAM & cluster

KV-01KV cache per tokenVERIFIED
kv_bytes/token = 2 × layers × kv_heads × head_dim × bytes_per_elem
Inputs
model anchor, kvCacheDtype
Defaults
FP16 KV (2 bytes); FP8 halves it
Source
Exact transformer KV-cache formula (GQA/MLA reduce kv_heads).
KV-02KV sizing context (fill model)COMPARABLE
ctx/session = fill × p95_ctx + (1 − fill) × avg_ctx (planForP95 → p95 for all)
Inputs
workType, contextTokens, planForP95
Defaults
per-work-type avg/p95/fill (e.g. chat 1.2K/8K/0.15)
Source
PRD §1.2 token profiles; avoids reserving full context per session (analysis #3).
KV-03Prefix-cache budget (agentic)ESTIMATED
prefix_cache_vram = max(0, usable×gpuCount − weights − pinned_KV); effective_cache_hit = min(profile_hit, budget ÷ (sessions × avg_ctx × kv/token)); shortfall raises re-prefill load → prefill tps (can bind)
Inputs
workType(agentic), prefixCacheHitRate
Defaults
idle-session context is a budget, not a reservation; profile hit 0.6 is the cap
Source
vLLM/SGLang prefix caching — cached prefixes live in leftover VRAM; misses re-prefill.
KV-04Cache ↔ prefill fixed pointESTIMATED
iterate gpu_count → cache budget → effective_cache_hit → prefill tps → gpu_count until stable (≤20 iters); choose the min-cost stable point
Inputs
workType(agentic)
Defaults
≤20 iterations; converges in 1–2 for MoE (cheap prefill)
Source
Budget depends on gpu_count and prefill load depends on the budget — a fixed point.
VRAM-01VRAM utilizationESTIMATED
(weights + KV(concurrency × ctx) + overhead) ÷ (usable VRAM × GPUs)
Inputs
model, GPU, concurrency, context
Defaults
usable 90% of advertised; 15% overhead
Source
VRAM budget = weights + KV + activations (PRD §1.4).
CLUS-01GPU type selectionCOMPARABLE
scale-preference list, fewest GPUs then lowest capex; forced GPU overrides
Inputs
people, forceGpuId
Defaults
per scale bucket (consumer → B200)
Source
PRD §1.5 serving stacks by team size.
CLUS-02GPU countESTIMATED
max(VRAM, throughput) → TP groups → + headroom → + N+1/N+2
Inputs
headroomPct, redundancy
Defaults
20% headroom; N+1 large / N+2 xlarge
Source
Headroom + redundancy per PRD §2.4 step 7.
CLUS-03NodesESTIMATED
ceil(GPUs ÷ GPUs_per_node)
Inputs
GPU count
Defaults
8 GPUs/node
Source
Standard 8-GPU node.
CLUS-04Tensor-parallel degreeCOMPARABLE
min GPUs to hold the weights; override forces more
Inputs
tensorParallel
Defaults
auto (weights ÷ 0.8·usable VRAM)
Source
TP sized to fit weights per GPU.
ENG-01Serving engineVERIFIED
scale default; Apple → MLX/llama.cpp; CPU → llama.cpp
Inputs
people, forceGpuId
Defaults
Ollama → vLLM → vLLM/SGLang → TensorRT-LLM
Source
PRD §1.5 (TGI in maintenance since Mar 2026).
TINY-01Tiny-team single deviceCOMPARABLE
≤5 seats → one consumer GPU / Mac Studio; skip cluster math
Inputs
people
Defaults
smallest device that fits
Source
PRD §2.7 Very small — API/single device almost always beats self-host.
FIT-04Existing-hardware fitCOMPARABLE
achievable concurrent = min(VRAM-limited, throughput-limited); additional = sized − owned
Inputs
forceGpuId, existingGpuCount, includeCapex
Defaults
capex sunk unless included
Source
PRD §2.7 Existing hardware — constrain solver, report achievable.

Cost

UTIL-01Projected utilizationESTIMATED
sustained output tok/s ÷ (GPUs × per-GPU aggregate)
Inputs
token budget, cluster
Defaults
—
Source
The utilization that must clear break-even for self-host to win (analysis #2).
COST-01Self-host monthly (line items)ESTIMATED
capex_amort + power + colo + staff (rows); capex = (GPUs + node_chassis[datacenter]) × infra ÷ depr; staff = FTE(nodes) × $/FTE
Inputs
depreciationYears, usdPerKwh, pue, staffMonthlyUsd, targetUtilization
Defaults
4yr depr; PUE 1.4; infra ×2.45 on (GPU+node) hardware; nodeCapex $30K/node; staff 0.1 FTE ≤1 node then ~1 SRE/4 nodes × $15K/mo; power at max(projected, idle 20%)
Source
PRD §1.6 — GPU ≈36% of TCO for a full node; node chassis explicit; FTE ~$180K/yr loaded.
COST-02Rental monthlyVERIFIED
GPU_count × $/GPU-hr × 730
Inputs
forceGpuId
Defaults
live board; else curated fallback
Source
GetDeploying / live rental board (H100 ~$3.32/hr median, Sept 2026).
COST-03API monthlyCOMPARABLE
input×in_price + output×out_price − cache discount (× batch tier)
Inputs
modelRequirement, batchMode
Defaults
live token price; else blended fallback
Source
OpenRouter / models.dev token prices.
COST-04Provisioned throughputCOMPARABLE
PTUs = ceil(weighted peak tokens/min ÷ per-PTU rate); if weighted tok/min < vendor minimum (15 PTU) → cost = null (below_minimum)
Inputs
batchMode
Defaults
Azure PTU: 3K tok/min/PTU, min 15 (~45K tok/min), output weight 4×
Source
Azure PTU calculator; null below minimum; flagged uncertain outside 0.5–20× of API.
COST-PWR-02ColocationCOMPARABLE
power_kW × colocation $/kW/mo (added to power)
Inputs
colocationUsdPerKwMonth
Defaults
none unless set
Source
Datacenter colo pricing.
COST-BEBreak-even utilizationESTIMATED
utilization where self-host $ = API $ at equal delivered tokens
Inputs
cost inputs
Defaults
—
Source
PRD Recommendation 3 — lead with break-even.
COST-3Y-02Growth-scaled 3-yearESTIMATED
3-year = monthly × 12 × (1 + (1+g) + (1+g)²); year-3 GPUs = base × (1+g)³
Inputs
growthPerYear
Defaults
no growth → ×36
Source
Integrates 3 years at compounding scale.

Decision

COST-BE-01API-preference floorCOMPARABLE
if raw cheapest is self-host/rent but API < floor → recommend API + show crossover
Inputs
deployment
Defaults
$20K/mo floor
Source
PRD §1.6 — below ~$15–20K/mo API spend, API wins.
COST-BE-02Crossover volumeESTIMATED
daily tokens where self-host (capped at target util) = API on the fixed cluster
Inputs
targetUtilization
Defaults
target util 55% ceiling
Source
Break-even reframed as a volume the buyer can reason about.
COST-DEC-01Compliance-forced pathCOMPARABLE
air-gap/residency/HIPAA/sovereign → cheaper of self-host/rent; premium vs cheapest
Inputs
compliance, region
Defaults
—
Source
PRD §1.6 compliance drivers; surfaces the flag, not legal advice.
COST-DEC-02User-preferred pathVERIFIED
deployment ≠ 'all' → that path wins; all four still costed
Inputs
deployment
Defaults
—
Source
User choice.
COST-DEC-03Cheapest pathESTIMATED
min of the four monthly costs
Inputs
cost inputs
Defaults
—
Source
Raw cheapest, before the API floor tie-break.

Modality & edge cases

SPCH-01Speech (STT) GPU-secondsVERIFIED
gpu_seconds/day = audio_hours × 3600 ÷ RTF; no token budget
Inputs
audioHoursPerDay
Defaults
Whisper large-v3-turbo RTF 216; API $0.006/audio-min
Source
PRD §1.2 speech (Whisper/Parakeet RTF).
BATCH-01Bursty batchVERIFIED
relax per-user target to ~5 tok/s; API at the batch tier
Inputs
batchMode
Defaults
50% batch discount
Source
AWS Bedrock / OpenAI batch tiers (~50% off).
VIS-01Vision image tokensCOMPARABLE
added input = images × (h×w/1024 + 2) (Qwen3-VL)
Inputs
imagesPerRequest, imageResolution
Defaults
1 image, 1024²
Source
PRD §1.2 modality; cross-model spread ~5× (258 Gemini / 1,334 Claude at 1024²).

Meta

MODEL-01Model tierCOMPARABLE
requirement → 2026 open-weight class + sizing anchor
Inputs
modelRequirement
Defaults
per requirement (e.g. simple → 8B, agentic → gpt-oss-120b)
Source
PRD §1.3 model-requirement mapping.
MODEL-02QuantizationCOMPARABLE
FP8 datacenter / INT4-AWQ small / MXFP4 gpt-oss / GGUF for llama.cpp
Inputs
quant, forceGpuId
Defaults
tier default; downgraded off CUDA-only formats on Apple/CPU
Source
PRD §1.3 quantization defaults (NVFP4 needs Blackwell).

Accuracy band: GPU count is ±30–50% pre-validation — real throughput depends on batch composition, prefix-cache hit rate, and engine tuning a static calculator can't fully capture. See the validation back-tests for measured error per anchor.