Methodology
How the capacity planner sizes
Every number the planner emits carries a rule id. Each rule below lists its formula, the inputs it consumes, its defaults, a confidence tag, and the source. This page is generated from the rules registry the engine runs on, so it can't drift from the code. Ruleset 2026.09.6.
Concurrency
CONC-01Arrival-rate concurrency (Little's Law)COMPARABLE
L = λ_peak × W; λ_peak = seats × activeShare × reqPerActiveUserHr/3600 × busyHourFactor; W = ttft + out/perUserTps + uncachedPrefill/prefill_rate; peak = Poisson-p95(L). Agentic: L = seats × activeShare × dutyCycle (continuously in-flight).
- Inputs
- people, workType, latency, busyHourFactor, peakConcurrencyPct
- Defaults
- busyHourFactor 0.17×24≈4.08; agent duty 0.6; peakConcurrencyPct is a direct override
- Source
- Little's Law; NVIDIA '2–3 active per 1,000'; iternal 1:10–1:20; Cisco 17% busy-hour; VMware 'requests actively processed by the engine'.
QUEUE-01External arrival-rate queue (Erlang-C)COMPARABLE
offered = arrivals/hr ÷ 3600 × handle_sec; servers = min c where Erlang-C mean wait ≤ target
- Inputs
- arrivalsPerHour, targetP95WaitSec, latency
- Defaults
- target wait 20s; handle = avg_output ÷ tok/s + 8s think
- Source
- Erlang-C; contact-center per-agent 2–6 concurrent chats.
CONC-03Agentic in-flight steps (pinned KV)COMPARABLE
L_request = Poisson-p95(λ_step × W_step); W_step = ttft + step_output/perUserTps + uncached_prefill/prefill_rate; step_output ≈ 250 (per-step, not per-task)
- Inputs
- people, workType(agentic), latency
- Defaults
- step_output 250 (PRD §1.2); only in-flight steps pin KV, not all active sessions
- Source
- TraceLab per-step output (Claude 252/Codex 184); agents keep few steps genuinely mid-decode at once.
MIX-01Mixed-workload envelopeESTIMATED
per-workload mean L=λW by weight; combine (sum by default, or largest if peaksCoincide:false) → Poisson-p95
- Inputs
- mixedWeights, peaksCoincide
- Defaults
- sum of per-workload means (they share the cluster) then p95
- Source
- PRD §2.7 mixed workloads — concurrent workloads add on one cluster.
Token budget
TOK-01Monthly input tokensCOMPARABLE
req/s × avg_input × seconds/month
- Inputs
- people, workType, prefixCacheHitRate, imagesPerRequest
- Defaults
- work-type req/hr, avg input; 730h month
- Source
- Azure LLM inference traces (Splitwise/DynamoLLM medians).
TOK-02Monthly output tokensCOMPARABLE
req/s × avg_output × seconds/month
- Inputs
- people, workType
- Defaults
- work-type avg output
- Source
- Azure trace output medians (chat 129, code-completion 13).
TOK-03Sustained output tok/sCOMPARABLE
active_users × req/s × avg_output
- Inputs
- people, workType
- Defaults
- active-user share per work type
- Source
- Derived from the token profile.
TOK-04Cached (prefix) tokensESTIMATED
monthly_input × prefix_cache_hit_rate
- Inputs
- prefixCacheHitRate
- Defaults
- agentic 60%, else 20%
- Source
- PRD §1.2 — prefix reuse dominant for agents/RAG.
Latency & throughput
LAT-01Interactivity + TTFT presetCOMPARABLE
per-user tok/s + TTFT from the latency tolerance preset
- Inputs
- latency, targetPerUserTps
- Defaults
- relaxed 15/2s · standard 25/1s · snappy 50/0.5s
- Source
- PRD §1.4 latency targets (chat >20 tok/s, TTFT <1s; code 100 tok/s; voice <500ms).
TPUT-01Per-GPU aggregate — measured corpusVERIFIED
look up (model, GPU, quant) in the benchmark corpus
- Inputs
- model, GPU, quant
- Defaults
- corpus rows (InferenceMAX/MLPerf/community)
- Source
- NVIDIA InferenceMAX, MLPerf v5.0, vLLM community.
TPUT-02Per-GPU aggregate — bandwidth-scaledESTIMATED
class_reference × (gpu_bandwidth ÷ reference_bandwidth)
- Inputs
- GPU memory bandwidth
- Defaults
- class reference per model size
- Source
- Decode is bandwidth-bound; scale a class reference by memory bandwidth.
TPUT-03Per-GPU aggregate — formula fallbackESTIMATED
single-stream bandwidth estimate × batch-8
- Inputs
- GPU bandwidth, active params, quant
- Defaults
- batch factor 8
- Source
- Roofline decode estimate when no corpus/class row exists.
TPUT-04Pareto interactivity adjustCOMPARABLE
at target > reference tok/s: aggregate × (ref_per_user ÷ target), floored at single-stream
- Inputs
- targetPerUserTps
- Defaults
- reference interactivity per corpus row
- Source
- PRD §1.4 throughput-vs-interactivity Pareto (up to 100×+ apart).
TPUT-05Required decode + binding constraintESTIMATED
required = max(concurrent × op, sustained); binding = vram | throughput | latency | prefill
- Inputs
- all sizing inputs
- Defaults
- —
- Source
- Whichever of VRAM / throughput / interactivity / prefill dominated the count.
TPUT-06Operating-point search (latency = floor)ESTIMATED
the latency target is a FLOOR: for each GPU search operating points perUserTps ∈ [target, measured-reference], recompute W (hence concurrency) and aggregate at each, pick the min-GPU point (tie → lower op)
- Inputs
- latency, targetPerUserTps
- Defaults
- candidates {target, corpus-reference}; never extrapolate above the measured reference
- Source
- Running faster than the target can cut GPUs (shorter KV occupancy); guarantees gpu_count(relaxed) ≤ standard ≤ snappy.
VRAM & cluster
KV-01KV cache per tokenVERIFIED
kv_bytes/token = 2 × layers × kv_heads × head_dim × bytes_per_elem
- Inputs
- model anchor, kvCacheDtype
- Defaults
- FP16 KV (2 bytes); FP8 halves it
- Source
- Exact transformer KV-cache formula (GQA/MLA reduce kv_heads).
KV-02KV sizing context (fill model)COMPARABLE
ctx/session = fill × p95_ctx + (1 − fill) × avg_ctx (planForP95 → p95 for all)
- Inputs
- workType, contextTokens, planForP95
- Defaults
- per-work-type avg/p95/fill (e.g. chat 1.2K/8K/0.15)
- Source
- PRD §1.2 token profiles; avoids reserving full context per session (analysis #3).
KV-03Prefix-cache budget (agentic)ESTIMATED
prefix_cache_vram = max(0, usable×gpuCount − weights − pinned_KV); effective_cache_hit = min(profile_hit, budget ÷ (sessions × avg_ctx × kv/token)); shortfall raises re-prefill load → prefill tps (can bind)
- Inputs
- workType(agentic), prefixCacheHitRate
- Defaults
- idle-session context is a budget, not a reservation; profile hit 0.6 is the cap
- Source
- vLLM/SGLang prefix caching — cached prefixes live in leftover VRAM; misses re-prefill.
KV-04Cache ↔ prefill fixed pointESTIMATED
iterate gpu_count → cache budget → effective_cache_hit → prefill tps → gpu_count until stable (≤20 iters); choose the min-cost stable point
- Inputs
- workType(agentic)
- Defaults
- ≤20 iterations; converges in 1–2 for MoE (cheap prefill)
- Source
- Budget depends on gpu_count and prefill load depends on the budget — a fixed point.
VRAM-01VRAM utilizationESTIMATED
(weights + KV(concurrency × ctx) + overhead) ÷ (usable VRAM × GPUs)
- Inputs
- model, GPU, concurrency, context
- Defaults
- usable 90% of advertised; 15% overhead
- Source
- VRAM budget = weights + KV + activations (PRD §1.4).
CLUS-01GPU type selectionCOMPARABLE
scale-preference list, fewest GPUs then lowest capex; forced GPU overrides
- Inputs
- people, forceGpuId
- Defaults
- per scale bucket (consumer → B200)
- Source
- PRD §1.5 serving stacks by team size.
CLUS-02GPU countESTIMATED
max(VRAM, throughput) → TP groups → + headroom → + N+1/N+2
- Inputs
- headroomPct, redundancy
- Defaults
- 20% headroom; N+1 large / N+2 xlarge
- Source
- Headroom + redundancy per PRD §2.4 step 7.
CLUS-03NodesESTIMATED
ceil(GPUs ÷ GPUs_per_node)
- Inputs
- GPU count
- Defaults
- 8 GPUs/node
- Source
- Standard 8-GPU node.
CLUS-04Tensor-parallel degreeCOMPARABLE
min GPUs to hold the weights; override forces more
- Inputs
- tensorParallel
- Defaults
- auto (weights ÷ 0.8·usable VRAM)
- Source
- TP sized to fit weights per GPU.
ENG-01Serving engineVERIFIED
scale default; Apple → MLX/llama.cpp; CPU → llama.cpp
- Inputs
- people, forceGpuId
- Defaults
- Ollama → vLLM → vLLM/SGLang → TensorRT-LLM
- Source
- PRD §1.5 (TGI in maintenance since Mar 2026).
TINY-01Tiny-team single deviceCOMPARABLE
≤5 seats → one consumer GPU / Mac Studio; skip cluster math
- Inputs
- people
- Defaults
- smallest device that fits
- Source
- PRD §2.7 Very small — API/single device almost always beats self-host.
FIT-04Existing-hardware fitCOMPARABLE
achievable concurrent = min(VRAM-limited, throughput-limited); additional = sized − owned
- Inputs
- forceGpuId, existingGpuCount, includeCapex
- Defaults
- capex sunk unless included
- Source
- PRD §2.7 Existing hardware — constrain solver, report achievable.
Cost
UTIL-01Projected utilizationESTIMATED
sustained output tok/s ÷ (GPUs × per-GPU aggregate)
- Inputs
- token budget, cluster
- Defaults
- —
- Source
- The utilization that must clear break-even for self-host to win (analysis #2).
COST-01Self-host monthly (line items)ESTIMATED
capex_amort + power + colo + staff (rows); capex = (GPUs + node_chassis[datacenter]) × infra ÷ depr; staff = FTE(nodes) × $/FTE
- Inputs
- depreciationYears, usdPerKwh, pue, staffMonthlyUsd, targetUtilization
- Defaults
- 4yr depr; PUE 1.4; infra ×2.45 on (GPU+node) hardware; nodeCapex $30K/node; staff 0.1 FTE ≤1 node then ~1 SRE/4 nodes × $15K/mo; power at max(projected, idle 20%)
- Source
- PRD §1.6 — GPU ≈36% of TCO for a full node; node chassis explicit; FTE ~$180K/yr loaded.
COST-02Rental monthlyVERIFIED
GPU_count × $/GPU-hr × 730
- Inputs
- forceGpuId
- Defaults
- live board; else curated fallback
- Source
- GetDeploying / live rental board (H100 ~$3.32/hr median, Sept 2026).
COST-03API monthlyCOMPARABLE
input×in_price + output×out_price − cache discount (× batch tier)
- Inputs
- modelRequirement, batchMode
- Defaults
- live token price; else blended fallback
- Source
- OpenRouter / models.dev token prices.
COST-04Provisioned throughputCOMPARABLE
PTUs = ceil(weighted peak tokens/min ÷ per-PTU rate); if weighted tok/min < vendor minimum (15 PTU) → cost = null (below_minimum)
- Inputs
- batchMode
- Defaults
- Azure PTU: 3K tok/min/PTU, min 15 (~45K tok/min), output weight 4×
- Source
- Azure PTU calculator; null below minimum; flagged uncertain outside 0.5–20× of API.
COST-PWR-02ColocationCOMPARABLE
power_kW × colocation $/kW/mo (added to power)
- Inputs
- colocationUsdPerKwMonth
- Defaults
- none unless set
- Source
- Datacenter colo pricing.
COST-BEBreak-even utilizationESTIMATED
utilization where self-host $ = API $ at equal delivered tokens
- Inputs
- cost inputs
- Defaults
- —
- Source
- PRD Recommendation 3 — lead with break-even.
COST-3Y-02Growth-scaled 3-yearESTIMATED
3-year = monthly × 12 × (1 + (1+g) + (1+g)²); year-3 GPUs = base × (1+g)³
- Inputs
- growthPerYear
- Defaults
- no growth → ×36
- Source
- Integrates 3 years at compounding scale.
Decision
COST-BE-01API-preference floorCOMPARABLE
if raw cheapest is self-host/rent but API < floor → recommend API + show crossover
- Inputs
- deployment
- Defaults
- $20K/mo floor
- Source
- PRD §1.6 — below ~$15–20K/mo API spend, API wins.
COST-BE-02Crossover volumeESTIMATED
daily tokens where self-host (capped at target util) = API on the fixed cluster
- Inputs
- targetUtilization
- Defaults
- target util 55% ceiling
- Source
- Break-even reframed as a volume the buyer can reason about.
COST-DEC-01Compliance-forced pathCOMPARABLE
air-gap/residency/HIPAA/sovereign → cheaper of self-host/rent; premium vs cheapest
- Inputs
- compliance, region
- Defaults
- —
- Source
- PRD §1.6 compliance drivers; surfaces the flag, not legal advice.
COST-DEC-02User-preferred pathVERIFIED
deployment ≠ 'all' → that path wins; all four still costed
- Inputs
- deployment
- Defaults
- —
- Source
- User choice.
COST-DEC-03Cheapest pathESTIMATED
min of the four monthly costs
- Inputs
- cost inputs
- Defaults
- —
- Source
- Raw cheapest, before the API floor tie-break.
Modality & edge cases
SPCH-01Speech (STT) GPU-secondsVERIFIED
gpu_seconds/day = audio_hours × 3600 ÷ RTF; no token budget
- Inputs
- audioHoursPerDay
- Defaults
- Whisper large-v3-turbo RTF 216; API $0.006/audio-min
- Source
- PRD §1.2 speech (Whisper/Parakeet RTF).
BATCH-01Bursty batchVERIFIED
relax per-user target to ~5 tok/s; API at the batch tier
- Inputs
- batchMode
- Defaults
- 50% batch discount
- Source
- AWS Bedrock / OpenAI batch tiers (~50% off).
VIS-01Vision image tokensCOMPARABLE
added input = images × (h×w/1024 + 2) (Qwen3-VL)
- Inputs
- imagesPerRequest, imageResolution
- Defaults
- 1 image, 1024²
- Source
- PRD §1.2 modality; cross-model spread ~5× (258 Gemini / 1,334 Claude at 1024²).
Meta
MODEL-01Model tierCOMPARABLE
requirement → 2026 open-weight class + sizing anchor
- Inputs
- modelRequirement
- Defaults
- per requirement (e.g. simple → 8B, agentic → gpt-oss-120b)
- Source
- PRD §1.3 model-requirement mapping.
MODEL-02QuantizationCOMPARABLE
FP8 datacenter / INT4-AWQ small / MXFP4 gpt-oss / GGUF for llama.cpp
- Inputs
- quant, forceGpuId
- Defaults
- tier default; downgraded off CUDA-only formats on Apple/CPU
- Source
- PRD §1.3 quantization defaults (NVFP4 needs Blackwell).
Accuracy band: GPU count is ±30–50% pre-validation — real throughput depends on batch composition, prefix-cache hit rate, and engine tuning a static calculator can't fully capture. See the validation back-tests for measured error per anchor.