inferenstack — the market read on AI inference · last observation 2026-09-20. See the boards
Agentic coding planner

The cheapest model good enough for agentic coding.

Coding capability × cost × tool-calling × whether it runs on your GPU — the join no other tool makes. Agents burn tokens, so $/M and decode speed matter more than for chat.

0.40
0 tok/s

Requires a GPU it fits on.

Capability vs cost · top-left is cheaper & better

13 models with a run cost
150 models match = best value (coding per $)
#ModelCaps
1
deepseek-ai/deepseek-v3.1
0.72
$0.332.2
2
stepfun-ai/step3
0.63
$0.302.1
3
deepseek-ai/deepseek-v3.1-terminus
0.72
$0.362.0
4
arcee-ai/trinity-large-thinking
0.65
$0.391.7
5
nex-agi/nex-n2-pro
0.65
$0.441.5
6
minimaxai/minimax-text-01
0.45
$0.431.1
7
inclusionai/ring-2.6-1t
0.74
API
$0.850.9
8
deepseek-ai/deepseek-v3.2-speciale
0.74
$0.850.9
9
deepseek-ai/deepseek-prover-v2-671b
0.70
$0.920.8
10
inclusionai/ring-1t
0.74
API
$0.980.8
11
moonshotai/kimi-k2-thinking
0.60
$0.860.7
12
meta-llama/llama-3.1-405b-fp8
0.55
$1.950.3
13
coherelabs/command-a-plus-05-2026-bf16
0.59
$4.380.1
14
willfalco/glm-5.2-exl3-tr3-3.36bpw
0.73
——
15
0xsero/deepseek-v4-flash-0731-reap
0.63
——
1–15 of 150 models
1 / 10

Coding scores, tool-calling flags, and prices are live data. Local $/M and speed = estimated decode at Q4_K_M / 8K on the selected GPU. Value = coding score per cheapest $/M (API or local, whichever is lower — the cheaper cell is highlighted).