Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

THE MODEL SELECTION SERIES

Best models for coding

Writing a function and fixing a repository are different jobs. Our Coding assistant profile gives 45% weight to LiveBench agentic coding, 30% to coding tasks, 15% to reasoning and 10% to instruction following.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

01 / WHAT TO CHECK

Look at the two coding columns separately

Coding measures contained code generation and completion. Agentic coding measures performance in a tool-using workflow across JavaScript, TypeScript and Python. A model that excels at the first may need a better harness or more attempts for the second.

02 / WHAT TO CHECK

Keep the development harness constant

Repository search, patch application, terminal access and the test loop change results. Use the same harness for every candidate, and record the model version and reasoning setting. Our tables preserve the source’s tested configuration.

03 / WHAT TO CHECK

Budget for the whole edit cycle

Add the codebase context sent on each turn and all billed output across planning, edits and tests. Cached context can reduce costs, but cache policies differ. The calculator intentionally shows an uncached baseline so its assumptions are visible.

04 / WHAT TO CHECK

Check what the benchmark does not cover

Run your own tests for framework knowledge, unfamiliar dependencies, migrations, security-sensitive code and multi-file changes. Benchmarks support a shortlist; your regression suite decides whether a change is acceptable.

YOUR WORKLOAD. YOUR PRIORITIES.

Find your model sweet spot.

How our ratings work ↗

Make the numbers yours Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.

Standard text API estimate. Include reasoning tokens in output.

A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.

THE MARKET MAP

Capability meets cost.

Efficient frontierOther modelsHigher + further left = more for less
020406080100$1$10$100$1,000Agent Fit / 100 ↑Estimated workload cost · USD · logarithmic scaleClaude Opus 5.5: 79.4 · $80DeepSeek V4.1 Flash: 78.8 · $2.14Claude Fable 5.1: 76.7 · $200Claude Fable 5: 74.8 · $200Muse Spark 1.3: 74.4 · $21Claude Opus 5: 73.8 · $100Kimi K3: 73.1 · $60GPT-5.6 Sol: 71.4 · $80GPT-6 Astra: 71.4 · $200Gemini 3.7 Flash: 71.1 · $15GLM-5.3: 70.9 · $13.68Claude Sonnet 5: 70.6 · $40Muse Spark 1.2: 70.1 · $21DeepSeek V4 Flash Vision Exp: 69.7 · $3.52Muse Spark 1.1: 69.6 · $21Qwen3.8 27B: 69.6 · $10.2GPT-5.5: 69.5 · $110Grok 4.6: 69.5 · $32GPT-6 Sol: 68.5 · $40GPT-5.6 Terra: 68.3 · $44Claude Opus 4.8: 67.9 · $100GPT-5.4: 67.7 · $55Gemini 3.8 Flash: 67.7 · $15DeepSeek V4 Pro 0813: 67.5 · $10.56Grok 4.7: 67.4 · $25.6Claude Opus 4.7: 67.2 · $100Grok 4.5: 66.2 · $32GLM-5.3 Flash: 66.2 · $2.5GPT-5.2 Codex: 65.6 · $45.5GPT-5.6 Luna: 65.5 · $4.4Gemini 3.5 Flash: 65.4 · $33GLM-5.2: 65.2 · $10.58Claude Opus 4.6: 65.1 · $100GPT-6 Luna: 64.6 · $2GPT-5.2: 64.1 · $45.5Gemini 3.1 Pro Preview: 63.3 · $44Gemini 3.6 Flash: 63.2 · $15DeepSeek V4 Flash 0731: 63.1 · $1.68Kimi K2.6: 63.0 · $17.5Inkling: 62.3 · $18.1Claude Sonnet 4.6: 62.0 · $60Qwen3.7 Max: 61.8 · $23.6GPT-5.4 nano: 61.2 · $4.5Kimi K2.7 Code: 60.8 · $13.66Claude Opus 4.5: 60.0 · $100Qwen3.6 Plus: 59.3 · $7.15Gemini 3.5 Flash-Lite: 58.9 · $8DeepSeek V4 Pro: 58.8 · $13.37Nemotron 3 Ultra: 57.2 · $10.8GPT-5.4 mini: 56.9 · $16.5MiniMax M3: 55.7 · $5.4Qwen3.6 27B: 55.1 · $8.6DeepSeek V4 Flash: 54.6 · $1.24Grok 4.3: 46.2 · $17.5
Claude Opus 5.579.4 Agent Fit$80
AI AGENT STORE ORIGINAL

More capability.
Less budget.

Practical Value Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. · Coding assistant

70% Agent Fit + 30% affordability. Rankings use your workload; filters do not change the reference cohort.

THE EVIDENCE, SIDE BY SIDE

AI model comparison table

56 of 56 models · select up to 4 to comparePrices in USD · scores explained with ?
AI model benchmarks, aggregate ratings, API pricing and context windows. Snapshot 2026-09-23.
CompareOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set.People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference.Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.Input / outputUSD / 1M tokens
01
Claude Opus 5.5Anthropic · max
79.461.083.2$801.00M$4 / $20Direct API
02
DeepSeek V4.1 FlashDeepSeek · maxOpen weights
78.883.581.1$2.141.05M$0.094 / $0.6OpenRouter catalog
03
Claude Fable 5.1Anthropic · max
76.754.387.52 sources83.41,498±8$2001.00M$10 / $50Direct API
04
Claude Fable 5Anthropic · source ID max; display xhighConfiguration discrepancy
74.852.983.0$2001.00M$10 / $50Direct API
05
Muse Spark 1.3Meta · xhigh
74.466.881.6$211.05M$1.25 / $4.25OpenRouter catalog
06
Claude Opus 5Anthropic · max
73.855.156.32 sources80.11,487±5$1001.00M$5 / $25Direct API
07
Kimi K3Moonshot AI · source defaultOpen weights
73.157.779.2$601.05M$3 / $15OpenRouter catalog
08
Qwen3.8 MaxAlibaba · source default · version mapping unconfirmed
71.678.5Not verifiedNot verified
09
GPT-5.6 SolOpenAI · source ID max; display xhighConfiguration discrepancy
71.455.481.1$801.05M$4 / $20Direct API
10
GPT-6 AstraOpenAI · max
71.450.553.12 sources82.21,480±12$2001.05M$10 / $50Direct API
11
Gemini 3.7 FlashGoogle · high
71.168.456.32 sources78.81,490±8 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
12
GLM-5.3Z.ai · source defaultOpen weights
70.969.476.1$13.681.31M$0.84 / $2.64OpenRouter catalog
13
Claude Sonnet 5Anthropic · xhigh
70.659.976.0$401.00M$2 / $10Direct API
14
Qwen3.8 Flash NextAlibaba · source defaultOpen weights
70.376.2Not verifiedNot verified
15
Muse Spark 1.2Meta · xhigh
70.163.868.82 sources78.01,500±11$211.05M$1.25 / $4.25OpenRouter catalog
16
DeepSeek V4 Flash Vision ExpDeepSeek · experimental · source defaultOpen weightsPreview
69.775.976.8$3.521.05M$0.22 / $0.66OpenRouter catalog
17
Muse Spark 1.1Meta · source ID xhigh; display highConfiguration discrepancy
69.663.475.3$211.05M$1.25 / $4.25OpenRouter catalog
18
Qwen3.8 27BAlibaba · source defaultOpen weights
69.671.975.3$10.21.00M$0.42 / $3OpenRouter catalog
19
GPT-5.5OpenAI · xhigh
69.550.380.2$1101.05M$5 / $30Direct API
20
Grok 4.6xAI · source default
69.560.878.0$32500K$2 / $6OpenRouter catalog
21
GPT-6 SolOpenAI · max
68.558.479.2$401.05M$2 / $10Direct API
22
GPT-5.6 TerraOpenAI · source ID max; display xhighConfiguration discrepancy
68.357.177.9$441.05M$2 / $12Direct API
23
Claude Opus 4.8Anthropic · max
67.950.976.2$1001.00M$5 / $25Direct API
24
GPT-5.4OpenAI · xhigh
67.754.878.0$551.05M$2.5 / $15Direct API
25
Gemini 3.8 FlashGoogle · high
67.766.150.02 sources75.81,493±9 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
26
DeepSeek V4 Pro 0813DeepSeek · source defaultOpen weights
67.569.977.4$10.561.05M$0.66 / $1.98OpenRouter catalog
27
Grok 4.7xAI · xhigh
67.460.277.4$25.6500K$1.6 / $4.8OpenRouter catalog
28
Claude Opus 4.7Anthropic · xhigh
67.250.476.5$1001.00M$5 / $25Direct API
29
Grok 4.5xAI · source default
66.258.575.8$32500K$2 / $6OpenRouter catalog
30
GLM-5.3 FlashZ.ai · source defaultOpen weights
66.274.071.6$2.51.31M$0.15 / $0.5OpenRouter catalog
31
GPT-5.2 CodexOpenAI · source default
65.654.174.0$45.5400K$1.75 / $14OpenRouter catalog
32
GPT-5.6 LunaOpenAI · source ID max; display xhighConfiguration discrepancy
65.572.573.6$4.41.05M$0.2 / $1.2Direct API
33
Gemini 3.5 FlashGoogle · high
65.457.112.52 sources74.61,478±4$331.05M$1.5 / $9OpenRouter catalog
34
GLM-5.2Z.ai · source defaultOpen weights
65.267.773.2$10.581.05M$0.6496 / $2.04OpenRouter catalog
35
Claude Opus 4.6Anthropic · high · adaptive thinking
65.149.056.32 sources74.51,505±4$1001.00M$5 / $25Direct API
36
GPT-6 LunaOpenAI · max
64.674.172.0$21.05M$0.1 / $0.5Direct API
37
GPT-5.2OpenAI · high · 2025-12-11
64.153.174.6$45.5400K$1.75 / $14OpenRouter catalog
38
Gemini 3.1 Pro PreviewGoogle · highPreview
63.353.777.0$441.05M$2 / $12OpenRouter catalog
39
Gemini 3.6 FlashGoogle · high
63.262.99.42 sources73.61,480±5$151.05M$0.75 / $3.75OpenRouter catalog
40
DeepSeek V4 Flash 0731DeepSeek · source defaultOpen weights
63.173.674.2$1.681.31M$0.04 / $0.64OpenRouter catalog
41
Kimi K2.6Moonshot AI · thinkingOpen weights
63.060.870.5$17.5262K$0.95 / $4OpenRouter catalog
42
InklingThinking Machines · xhighOpen weights
62.359.571.9$18.11.05M$1 / $4.05OpenRouter catalog
43
Claude Sonnet 4.6Anthropic · medium · adaptive thinking
62.049.973.0$601.00M$3 / $15Direct API
44
Qwen3.7 MaxAlibaba · source default
61.856.873.1$23.61.00M$1.48 / $4.43OpenRouter catalog
45
GPT-5.4 nanoOpenAI · xhigh
61.268.969.6$4.5400K$0.2 / $1.25Direct API
46
Kimi K2.7 CodeMoonshot AI · source defaultOpen weights
60.862.968.4$13.66262K$0.7062 / $3.3OpenRouter catalog
47
Claude Opus 4.5Anthropic · high · 64K thinking
60.045.472.6$100200K$5 / $25Direct API
48
Qwen3.6 PlusAlibaba · source default
59.366.468.9$7.151.00M$0.325 / $1.95OpenRouter catalog
58.965.663.9$81.05M$0.3 / $2.5OpenRouter catalog
50
DeepSeek V4 ProDeepSeek · source defaultOpen weights
58.862.171.6$13.371.05M$0.9553 / $1.91OpenRouter catalog
51
Nemotron 3 UltraNVIDIA · 550B A55B · source defaultOpen weights
57.261.567.4$10.8262K$0.6 / $2.4OpenRouter catalog
52
GPT-5.4 miniOpenAI · xhigh
56.957.466.4$16.5400K$0.75 / $4.5Direct API
53
MiniMax M3MiniMax · source default
55.764.467.3$5.41.05M$0.3 / $1.2OpenRouter catalog
54
Qwen3.6 27BAlibaba · source defaultOpen weights
55.162.364.0$8.6262K$0.32 / $2.7OpenRouter catalog
55
DeepSeek V4 FlashDeepSeek · source defaultOpen weights
54.668.265.5$1.241.05M$0.0886 / $0.1772OpenRouter catalog
56
Grok 4.3xAI · source default
46.249.162.2$17.51.00M$1.25 / $2.5OpenRouter catalog

— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.

3 models in your shortlistCompare side by side ↓

LOOK BEYOND ONE NUMBER

Your shortlist, under the microscope.

Change models ↑
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5GPT-6 SolDeepSeek V4.1 Flash
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxGPT-6 SolmaxDeepSeek V4.1 Flashmax
Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.79.468.578.8
Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.61.058.483.5
Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Insufficient evidenceInsufficient evidenceInsufficient evidence
Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.$80$40$2.14
Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.1.00M1.05M1.05M
Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly.SupportedSupportedSupported
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.92.288.786.7
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.89.381.880.0
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.71.752.977.3
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.97.196.493.3
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.80.381.279.3
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.86.385.381.2
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.65.768.670.0
Input / output per 1M$4 / $20$2 / $10$0.094 / $0.6
AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗