Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

THE MODEL SELECTION SERIES

LLM API pricing calculator

API prices are usually quoted per million tokens, but your budget depends on how much context you send and how much output the model bills. Change the calculator to compare the same workload across providers.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

01 / WHAT TO CHECK

Read input and output prices together

A model priced at $2 input and $10 output per million tokens costs $40 for 1,000 calls of 10,000 input and 2,000 output tokens each. Different tokenizers can turn the same text into different token counts.

02 / WHAT TO CHECK

Reasoning, caching and long context change the bill

Include billed reasoning tokens in output. Cache reads may be cheaper while cache writes can cost more. Some providers charge higher rates for large prompts or premium service tiers. Our estimate excludes those adjustments; check the linked source before committing a budget.

03 / WHAT TO CHECK

Direct API and routed prices are different offers

We use verified direct prices for selected OpenAI and Anthropic models and explicitly labeled OpenRouter catalog rates for other models. A catalog entry can route across endpoints with different prices and limits. Each profile shows the exact catalog identifier and any observed price difference.

04 / WHAT TO CHECK

Use Practical Value as a preference, not a financial forecast

The default score combines 70% task-weighted capability with 30% affordability rank within the snapshot. Move the slider to express your tradeoff. It does not predict task success, retries, hosting charges or actual monthly savings.

YOUR WORKLOAD. YOUR PRIORITIES.

Find your model sweet spot.

How our ratings work ↗

Make the numbers yours Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.

Standard text API estimate. Include reasoning tokens in output.

A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.

THE MARKET MAP

Capability meets cost.

Efficient frontierOther modelsHigher + further left = more for less
020406080100$1$10$100$1,000Agent Fit / 100 ↑Estimated workload cost · USD · logarithmic scaleDeepSeek V4.1 Flash: 78.9 · $2.14DeepSeek V4 Flash Vision Exp: 75.1 · $3.52DeepSeek V4 Flash 0731: 71.1 · $1.68GPT-6 Luna: 67.7 · $2GPT-5.6 Luna: 70.4 · $4.4GLM-5.3 Flash: 67.2 · $2.5Qwen3.8 27B: 73.5 · $10.2DeepSeek V4 Pro 0813: 73.3 · $10.56GPT-5.4 nano: 67.7 · $4.5DeepSeek V4 Flash: 61.6 · $1.24Gemini 3.7 Flash: 76.1 · $15GLM-5.3: 73.7 · $13.68Muse Spark 1.3: 79.3 · $21GLM-5.2: 68.5 · $10.58Gemini 3.8 Flash: 73.3 · $15Qwen3.6 Plus: 63.9 · $7.15MiniMax M3: 63.1 · $5.4Muse Spark 1.2: 76.3 · $21DeepSeek V4 Pro: 67.1 · $13.37Gemini 3.6 Flash: 70.3 · $15Muse Spark 1.1: 74.0 · $21Nemotron 3 Ultra: 63.8 · $10.8Gemini 3.5 Flash-Lite: 59.5 · $8Qwen3.6 27B: 60.0 · $8.6Kimi K2.7 Code: 64.8 · $13.66Grok 4.6: 75.3 · $32Grok 4.7: 73.7 · $25.6Inkling: 68.9 · $18.1Kimi K2.6: 66.9 · $17.5Grok 4.5: 73.1 · $32Qwen3.7 Max: 70.4 · $23.6GPT-6 Sol: 74.7 · $40Claude Sonnet 5: 73.3 · $40GPT-5.4 mini: 62.5 · $16.5GPT-5.6 Terra: 74.1 · $44Claude Opus 5.5: 79.4 · $80Gemini 3.5 Flash: 70.8 · $33Kimi K3: 77.4 · $60Gemini 3.1 Pro Preview: 73.2 · $44GPT-5.4: 74.4 · $55GPT-5.6 Sol: 77.1 · $80GPT-5.2 Codex: 69.9 · $45.5GPT-5.2: 69.8 · $45.5Claude Opus 5: 75.7 · $100Claude Fable 5.1: 79.6 · $200Grok 4.3: 56.0 · $17.5Claude Fable 5: 79.0 · $200GPT-6 Astra: 78.6 · $200Claude Sonnet 4.6: 69.4 · $60GPT-5.5: 75.8 · $110Claude Opus 4.8: 73.0 · $100Claude Opus 4.7: 72.9 · $100Claude Opus 4.6: 70.5 · $100Claude Opus 4.5: 66.7 · $100
DeepSeek V4.1 Flash78.9 Agent Fit$2.14
AI AGENT STORE ORIGINAL

More capability.
Less budget.

Practical Value Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. · General agent

70% Agent Fit + 30% affordability. Rankings use your workload; filters do not change the reference cohort.

THE EVIDENCE, SIDE BY SIDE

AI model comparison table

56 of 56 models · select up to 4 to comparePrices in USD · scores explained with ?
AI model benchmarks, aggregate ratings, API pricing and context windows. Snapshot 2026-09-23.
CompareOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set.People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference.Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.Input / outputUSD / 1M tokens
01
DeepSeek V4.1 FlashDeepSeek · maxOpen weights
78.983.581.1$2.141.05M$0.094 / $0.6OpenRouter catalog
02
DeepSeek V4 Flash Vision ExpDeepSeek · experimental · source defaultOpen weightsPreview
75.179.876.8$3.521.05M$0.22 / $0.66OpenRouter catalog
03
DeepSeek V4 Flash 0731DeepSeek · source defaultOpen weights
71.179.274.2$1.681.31M$0.04 / $0.64OpenRouter catalog
04
GPT-6 LunaOpenAI · max
67.776.272.0$21.05M$0.1 / $0.5Direct API
05
GPT-5.6 LunaOpenAI · source ID max; display xhighConfiguration discrepancy
70.475.973.6$4.41.05M$0.2 / $1.2Direct API
06
GLM-5.3 FlashZ.ai · source defaultOpen weights
67.274.871.6$2.51.31M$0.15 / $0.5OpenRouter catalog
07
Qwen3.8 27BAlibaba · source defaultOpen weights
73.574.775.3$10.21.00M$0.42 / $3OpenRouter catalog
08
DeepSeek V4 Pro 0813DeepSeek · source defaultOpen weights
73.373.977.4$10.561.05M$0.66 / $1.98OpenRouter catalog
09
GPT-5.4 nanoOpenAI · xhigh
67.773.469.6$4.5400K$0.2 / $1.25Direct API
10
DeepSeek V4 FlashDeepSeek · source defaultOpen weights
61.673.165.5$1.241.05M$0.0886 / $0.1772OpenRouter catalog
11
Gemini 3.7 FlashGoogle · high
76.171.956.32 sources78.81,490±8 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
12
GLM-5.3Z.ai · source defaultOpen weights
73.771.476.1$13.681.31M$0.84 / $2.64OpenRouter catalog
13
Muse Spark 1.3Meta · xhigh
79.370.281.6$211.05M$1.25 / $4.25OpenRouter catalog
14
GLM-5.2Z.ai · source defaultOpen weights
68.570.173.2$10.581.05M$0.6496 / $2.04OpenRouter catalog
15
Gemini 3.8 FlashGoogle · high
73.370.050.02 sources75.81,493±9 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
16
Qwen3.6 PlusAlibaba · source default
63.969.668.9$7.151.00M$0.325 / $1.95OpenRouter catalog
17
MiniMax M3MiniMax · source default
63.169.667.3$5.41.05M$0.3 / $1.2OpenRouter catalog
18
Muse Spark 1.2Meta · xhigh
76.368.168.82 sources78.01,500±11$211.05M$1.25 / $4.25OpenRouter catalog
19
DeepSeek V4 ProDeepSeek · source defaultOpen weights
67.167.971.6$13.371.05M$0.9553 / $1.91OpenRouter catalog
20
Gemini 3.6 FlashGoogle · high
70.367.99.42 sources73.61,480±5$151.05M$0.75 / $3.75OpenRouter catalog
21
Muse Spark 1.1Meta · source ID xhigh; display highConfiguration discrepancy
74.066.575.3$211.05M$1.25 / $4.25OpenRouter catalog
22
Nemotron 3 UltraNVIDIA · 550B A55B · source defaultOpen weights
63.866.167.4$10.8262K$0.6 / $2.4OpenRouter catalog
59.566.063.9$81.05M$0.3 / $2.5OpenRouter catalog
24
Qwen3.6 27BAlibaba · source defaultOpen weights
60.065.864.0$8.6262K$0.32 / $2.7OpenRouter catalog
25
Kimi K2.7 CodeMoonshot AI · source defaultOpen weights
64.865.868.4$13.66262K$0.7062 / $3.3OpenRouter catalog
26
Grok 4.6xAI · source default
75.364.978.0$32500K$2 / $6OpenRouter catalog
27
Grok 4.7xAI · xhigh
73.764.677.4$25.6500K$1.6 / $4.8OpenRouter catalog
28
InklingThinking Machines · xhighOpen weights
68.964.171.9$18.11.05M$1 / $4.05OpenRouter catalog
29
Kimi K2.6Moonshot AI · thinkingOpen weights
66.963.570.5$17.5262K$0.95 / $4OpenRouter catalog
30
Grok 4.5xAI · source default
73.163.475.8$32500K$2 / $6OpenRouter catalog
31
Qwen3.7 MaxAlibaba · source default
70.462.973.1$23.61.00M$1.48 / $4.43OpenRouter catalog
32
GPT-6 SolOpenAI · max
74.762.779.2$401.05M$2 / $10Direct API
33
Claude Sonnet 5Anthropic · xhigh
73.361.876.0$401.00M$2 / $10Direct API
34
GPT-5.4 miniOpenAI · xhigh
62.561.366.4$16.5400K$0.75 / $4.5Direct API
35
GPT-5.6 TerraOpenAI · source ID max; display xhighConfiguration discrepancy
74.161.277.9$441.05M$2 / $12Direct API
36
Claude Opus 5.5Anthropic · max
79.461.083.2$801.00M$4 / $20Direct API
37
Gemini 3.5 FlashGoogle · high
70.860.912.52 sources74.61,478±4$331.05M$1.5 / $9OpenRouter catalog
38
Kimi K3Moonshot AI · source defaultOpen weights
77.460.779.2$601.05M$3 / $15OpenRouter catalog
39
Gemini 3.1 Pro PreviewGoogle · highPreview
73.260.677.0$441.05M$2 / $12OpenRouter catalog
40
GPT-5.4OpenAI · xhigh
74.459.478.0$551.05M$2.5 / $15Direct API
41
GPT-5.6 SolOpenAI · source ID max; display xhighConfiguration discrepancy
77.159.381.1$801.05M$4 / $20Direct API
42
GPT-5.2 CodexOpenAI · source default
69.957.174.0$45.5400K$1.75 / $14OpenRouter catalog
43
GPT-5.2OpenAI · high · 2025-12-11
69.857.174.6$45.5400K$1.75 / $14OpenRouter catalog
44
Claude Opus 5Anthropic · max
75.756.456.32 sources80.11,487±5$1001.00M$5 / $25Direct API
45
Claude Fable 5.1Anthropic · max
79.656.387.52 sources83.41,498±8$2001.00M$10 / $50Direct API
46
Grok 4.3xAI · source default
56.055.962.2$17.51.00M$1.25 / $2.5OpenRouter catalog
47
Claude Fable 5Anthropic · source ID max; display xhighConfiguration discrepancy
79.055.883.0$2001.00M$10 / $50Direct API
48
GPT-6 AstraOpenAI · max
78.655.653.12 sources82.21,480±12$2001.05M$10 / $50Direct API
49
Claude Sonnet 4.6Anthropic · medium · adaptive thinking
69.455.173.0$601.00M$3 / $15Direct API
50
GPT-5.5OpenAI · xhigh
75.854.880.2$1101.05M$5 / $30Direct API
51
Claude Opus 4.8Anthropic · max
73.054.576.2$1001.00M$5 / $25Direct API
52
Claude Opus 4.7Anthropic · xhigh
72.954.476.5$1001.00M$5 / $25Direct API
53
Claude Opus 4.6Anthropic · high · adaptive thinking
70.552.856.32 sources74.51,505±4$1001.00M$5 / $25Direct API
54
Claude Opus 4.5Anthropic · high · 64K thinking
66.750.172.6$100200K$5 / $25Direct API
55
Qwen3.8 Flash NextAlibaba · source defaultOpen weights
76.276.2Not verifiedNot verified
56
Qwen3.8 MaxAlibaba · source default · version mapping unconfirmed
77.078.5Not verifiedNot verified

— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.

3 models in your shortlistCompare side by side ↓

LOOK BEYOND ONE NUMBER

Your shortlist, under the microscope.

Change models ↑
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5GPT-6 SolDeepSeek V4.1 Flash
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxGPT-6 SolmaxDeepSeek V4.1 Flashmax
Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.79.474.778.9
Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.61.062.783.5
Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Insufficient evidenceInsufficient evidenceInsufficient evidence
Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.$80$40$2.14
Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.1.00M1.05M1.05M
Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly.SupportedSupportedSupported
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.92.288.786.7
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.89.381.880.0
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.71.752.977.3
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.97.196.493.3
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.80.381.279.3
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.86.385.381.2
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.65.768.670.0
Input / output per 1M$4 / $20$2 / $10$0.094 / $0.6
AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗