Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

MODEL INTELLIGENCE · SEPTEMBER 2026

AI models, compared.
Choose with confidence.

The model market, made understandable. Explore benchmarks, compare costs, and find the right intelligence for your next AI agent.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Our snapshot take

Claude Fable 5.1 leads the default Agent Fit mix at 79.6. Change the workload and priorities below to find where capability is worth its cost.

Explore the evidence ↓

YOUR WORKLOAD. YOUR PRIORITIES.

Find your model sweet spot.

How our ratings work ↗

Make the numbers yours Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.

Standard text API estimate. Include reasoning tokens in output.

A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.

THE MARKET MAP

Capability meets cost.

Efficient frontierOther modelsHigher + further left = more for less
020406080100$1$10$100$1,000Agent Fit / 100 ↑Estimated workload cost · USD · logarithmic scaleClaude Fable 5.1: 79.6 · $200Claude Opus 5.5: 79.4 · $80Muse Spark 1.3: 79.3 · $21Claude Fable 5: 79.0 · $200DeepSeek V4.1 Flash: 78.9 · $2.14GPT-6 Astra: 78.6 · $200Kimi K3: 77.4 · $60GPT-5.6 Sol: 77.1 · $80Muse Spark 1.2: 76.3 · $21Gemini 3.7 Flash: 76.1 · $15GPT-5.5: 75.8 · $110Claude Opus 5: 75.7 · $100Grok 4.6: 75.3 · $32DeepSeek V4 Flash Vision Exp: 75.1 · $3.52GPT-6 Sol: 74.7 · $40GPT-5.4: 74.4 · $55GPT-5.6 Terra: 74.1 · $44Muse Spark 1.1: 74.0 · $21GLM-5.3: 73.7 · $13.68Grok 4.7: 73.7 · $25.6Qwen3.8 27B: 73.5 · $10.2Gemini 3.8 Flash: 73.3 · $15Claude Sonnet 5: 73.3 · $40DeepSeek V4 Pro 0813: 73.3 · $10.56Gemini 3.1 Pro Preview: 73.2 · $44Grok 4.5: 73.1 · $32Claude Opus 4.8: 73.0 · $100Claude Opus 4.7: 72.9 · $100DeepSeek V4 Flash 0731: 71.1 · $1.68Gemini 3.5 Flash: 70.8 · $33Claude Opus 4.6: 70.5 · $100Qwen3.7 Max: 70.4 · $23.6GPT-5.6 Luna: 70.4 · $4.4Gemini 3.6 Flash: 70.3 · $15GPT-5.2 Codex: 69.9 · $45.5GPT-5.2: 69.8 · $45.5Claude Sonnet 4.6: 69.4 · $60Inkling: 68.9 · $18.1GLM-5.2: 68.5 · $10.58GPT-5.4 nano: 67.7 · $4.5GPT-6 Luna: 67.7 · $2GLM-5.3 Flash: 67.2 · $2.5DeepSeek V4 Pro: 67.1 · $13.37Kimi K2.6: 66.9 · $17.5Claude Opus 4.5: 66.7 · $100Kimi K2.7 Code: 64.8 · $13.66Qwen3.6 Plus: 63.9 · $7.15Nemotron 3 Ultra: 63.8 · $10.8MiniMax M3: 63.1 · $5.4GPT-5.4 mini: 62.5 · $16.5DeepSeek V4 Flash: 61.6 · $1.24Qwen3.6 27B: 60.0 · $8.6Gemini 3.5 Flash-Lite: 59.5 · $8Grok 4.3: 56.0 · $17.5
Claude Opus 5.579.4 Agent Fit$80
AI AGENT STORE ORIGINAL

More capability.
Less budget.

Practical Value Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. · General agent

70% Agent Fit + 30% affordability. Rankings use your workload; filters do not change the reference cohort.

THE EVIDENCE, SIDE BY SIDE

AI model comparison table

56 of 56 models · select up to 4 to comparePrices in USD · scores explained with ?
AI model benchmarks, aggregate ratings, API pricing and context windows. Snapshot 2026-09-23.
CompareOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set.People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference.Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.Input / outputUSD / 1M tokens
01
Claude Fable 5.1Anthropic · max
79.656.387.52 sources83.41,498±8$2001.00M$10 / $50Direct API
02
Claude Opus 5.5Anthropic · max
79.461.083.2$801.00M$4 / $20Direct API
03
Muse Spark 1.3Meta · xhigh
79.370.281.6$211.05M$1.25 / $4.25OpenRouter catalog
04
Claude Fable 5Anthropic · source ID max; display xhighConfiguration discrepancy
79.055.883.0$2001.00M$10 / $50Direct API
05
DeepSeek V4.1 FlashDeepSeek · maxOpen weights
78.983.581.1$2.141.05M$0.094 / $0.6OpenRouter catalog
06
GPT-6 AstraOpenAI · max
78.655.653.12 sources82.21,480±12$2001.05M$10 / $50Direct API
07
Kimi K3Moonshot AI · source defaultOpen weights
77.460.779.2$601.05M$3 / $15OpenRouter catalog
08
GPT-5.6 SolOpenAI · source ID max; display xhighConfiguration discrepancy
77.159.381.1$801.05M$4 / $20Direct API
09
Qwen3.8 MaxAlibaba · source default · version mapping unconfirmed
77.078.5Not verifiedNot verified
10
Muse Spark 1.2Meta · xhigh
76.368.168.82 sources78.01,500±11$211.05M$1.25 / $4.25OpenRouter catalog
11
Qwen3.8 Flash NextAlibaba · source defaultOpen weights
76.276.2Not verifiedNot verified
12
Gemini 3.7 FlashGoogle · high
76.171.956.32 sources78.81,490±8 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
13
GPT-5.5OpenAI · xhigh
75.854.880.2$1101.05M$5 / $30Direct API
14
Claude Opus 5Anthropic · max
75.756.456.32 sources80.11,487±5$1001.00M$5 / $25Direct API
15
Grok 4.6xAI · source default
75.364.978.0$32500K$2 / $6OpenRouter catalog
16
DeepSeek V4 Flash Vision ExpDeepSeek · experimental · source defaultOpen weightsPreview
75.179.876.8$3.521.05M$0.22 / $0.66OpenRouter catalog
17
GPT-6 SolOpenAI · max
74.762.779.2$401.05M$2 / $10Direct API
18
GPT-5.4OpenAI · xhigh
74.459.478.0$551.05M$2.5 / $15Direct API
19
GPT-5.6 TerraOpenAI · source ID max; display xhighConfiguration discrepancy
74.161.277.9$441.05M$2 / $12Direct API
20
Muse Spark 1.1Meta · source ID xhigh; display highConfiguration discrepancy
74.066.575.3$211.05M$1.25 / $4.25OpenRouter catalog
21
GLM-5.3Z.ai · source defaultOpen weights
73.771.476.1$13.681.31M$0.84 / $2.64OpenRouter catalog
22
Grok 4.7xAI · xhigh
73.764.677.4$25.6500K$1.6 / $4.8OpenRouter catalog
23
Qwen3.8 27BAlibaba · source defaultOpen weights
73.574.775.3$10.21.00M$0.42 / $3OpenRouter catalog
24
Gemini 3.8 FlashGoogle · high
73.370.050.02 sources75.81,493±9 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
25
Claude Sonnet 5Anthropic · xhigh
73.361.876.0$401.00M$2 / $10Direct API
26
DeepSeek V4 Pro 0813DeepSeek · source defaultOpen weights
73.373.977.4$10.561.05M$0.66 / $1.98OpenRouter catalog
27
Gemini 3.1 Pro PreviewGoogle · highPreview
73.260.677.0$441.05M$2 / $12OpenRouter catalog
28
Grok 4.5xAI · source default
73.163.475.8$32500K$2 / $6OpenRouter catalog
29
Claude Opus 4.8Anthropic · max
73.054.576.2$1001.00M$5 / $25Direct API
30
Claude Opus 4.7Anthropic · xhigh
72.954.476.5$1001.00M$5 / $25Direct API
31
DeepSeek V4 Flash 0731DeepSeek · source defaultOpen weights
71.179.274.2$1.681.31M$0.04 / $0.64OpenRouter catalog
32
Gemini 3.5 FlashGoogle · high
70.860.912.52 sources74.61,478±4$331.05M$1.5 / $9OpenRouter catalog
33
Claude Opus 4.6Anthropic · high · adaptive thinking
70.552.856.32 sources74.51,505±4$1001.00M$5 / $25Direct API
34
Qwen3.7 MaxAlibaba · source default
70.462.973.1$23.61.00M$1.48 / $4.43OpenRouter catalog
35
GPT-5.6 LunaOpenAI · source ID max; display xhighConfiguration discrepancy
70.475.973.6$4.41.05M$0.2 / $1.2Direct API
36
Gemini 3.6 FlashGoogle · high
70.367.99.42 sources73.61,480±5$151.05M$0.75 / $3.75OpenRouter catalog
37
GPT-5.2 CodexOpenAI · source default
69.957.174.0$45.5400K$1.75 / $14OpenRouter catalog
38
GPT-5.2OpenAI · high · 2025-12-11
69.857.174.6$45.5400K$1.75 / $14OpenRouter catalog
39
Claude Sonnet 4.6Anthropic · medium · adaptive thinking
69.455.173.0$601.00M$3 / $15Direct API
40
InklingThinking Machines · xhighOpen weights
68.964.171.9$18.11.05M$1 / $4.05OpenRouter catalog
41
GLM-5.2Z.ai · source defaultOpen weights
68.570.173.2$10.581.05M$0.6496 / $2.04OpenRouter catalog
42
GPT-5.4 nanoOpenAI · xhigh
67.773.469.6$4.5400K$0.2 / $1.25Direct API
43
GPT-6 LunaOpenAI · max
67.776.272.0$21.05M$0.1 / $0.5Direct API
44
GLM-5.3 FlashZ.ai · source defaultOpen weights
67.274.871.6$2.51.31M$0.15 / $0.5OpenRouter catalog
45
DeepSeek V4 ProDeepSeek · source defaultOpen weights
67.167.971.6$13.371.05M$0.9553 / $1.91OpenRouter catalog
46
Kimi K2.6Moonshot AI · thinkingOpen weights
66.963.570.5$17.5262K$0.95 / $4OpenRouter catalog
47
Claude Opus 4.5Anthropic · high · 64K thinking
66.750.172.6$100200K$5 / $25Direct API
48
Kimi K2.7 CodeMoonshot AI · source defaultOpen weights
64.865.868.4$13.66262K$0.7062 / $3.3OpenRouter catalog
49
Qwen3.6 PlusAlibaba · source default
63.969.668.9$7.151.00M$0.325 / $1.95OpenRouter catalog
50
Nemotron 3 UltraNVIDIA · 550B A55B · source defaultOpen weights
63.866.167.4$10.8262K$0.6 / $2.4OpenRouter catalog
51
MiniMax M3MiniMax · source default
63.169.667.3$5.41.05M$0.3 / $1.2OpenRouter catalog
52
GPT-5.4 miniOpenAI · xhigh
62.561.366.4$16.5400K$0.75 / $4.5Direct API
53
DeepSeek V4 FlashDeepSeek · source defaultOpen weights
61.673.165.5$1.241.05M$0.0886 / $0.1772OpenRouter catalog
54
Qwen3.6 27BAlibaba · source defaultOpen weights
60.065.864.0$8.6262K$0.32 / $2.7OpenRouter catalog
59.566.063.9$81.05M$0.3 / $2.5OpenRouter catalog
56
Grok 4.3xAI · source default
56.055.962.2$17.51.00M$1.25 / $2.5OpenRouter catalog

— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.

3 models in your shortlistCompare side by side ↓

LOOK BEYOND ONE NUMBER

Your shortlist, under the microscope.

Change models ↑
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5GPT-6 SolDeepSeek V4.1 Flash
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxGPT-6 SolmaxDeepSeek V4.1 Flashmax
Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.79.474.778.9
Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.61.062.783.5
Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Insufficient evidenceInsufficient evidenceInsufficient evidence
Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.$80$40$2.14
Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.1.00M1.05M1.05M
Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly.SupportedSupportedSupported
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.92.288.786.7
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.89.381.880.0
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.71.752.977.3
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.97.196.493.3
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.80.381.279.3
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.86.385.381.2
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.65.768.670.0
Input / output per 1M$4 / $20$2 / $10$0.094 / $0.6

START WITH YOUR QUESTION

Choose a model for the work you do.

Popular model comparisons

LESS JARGON. BETTER DECISIONS.

AI model questions,
plain-English answers.

Explore the benchmark field guide ↗
What is the best AI model for an AI agent?+

There is no universal winner. Compare the task profile, tool support, reasoning configuration and total workload cost. Our Agent Fit score is a task-weighted shortlist; validate it using real tasks and your own tools.

How does AI Agent Store combine model benchmarks?+

Agent Fit uses transparent weights over seven LiveBench categories. Evidence Consensus averages percentile ranks from LiveBench and Arena for nine explicitly matched configurations. Practical Value combines Agent Fit with a cost percentile derived from sourced API prices. Missing evidence is never filled with a guessed score.

Why do some models have no Consensus score?+

The snapshot has no verified matching Arena configuration for those models. Reasoning effort and version matter, so a result for a different setting is not silently reused. A missing score does not mean the model is weak.

Are these AI model rankings live?+

This is a reviewed snapshot dated 23 September 2026, not a live feed. LiveBench uses its 25 June 2026 question set with model observations collected on the snapshot date. Arena and other sources have their own update dates, shown in the methodology.

What does an AI model context window mean?+

It is the advertised token capacity available to a request and its response. A large window lets you supply more information, but does not prove the model will retrieve or reason about every detail reliably.

Does a high benchmark score guarantee a reliable agent?+

No. Reliability also depends on prompts, tool design, permissions, retrieval, retry policies and evaluation. These pages compare models and reported evidence, not complete agent systems.

TURN YOUR SHORTLIST INTO SOMETHING USEFUL

A model is the beginning.
Give it a job to do.

Build an AI worker ↗
AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗