Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

A FIELD GUIDE TO THE NUMBERS

Less benchmark jargon.
More useful answers.

A benchmark asks a model to do a particular kind of work. Knowing what the test asks is more useful than knowing who came first.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

AI AGENT STORE RATING

Agent Fit

Our task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.

AI AGENT STORE RATING

Practical Value

Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.

AI AGENT STORE RATING

Evidence Consensus

Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.

MEASURE EXPLAINED

LiveBench

Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set.

MEASURE EXPLAINED

Arena preference

People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference.

MEASURE EXPLAINED

Reasoning

Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.

MEASURE EXPLAINED

Coding

Can it write or complete code that passes tests? These are contained programming problems, not entire software projects.

MEASURE EXPLAINED

Agentic coding

Can it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.

MEASURE EXPLAINED

Mathematics

Can it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.

MEASURE EXPLAINED

Data analysis

Can it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.

MEASURE EXPLAINED

Language

Can it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.

MEASURE EXPLAINED

Instruction following

Can it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.

MEASURE EXPLAINED

Context window

The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.

MEASURE EXPLAINED

Workload cost

Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.

MEASURE EXPLAINED

Tool calling

The catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly.

MEASURE EXPLAINED

BFCL V4

Berkeley tests selecting functions, supplying arguments, multi-turn tool use, search and memory. Its older model configurations are shown separately and never substituted for newer releases.

A DIFFERENT TEST. A DIFFERENT SNAPSHOT.

Tool-use evidence: Berkeley BFCL V4

These published configurations are from an older, separate evaluation, last updated 12 April 2026. They are useful context for understanding tool use; they do not measure the latest versions in our model table and are excluded from our aggregate ratings. FC means native function calling; Prompt means text-based prompting.

Published BFCL V4 results · Source and methodology ↗
Tested configurationOverall accuracy Berkeley tests selecting functions, supplying arguments, multi-turn tool use, search and memory. Its older model configurations are shown separately and never substituted for newer releases.Multi-turnMemory
Claude-Opus-4-5-20251101 (FC)77.47%68.38%73.76%
Claude-Sonnet-4-5-20250929 (FC)73.24%61.37%64.95%
Gemini-3-Pro-Preview (Prompt)72.51%60.75%61.72%
GLM-4.6 (FC thinking)72.38%68.00%55.70%
Grok-4-1-fast-reasoning (FC)69.57%58.87%53.98%
Claude-Haiku-4-5-20251001 (FC)68.70%53.62%54.41%
Gemini-3-Pro-Preview (FC)68.14%63.12%54.84%
o3-2025-04-16 (Prompt)63.05%62.25%51.83%
Grok-4-0709 (Prompt)62.97%47.00%50.54%
Grok-4-0709 (FC)61.38%33.88%55.91%

GROUNDED ANSWERS NEED THEIR OWN TEST

Hallucination is not one universal number.

Vectara’s document-summary evaluation checks whether an answer stays supported by the supplied document. In its 11 May 2026 snapshot, GPT-5.4 nano’s reported hallucination rate is 3.1%, GPT-5.4 mini’s is 5.5%, GPT-5.4’s is 7.0% and GPT-5.5’s is 9.3%. These observations use Vectara’s named configurations, not the reasoning configurations in our LiveBench table.

These rates do not predict factual accuracy on open-ended questions or autonomous tasks. Answer rate and the evaluation judge also matter. We therefore keep them out of our scores. Read the dataset, configurations and limitations ↗

Now put the evidence to work.

Compare models with explanations beside every metric.

Explore the model comparison ↗
AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗