Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

OPEN FORMULAS. VISIBLE TRADEOFFS.

Know what’s
behind the number.

Our ratings are a transparent interpretation of published evidence. We do not run these third-party benchmarks ourselves, and no single number captures everything a model can do.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

1. Agent Fit: capability for a specific job

Agent Fit is our weighted mean of LiveBench category scores, on a 0–100 scale. The source reports 23 tasks grouped into seven categories. We first average the task scores within each category, then apply the chosen profile’s weights. All 56 configurations use the same 2026-06-25 question set. This date identifies the benchmark release; it is not the release date of every model.

Agent Fit = Σ(category score × profile weight) / 100

A higher score means better performance on this selected task mix. It is not a percentage of successful agent runs. The weights are our editorial choices, made explicit below.

CategoryGeneral agentCoding assistantResearch & analysisWriting & support
Reasoning30%15%35%15%
Coding10%30%0%0%
Agentic coding20%45%0%0%
Mathematics0%0%0%0%
Data analysis15%0%30%0%
Language0%0%15%40%
Instruction following25%10%20%45%

LiveBench Overall remains visible as the equal-weight mean of all seven categories. Mathematics has no separate weight in our task profiles; it remains visible for readers who need quantitative problem solving.

2. Practical Value: capability with a budget

We combine Agent Fit with affordability. Affordability is the inverse percentile rank of the estimated workload cost among all priced models in this snapshot. The default gives 70% to capability and 30% to affordability. You can change both the token workload and this tradeoff.

Practical Value = Agent Fit × (1 − cost weight) + affordability percentile × cost weight

The monthly estimate is calls × (input tokens × input rate + output tokens × output rate) / 1,000,000. For example, $2 input and $10 output rates with 10,000 input tokens, 2,000 output tokens and 1,000 calls produce a $40 estimate.

Include all billed reasoning in output. Estimates exclude caching, cache writes, search or other tool fees, retries, image/audio charges, hosting, taxes, long-context premiums and special service tiers. Input and output token counts are assumed identical across models even though tokenizers differ. We do not estimate success-adjusted cost without real workflow measurements.

We prefer checked direct standard API rates for the OpenAI and Anthropic models listed in the snapshot. Other prices are labeled OpenRouter catalog rates. Model pages retain both observations when direct and catalog prices differ. Context and tool-support fields refer to the catalog endpoint; verify them with the endpoint you purchase.

3. Evidence Consensus: agreement across two lenses

Nine configurations have an explicitly matched Arena text-preference result and a LiveBench result. Within that fixed shared cohort, we calculate a percentile for each benchmark and average the two percentiles equally. This prevents directly averaging an Arena rating around 1,500 with a benchmark percentage around 80.

Consensus = (LiveBench percentile + Arena percentile) / 2

Percentile = 100 × (number below + (number tied − 1) / 2) / (cohort size − 1). Identical values share a midrank. A one-model cohort or an all-equal cohort receives 50. Affordability uses the same rule with lower costs ranking higher.

This is a nine-configuration comparison, not a whole-market intelligence ranking. Missing pairs receive no Consensus score. Filter controls never change the reference cohort. An updated snapshot can change percentiles even if an individual raw score stays the same.

Arena measures human preference, which can reward style. LiveBench measures performance on objective tasks. They are different signals, and combining them is an editorial choice. We display Arena’s published ± range and vote count on profiles; those uncertainty values are not propagated into a confidence interval for our aggregate. Preliminary results remain labeled. Small rank differences should not be treated as decisive.

Model identity and missing evidence

Each model profile preserves the exact LiveBench row identifier, reasoning configuration, catalog ID, and matched Arena identifier when present. Only explicit configuration matches enter Consensus. A default or unknown reasoning setting is not assumed to equal “high” or “max.”

Some LiveBench identifiers and display labels disagree about effort. We expose that discrepancy and exclude those configurations from Consensus. Qwen3.8 Max’s version-specific catalog mapping is unconfirmed, so its price and catalog capabilities are withheld. No approximate value is silently inserted.

Missing means unavailable or unverified, never zero. A model without a price receives neither a cost estimate nor a Practical Value score. The same applies when the supplied input plus billed output exceeds its advertised catalog context window. Tool-calling support is a declared feature, not a tested reliability rating. “Open weights” describes availability of parameters, not a blanket license to use them.

The efficient frontier connects models for which no visible model is both cheaper and at least as capable, with one strict improvement. It depends on the selected metric, workload and filters. It is not a recommendation to choose an extreme without testing.

Source register

Download this reviewed snapshot (JSON) to inspect the observations, source links and source-file hashes. Third-party observations remain attributed to their publishers.

Snapshot reviewed on 2026-09-23. Eight research and data sources inform this section; they do not all contribute to every score. Source update dates and our retrieval date are separate. Artificial Analysis was researched as a design reference; its proprietary scores and speed measurements are not republished or used in these formulas.

Independent evaluation

LiveBench

23 tasks in 7 categories. Release date identifies the question set; model rows were retrieved on the snapshot date.

Source date / release: 2026-06-25 · Retrieved: 2026-09-23Public source data ↗
Human preference

Arena

Human preference scores, vote counts and published uncertainty. Only explicitly matching reasoning configurations enter Consensus.

Source date / release: 2026-09-13 · Retrieved: 2026-09-23
API catalog

OpenRouter

Catalog API prices, context windows, modalities and declared tool support. Routed prices can differ from direct providers.

Source date / release: 2026-09-23 · Retrieved: 2026-09-23
Official provider

OpenAI pricing

Standard text input and output rates; long-context and service-tier premiums may apply.

Source date / release: 2026-09-23 · Retrieved: 2026-09-23
Official provider

Anthropic pricing

Standard direct API input and output rates; caching, regional and tool fees are separate.

Source date / release: 2026-09-23 · Retrieved: 2026-09-23
Pricing reference

Google pricing

Explains Gemini paid tiers, thinking-token billing and long-context pricing. Catalog prices are shown in this snapshot.

Source date / release: 2026-09-23 · Retrieved: 2026-09-23
Independent evaluation

Berkeley Function Calling Leaderboard

Separate tool-use evaluation for older model versions. Not merged into current-model ratings.

Source date / release: 2026-04-12 · Retrieved: 2026-09-23Public source data ↗
Independent evaluation

Vectara hallucination leaderboard

Document-summary hallucination evaluation. A different task and configuration; contextual evidence only.

Source date / release: 2026-05-11 · Retrieved: 2026-09-23

What these rankings cannot tell you

The selection contains models covered by the chosen LiveBench release, including older versions still useful for comparison. It is not an exhaustive inventory of every model. Newly launched models, private models and specialized image/audio models may be absent. Current availability and retirement status must be checked with the provider.

We do not publish a throughput or latency ranking here because this snapshot lacks a consistently matched and reusable performance dataset. Those measurements depend on endpoint, load, prompt length, caching and reasoning effort. We also do not treat the older Berkeley and Vectara evaluations as measurements of newer model releases.

Benchmark questions can become familiar to models, category coverage is incomplete, and published harnesses may differ from yours. These scores say nothing definitive about security, privacy, data residency, production reliability or suitability for a sensitive application.

Updates and corrections

This is a versioned, reviewed snapshot, not an automatically refreshing feed. Updates require checking source versions, model mappings, rates and calculations before publication. The current scoring methodology is version 1.0. Source changes should never silently rewrite an older observation’s date.

To report a correction, contact AI Agent Store with the model name, field, current source URL and observation date. The calculator and rating formulas contain no paid-placement or affiliate-commission input.

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗