Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗
LOW → MEDIUM → HIGH → EXTRA HIGH → MAX

Find the thinking level
that earns its cost.

How much does extra reasoning help? Compare the same model at different effort settings. Spot the useful gains, the expensive plateaus and the times more effort scores lower.

36models
45matched sweeps
156tested configurations
4benchmark sources
Open GPT-6 Sol’s dedicated page ↗
SAME MODEL. DIFFERENT THINKING.

Find GPT-6 Sol’s useful effort.

Download CSV ↓
Hex DataBench v1.1

Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.

Analytics in the Hex agent harness · Leaderboard updated: 2026-09-22 · Reviewed September 23, 2026Check the source ↗
YOUR NEAR-BEST PICKExtra high

Lowest measured cost within 1.0 points of the best tested analytics score.

61.3 / 100 · $0.5817 / source task
Exact best1.0 pts10 pts

Set 0 for the highest score. Increase the gap to explore cheaper or lighter options.

THE EFFORT CURVE

Does more thinking help?

● Near-best pickOrange = score decreased

025507510050.3Low51.3Medium54.7High61.3Xhigh57.0Max
025507510050.3Low51.3Medium54.7High61.3Xhigh57.0Max

Scores out of 100 · Full 0–100 scale · tap a point to select mode B. Cyan band: within 1.0 points of the best tested score.

WHAT YOU SPEND

Cost of more effort

Measured average USD per benchmark task.

Low$0.1748
Medium$0.2708
High$0.3874
Extra highNear-best pick$0.5817
Max$0.9845

Within this benchmark and harness. Your task mix, tools, caching and endpoint can change the result.

Score difference · B minus A-4.3 pts

Noticeable gap · B scored lower

Measured cost · B / A1.69×

Same benchmark workload

Median duration · B / A1.89×

Measured in the source harness

Gap labels are editorial: under 1 point is small, 1–under 5 noticeable, 5+ large. They do not establish statistical significance. A tied score does not prove equivalent behavior.

WHERE MORE SCORED LOWER

Watch these effort increases.

Extra highMax-4.3 pts61.357.0 on analytics score

Observed reversals within this sweep. They can reflect task fit, sampling, harness behavior or errors; the score alone does not prove “overthinking.”
EVERY TESTED MODE

The numbers behind the choice

5 modes · missing levels are not interpolated

Low

Tested mode
50.3 / 100

-11.0 pts vs Extra high

$0.1748 / source task

98 s median duration

Medium

Tested mode
51.3 / 100

-10.0 pts vs Extra high

$0.2708 / source task

146 s median duration

High

Tested mode
54.7 / 100

-6.7 pts vs Extra high

$0.3874 / source task

238 s median duration

Extra high

Near-best pick
61.3 / 100

0.0 pts vs Extra high

$0.5817 / source task

327 s median duration

Max

Another mode trades better
57.0 / 100

-4.3 pts vs Extra high

$0.9845 / source task

618 s median duration

Full benchmark table & exact configuration IDs

Swipe inside the table for all measures. Figures from different study selections are not directly comparable.

GPT-6 Sol · Hex DataBench v1.1 · Analytics in the Hex agent harness
EffortAnalytics scoreCost / taskMedian secondsOutput tokensExact source ID
Low50.3$0.174897.6GPT-6 Sol · Low
Medium51.3$0.2708146.2GPT-6 Sol · Medium
High54.7$0.3874238.0GPT-6 Sol · High
Extra high61.3$0.5817326.7GPT-6 Sol · XHigh
Max57.0$0.9845617.5GPT-6 Sol · Max
SCALE THE OBSERVED COST

Effort cost difference calculator

What would A versus B cost if you repeated this benchmark’s average workload?

A · Extra high$581.65
B · Max$984.49
Difference$402.83more with B

A source-workload scenario using average measured USD costs. Source prices, caching and tools are already reflected in those observed costs. Your agent’s costs may differ.

YOUR WORKLOAD

Reasoning-token cost calculator

Estimate the extra bill from thinking tokens. The starting token counts are editable examples, not measurements of GPT-6 Sol. Rates start from our September 23 model snapshot.

A · Extra high$40estimated monthly cost
B · Max$80estimated monthly cost

Mode B costs $40 more per month with these assumptions.

Standard text-token arithmetic. Excludes retries, tools, caching, long-context premiums, hosting and taxes. Includes thinking in output billing; do not also include it in “visible output.” Check endpoint limits separately. Changing mode does not automatically predict tokens or quality.

START WITH THE MODEL YOU USE

36 models. Their actual effort curves.

36 models shown
Anthropic · 1 sweep

Claude Fable 5

LowMediumHighMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 3.3 points.

Hex DataBench v1.1
Anthropic · 1 sweep

Claude Fable 5.1

LowMediumHighExtra highMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 10.7 points.

Hex DataBench v1.1
Anthropic · 2 sweeps

Claude Opus 4.5

LowMediumHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 11.9 points.

LiveBench
Anthropic · 1 sweep

Claude Opus 4.7

LowMediumHighExtra high

Extra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 8.1 points.

LiveBench
Anthropic · 3 sweeps

Claude Opus 4.8

LowMediumHighExtra highMax

Extra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 3.7 points.

Hex DataBench v1.1 · LiveBench · OckBench
Anthropic · 2 sweeps

Claude Opus 5

LowMediumHighExtra highMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 4.0 points.

Hex DataBench v1.1 · OckBench
Anthropic · 2 sweeps

Claude Opus 5.5

LowMediumHighExtra highMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 16.5 points.

Hex DataBench v1.1 · LiveBench
Anthropic · 1 sweep

Claude Sonnet 4.6

LowMediumHigh

Medium is the lowest-effort option within 1 point of the best general agent fit in LiveBench. High trails the best tested score by 0.2 points.

LiveBench
Anthropic · 1 sweep

Claude Sonnet 5

LowMediumHighExtra highMax

Medium is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 2.0 points.

Hex DataBench v1.1
DeepSeek · 1 sweep

DeepSeek V4 Flash

NoneHighMax

Max is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 62.5 points.

OckBench
DeepSeek · 1 sweep

DeepSeek V4 Pro

NoneHighMax

High is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 59.0 points.

OckBench
DeepSeek · 1 sweep

DeepSeek V4.1 Flash

LowHigh

Low is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 0.3 points.

Hex DataBench v1.1
Google · 1 sweep

Gemini 3 Flash Preview

MinimalHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 22.1 points.

LiveBench
Google · 1 sweep

Gemini 3 Pro Preview

LowHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 13.3 points.

LiveBench
Google · 1 sweep

Gemini 3.1 Pro Preview

LowMedium

Medium is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 8.0 points.

OckBench
OpenAI · 1 sweep

GPT-5

MinimalLowHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 21.3 points.

LiveBench
OpenAI · 1 sweep

GPT-5 mini

MinimalLowHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 28.7 points.

LiveBench
OpenAI · 1 sweep

GPT-5 nano

LowHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 15.1 points.

LiveBench
OpenAI · 1 sweep

GPT-5.1

LowMediumHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 15.3 points.

LiveBench
OpenAI · 1 sweep

GPT-5.2

LowMediumHigh

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 9.8 points.

LiveBench
OpenAI · 1 sweep

GPT-5.3 Codex

HighExtra high

High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. Extra high trails the best tested score by 0.8 points.

LiveBench
OpenAI · 2 sweeps

GPT-5.4

NoneLowMediumHighExtra high

Extra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 7.0 points.

LiveBench · OckBench
OpenAI · 1 sweep

GPT-5.4 mini

LowMediumHighExtra high

Extra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 23.1 points.

LiveBench
OpenAI · 1 sweep

GPT-5.4 nano

LowMediumHighExtra high

Extra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 25.4 points.

LiveBench
OpenAI · 3 sweeps

GPT-5.5

NoneLowMediumHighExtra high

Extra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 11.0 points.

LiveBench · OckBench · Stet coding case study
OpenAI · 1 sweep

GPT-5.6 Luna

LowMediumHighExtra high

Extra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 16.3 points.

Hex DataBench v1.1
OpenAI · 1 sweep

GPT-5.6 Sol

LowMediumHighExtra high

Medium is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 10.7 points.

Hex DataBench v1.1
OpenAI · 1 sweep

GPT-5.6 Terra

LowMediumHighExtra high

Extra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 12.3 points.

Hex DataBench v1.1
OpenAI · 1 sweep

GPT-6 Astra

LowMediumHighExtra highMax

Low is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 3.0 points.

Hex DataBench v1.1
OpenAI · 1 sweep

GPT-6 Luna

LowMediumHighExtra highMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 22.0 points.

Hex DataBench v1.1
OpenAI · 1 sweep

GPT-6 Sol

LowMediumHighExtra highMax

Extra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 4.3 points.

Hex DataBench v1.1
OpenAI · 1 sweep

o3

MediumHigh

High is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 1.5 points.

LiveBench
OpenAI · 1 sweep

o3-mini

LowMediumHigh

High is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 11.6 points.

LiveBench
OpenAI · 1 sweep

o4-mini

MediumHigh

High is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 4.3 points.

LiveBench
Z.ai · 2 sweeps

GLM 5.2

OffHighMax

High is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 4.3 points.

Hex DataBench v1.1 · OckBench
Z.ai · 1 sweep

GLM 5.3

LowHighMax

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 7.7 points.

Hex DataBench v1.1

Reasoning effort, explained

Is higher reasoning effort always better?

No. Some published sweeps show a lower score at a higher effort. For example, GPT-6 Sol scores 61.3 at extra high and 57.0 at max in Hex DataBench v1.1. That is a 4.3-point observed decline in this analytics harness. It does not prove that max is worse on every task.

Which reasoning effort should I choose for coding, research or analytics?

Choose a matching benchmark and task measure first. LiveBench offers coding, research and other task-weighted profiles; OckBench separates math, coding and science; DataBench covers analytics agents. Our near-best pick selects the lowest measured cost, otherwise lowest measured output tokens, otherwise lowest tested effort within your allowed score gap. Validate the shortlist on your own tasks.

What does extra high or xhigh mean?

Xhigh means extra high reasoning effort. It asks a supported model to invest more effort than high, but it is not a fixed token budget. Available settings and defaults depend on the model, provider and tool. The same label is not a comparable quantity across providers.

Is a one-point benchmark difference meaningful?

It depends on the benchmark, task count and repeatability. We call gaps under 1 point small, 1 to under 5 noticeable, and 5 or more large as a reading aid. Those labels and the default 1-point tolerance are editorial choices, not statistical significance tests. The sources generally do not supply per-mode confidence intervals.

Why are some effort levels or costs missing?

Only explicitly identified, published configurations enter a sweep. Missing modes are not estimated, and unknown defaults are excluded. We show dollar costs or duration only when the same source measured them. Output tokens alone cannot determine the full bill.

Can I estimate the cost of low versus high effort?

Yes. The measured-cost calculator scales the average cost of a benchmark task. The separate reasoning-token calculator uses your own token assumptions and editable input/output prices. A mode label alone cannot predict its token usage, cost or score on your workload.

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗