Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗
Anthropic · REASONING EFFORT GUIDE

Claude Opus 5
How much thinking?

Compare Low vs Medium vs High vs Extra high vs Max where tested. Find the effort setting that fits your task, and see how much extra reasoning changes the result.

SAME MODEL. DIFFERENT THINKING.

Find Claude Opus 5’s useful effort.

Download CSV ↓
Hex DataBench v1.1

Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.

Analytics in the Hex agent harness · Leaderboard updated: 2026-09-22 · Reviewed September 23, 2026Check the source ↗
YOUR NEAR-BEST PICKMax

Lowest measured cost within 1.0 points of the best tested analytics score.

19.7 / 100 · $2.11 / source task
Exact best1.0 pts10 pts

Set 0 for the highest score. Increase the gap to explore cheaper or lighter options.

THE EFFORT CURVE

Does more thinking help?

● Near-best pickOrange = score decreased

025507510015.7Low17.0Medium18.0High17.3Xhigh19.7Max
025507510015.7Low17.0Medium18.0High17.3Xhigh19.7Max

Scores out of 100 · Full 0–100 scale · tap a point to select mode B. Cyan band: within 1.0 points of the best tested score.

WHAT YOU SPEND

Cost of more effort

Measured average USD per benchmark task.

Low$0.7446
Medium$1.1
High$1.61
Extra high$1.89
MaxNear-best pick$2.11

Within this benchmark and harness. Your task mix, tools, caching and endpoint can change the result.

Score difference · B minus A+4.0 pts

Noticeable gap · B scored higher

Measured cost · B / A2.83×

Same benchmark workload

Median duration · B / A4.00×

Measured in the source harness

Gap labels are editorial: under 1 point is small, 1–under 5 noticeable, 5+ large. They do not establish statistical significance. A tied score does not prove equivalent behavior.

WHERE MORE SCORED LOWER

Watch these effort increases.

HighExtra high-0.7 pts18.017.3 on analytics score

Observed reversals within this sweep. They can reflect task fit, sampling, harness behavior or errors; the score alone does not prove “overthinking.”
EVERY TESTED MODE

The numbers behind the choice

5 modes · missing levels are not interpolated

Low

Tested mode
15.7 / 100

0.0 pts vs Low

$0.7446 / source task

98 s median duration

Medium

Tested mode
17.0 / 100

+1.3 pts vs Low

$1.1 / source task

149 s median duration

High

Tested mode
18.0 / 100

+2.3 pts vs Low

$1.61 / source task

246 s median duration

Extra high

Another mode trades better
17.3 / 100

+1.7 pts vs Low

$1.89 / source task

337 s median duration

Max

Near-best pick
19.7 / 100

+4.0 pts vs Low

$2.11 / source task

393 s median duration

Full benchmark table & exact configuration IDs

Swipe inside the table for all measures. Figures from different study selections are not directly comparable.

Claude Opus 5 · Hex DataBench v1.1 · Analytics in the Hex agent harness
EffortAnalytics scoreCost / taskMedian secondsOutput tokensExact source ID
Low15.7$0.744698.2Opus 5 · Low
Medium17.0$1.1149.5Opus 5 · Medium
High18.0$1.61246.5Opus 5 · High
Extra high17.3$1.89337.5Opus 5 · XHigh
Max19.7$2.11392.7Opus 5 · Max
SCALE THE OBSERVED COST

Effort cost difference calculator

What would A versus B cost if you repeated this benchmark’s average workload?

A · Low$744.62
B · Max$2,109.06
Difference$1,364.44more with B

A source-workload scenario using average measured USD costs. Source prices, caching and tools are already reflected in those observed costs. Your agent’s costs may differ.

YOUR WORKLOAD

Reasoning-token cost calculator

Estimate the extra bill from thinking tokens. The starting token counts are editable examples, not measurements of Claude Opus 5. Rates start from our September 23 model snapshot.

A · Low$100estimated monthly cost
B · Max$200estimated monthly cost

Mode B costs $100 more per month with these assumptions.

Standard text-token arithmetic. Excludes retries, tools, caching, long-context premiums, hosting and taxes. Includes thinking in output billing; do not also include it in “visible output.” Check endpoint limits separately. Changing mode does not automatically predict tokens or quality.

READ ACROSS THE EVIDENCE

Claude Opus 5 effort benchmarks at a glance

Each card is its own comparable sweep. Scores from different benchmarks or question sets measure different work and are not averaged together.

Leaderboard updated · 2026-09-22

Analytics · DataBench v1.1

Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.

Low
15.7 · $0.7446 / task
Medium
17.0 · $1.1 / task
High
18.0 · $1.61 / task
Extra high
17.3 · $1.89 / task
Max · pick
19.7 · $2.11 / task

Near-best pick: Max at 1-point tolerance. Score spread: 4.0 points.

Explore this sweep ↗

Original source ↗

PUBLISHED BENCHMARK

Math, code & science · OckBench

Correct answers out of 200 problems. Domain proportions are 100 math, 60 coding and 40 science.

Low
85.0
Medium
85.0
High · pick
94.0
Extra high
95.0
Max
94.5

Near-best pick: High at 1-point tolerance. Score spread: 10.0 points.

Explore this sweep ↗

Original source ↗

Reasoning effort, explained

Which Claude Opus 5 thinking level is most efficient?

Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 4.0 points. This uses a 1-point practical tolerance in the first listed study. Change the evidence, task measure and tolerance above to get a recommendation for another benchmark. It is a shortlist for testing your own workload.

Which Claude Opus 5 effort settings are compared?

Low, Medium, High, Extra high, Max appear in the collected studies. Each chart only includes modes tested together in that study. These published labels do not establish that every endpoint or application exposes every setting.

Does more effort improve Claude Opus 5?

In Hex DataBench v1.1, High to Extra high lowers the observed score by 0.7 points. The overall tested range is 4.0 points. These are observed differences, not proof that extra thinking caused the decline.

How much more does high effort cost for Claude Opus 5?

Use the A/B selector and measured-cost calculator when the source reports dollar costs. For your own agent, enter average input, visible output and thinking tokens in the token calculator. Higher effort is adaptive, so the label itself cannot tell you the bill. Missing costs stay missing.

Keep comparing

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗