Low
Tested mode0.0 pts vs Low
Compare Low vs Medium vs High where tested. Find the effort setting that fits your task, and see how much extra reasoning changes the result.
AI Agent Store task-weighted category mix. An editorial selection aid, not a measured agent success rate; weights match our published Agent Fit methodology.
2025-12-11 version · Question set: 2026-01-08 · Reviewed September 23, 2026Check the source ↗Lowest tested effort within 1.0 points of the best tested general agent fit.
70.1 / 100Set 0 for the highest score. Increase the gap to explore cheaper or lighter options.
● Near-best pickOrange = score decreased
Scores out of 100 · Full 0–100 scale · tap a point to select mode B. Cyan band: within 1.0 points of the best tested score.
The recommendation uses the lowest tested effort within your score tolerance. That does not prove it is the cheapest or fastest.
Use the token calculator below with your own usage measurements. Effort labels alone do not tell us a token bill.
Estimate your own token costs ↓Large gap · B scored higher
Best minus lowest tested score
High
Gap labels are editorial: under 1 point is small, 1–under 5 noticeable, 5+ large. They do not establish statistical significance. A tied score does not prove equivalent behavior.
0.0 pts vs Low
+7.4 pts vs Low
+9.8 pts vs Low
Swipe inside the table for all measures. Figures from different study selections are not directly comparable.
| Effort | Overall benchmark | General agent fit | Coding agent fit | Research agent fit | Writing & support fit | Reasoning | Coding | Agentic Coding | Mathematics | Data Analysis | Language | Instruction following | Cost / task | Median seconds | Output tokens | Exact source ID |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Low | 65.3 | 60.3 | 60.8 | 62.2 | 63.3 | 71.2 | 73.9 | 50.0 | 84.7 | 52.7 | 70.2 | 54.5 | — | — | — | gpt-5.2-2025-12-11-low |
| Medium | 71.8 | 67.7 | 63.3 | 73.3 | 68.5 | 84.2 | 72.1 | 51.7 | 92.1 | 70.4 | 74.9 | 57.5 | — | — | — | gpt-5.2-2025-12-11-medium |
| High | 74.8 | 70.1 | 64.7 | 76.9 | 72.2 | 83.2 | 76.1 | 51.7 | 93.2 | 78.2 | 79.8 | 61.8 | — | — | — | gpt-5.2-2025-12-11-high |
What would A versus B cost if you repeated this benchmark’s average workload?
This source has no comparable dollar-cost observation. Use the token calculator below with your own usage; token counts alone cannot supply a full bill.
Estimate the extra bill from thinking tokens. The starting token counts are editable examples, not measurements of GPT-5.2. Rates start from our September 23 model snapshot.
Mode B costs $56 more per month with these assumptions.
Standard text-token arithmetic. Excludes retries, tools, caching, long-context premiums, hosting and taxes. Includes thinking in output billing; do not also include it in “visible output.” Check endpoint limits separately. Changing mode does not automatically predict tokens or quality.
Each card is its own comparable sweep. Scores from different benchmarks or question sets measure different work and are not averaged together.
AI Agent Store task-weighted category mix. An editorial selection aid, not a measured agent success rate; weights match our published Agent Fit methodology.
Near-best pick: High at 1-point tolerance. Score spread: 9.8 points.
Explore this sweep ↗High is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 9.8 points. This uses a 1-point practical tolerance in the first listed study. Change the evidence, task measure and tolerance above to get a recommendation for another benchmark. It is a shortlist for testing your own workload.
Low, Medium, High appear in the collected studies. Each chart only includes modes tested together in that study. These published labels do not establish that every endpoint or application exposes every setting.
In the first listed LiveBench sweep, the tested score range is 9.8 points and no adjacent effort increase lowers the score. This does not guarantee a gain on every task or in another harness.
Use the A/B selector and measured-cost calculator when the source reports dollar costs. For your own agent, enter average input, visible output and thinking tokens in the token calculator. Higher effort is adaptive, so the label itself cannot tell you the bill. Missing costs stay missing.