Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

ALIBABA · OPEN WEIGHTS

Qwen3.8 Flash Next

Benchmarks, pricing and tradeoffs for agent builders.
Tested configuration: source default.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Add to a custom comparison ↗
Agent Fit76.2General agent profile
Practical Value70% capability · 30% cost
Estimated monthly API cost1,000 calls · 10K input + 2K output
Catalog context windowNot verifiedAdvertised capacity

Capability profile

Explain these tests ↗
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.87.4 / 100
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.72.6 / 100
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.61.6 / 100
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.85.8 / 100
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.74.2 / 100
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.74.6 / 100
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.77.1 / 100

What the evidence says

Qwen3.8 Flash Next scores 76.2 on the equal-weight LiveBench category average. Its highest measured category is reasoning (87.4); its lowest is agentic coding (61.6). Categories differ in difficulty, so these are test results, not universal strengths or weaknesses.

Our General agent profile produces an Agent Fit score of 76.2 by emphasizing reasoning, instruction following, agentic coding, coding and data analysis. See the exact weights or adjust the profile in the interactive comparison.

Independent preference evidence

A matching Arena configuration is not verified in this snapshot. Evidence Consensus is withheld. A result for another reasoning setting or version is not substituted.

API features to verify

The catalog has no verified tool-support observation. These are declared features, not evaluations of tool accuracy or visual understanding.

Weights are publicly available for this model family. Review the exact model license and serving requirements before deployment. Hosted API costs below are not self-hosting estimates.

Qwen3.8 Flash Next API pricing

We have not verified an exact catalog version match, so pricing and workload cost are withheld. Do not substitute a similarly named release when budgeting.

Trace this model to the source

LiveBench row ID
qwen3.8-flash-next
Question-set release
2026-06-25 · Source CSV ↗
Catalog ID
Exact version not verified
Arena configuration
No verified matching observation
Reviewed snapshot
2026-09-23

Compare alternatives

For production decisions, test representative tasks with a fixed harness and success criteria. Neither a large context window nor a strong benchmark score guarantees reliability in your own agent.

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗