Qwen3.8 Max vs GPT-5.6 Sol & Kimi K3: Full Comparison
Explore Qwen3.8 Max, Alibaba's latest flagship AI model featuring a 2.4T-parameter MoE architecture, 1M-token context window, multimodal capabilities, official benchmark results, pricing, API integration, and how it compares with GPT-5.6 Sol, Claude Fable 5, and Kimi K3.
For two weeks, Alibaba's newest flagship had no benchmark table. Just a parameter count and a claim: "second only to Claude Fable 5."
Qwen 3.8 Max was previewed on July 19, 2026, at the World AI Conference in Shanghai as a preview endpoint. No scores shipped with the announcement. Every early writeup ran on vibes and vendor promises. Then August 3 arrived.
Alibaba published a full benchmark table with the official release. The numbers showed genuine strengths across terminal agents, multimodal work, and instruction following, alongside real weaknesses on the hardest software engineering benchmarks. This is what that table actually means for teams deciding whether to switch.
What Qwen3.8 Max Actually Is
Qwen3.8-Max was officially announced on August 3, 2026. The hosted model became available through Qwen Studio and QwenCloud at launch, with open weights announced but not yet published.
It carries 2.4 trillion total parameters in a sparse Mixture-of-Experts architecture, with roughly 95 billion parameters active per token. The active parameter count is what drives serving cost. You pay for 95B compute, not 2.4T.
Qwen says Qwen3.8 Max builds on the Qwen3.5 architecture and scales to 2.4T total parameters. Qwen calls it the first Max-class open-weight model, but the tense matters. The release article says the weights arrive the following week on Hugging Face and ModelScope.
Official Qwen Cloud integration metadata lists a 983,616-token context window and 131,072-token maximum output. Thinking is always enabled, with low, medium, and xhigh reasoning settings.
Qwen3.8 Max architecture overview
Architecture: How a 2.4T Model Runs on Practical Hardware
Qwen3.8 Max uses the same sparse MoE design philosophy that defines the entire 2026 open-weight frontier. Most parameters stay idle. Only a small fraction activates per token.
The sparse Mixture-of-Experts architecture activates roughly 95 billion parameters per token. Active parameters are the number that drives serving cost. A 2.4T total model running at 95B active per token costs roughly the same to serve as a dense 100B model. The knowledge base is 2.4T deep. The compute bill reflects 95B.
The context window sits at approximately one million tokens with up to 128,000 output tokens per call. Thinking is always enabled at launch, with low, medium, and xhigh reasoning effort selectable per request. The xhigh setting is the documented default for the preview.
Multimodal input is confirmed at launch. Text plus visual inputs are confirmed. Coverage differs on the full list; video, documents, speech, and image generation have all been reported. Alibaba has not published a full spec sheet for modality support.
The model supports OpenAI-compatible and Anthropic-compatible API specifications, making migration from existing infrastructure straightforward for most teams.
Model performance
The Benchmark Table: What Is Verified and What Is Not
This is where most coverage of Qwen3.8 Max gets sloppy. Alibaba published scores. Those scores come from Alibaba's own evaluation runs, not independent third-party replication. The distinction matters.
Here is what the official table shows, labeled accurately:
Coding and Agent Benchmarks (Alibaba-run):
Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, 67.7 on SWE-bench Pro, 56.6 on DeepSWE 1.1, and 93.0 on PaperBench. IFBench lands at 82.8.
Additional coding results include 74.8 on CoWorkBench and 81.9 on WideSearch, alongside 92.6 on GPQA Diamond and 92.9 on MRCR v2 at 256K.
Multimodal Benchmarks (Alibaba-run):
Multimodal results include 82.3 on MMMU-Pro, 86.1 on OSWorld-Verified, 92.1 on OmniDocBench 1.5, and 90.4 on Video-MME with subtitles.
| Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max |
|---|---|---|---|---|---|
| Coding Agent | |||||
| Terminal Bench 2.1 | 84.6 | 84.6 | 88.8 | 74.5 | 86.6 |
| SWE-bench Pro | 69.2 | 80.0 | 64.6 | 60.6 | 67.7 |
| DeepSWE 1.1 | 59.0 | 70.0 | 73.0 | 21.6 | 56.6 |
| NL2Repo-Bench | 69.4 | -- | -- | 47.2 | 55.9 |
| FrontierSWE | 70.0 | 88.8 | -- | 40.7 | 73.5 |
| MLS-Bench-Lite | 42.8 | 49.9 | 46.2 | 31.7 | 41.0 |
| PaperBench | 80.3 | 88.8 | 90.5 | 64.8 | 93.0 |
| AndroidBench | 69.8 | 84.5 | 74.0 | 56.5 | 75.1 |
| QwenSWEBench | 84.0 | 86.3 | 73.5 | 63.4 | 80.7 |
| QwenQoderBench | 62.7 | 63.1 | 53.8 | 36.8 | 58.4 |
| QwenReactBench | 1694 | 1770 | 1564 | 1538 | 1724 |
| QwenSVGBench | 1648 | 1690 | 1758 | 1499 | 1713 |
| General Agent | |||||
| CoWorkBench | 72.3 | 75.9 | 71.5 | 64.6 | 74.8 |
| WorkSpaceBench | 66.8 | 68.7 | 65.6 | 61.4 | 67.7 |
| JobBench | 48.4 | 57.4 | 45.4 | 31.3 | 53.4 |
| SkillsBench | 65.1 | 70.9 | 73.5 | 61.2 | 70.2 |
| Agents' Last Exam (Pass / Score) | 27.0 / 45.1 | -- / -- | 30.6 / 53.6 | 11.8 / 31.1 | 27.0 / 52.4 |
| Automation-Bench (Pass@1) | 27.2 | 29.1 | 29.7 | 14.2 | 27.3 |
| Toolathlon Verified (Pass@1) | 76.2 | 77.9 | 74.9 | 49.7 | 72.5 |
| WideSearch | 72.9 | 81.2 | -- | 75.2 | 81.9 |
| HLE w/ tools | 57.9 | 64.5 | 58.0 | 53.5 | 56.2 |
| General Capabilities | |||||
| GPQA Diamond | 92.0 | 92.6 | 94.1 | 92.4 | 92.6 |
| HLE | 45.7 | 53.3 | 47.2 | 41.4 | 43.6 |
| IFBench | 62.2 | 63.5 | 72.7 | 79.1 | 82.8 |
| $OneMillion-Bench (expert score) | 41.8 | 55.9 | 53.8 | 44.4 | 52.5 |
| HealthBench | 52.4 | -- | 55.3 | 54.5 | 60.2 |
| PLawBench | 69.6 | 70.2 | 72.3 | 58.9 | 73.2 |
| PRBench-Legal | 52.7 | 57.6 | 57.6 | 48.5 | 57.6 |
| PRBench-Finance | 51.9 | 55.8 | 55.5 | 46.8 | 58.3 |
| MRCR v2 256K (8-needle) | 83.2 | -- | 93.8 | 86.7 | 92.9 |
| LongBench v2 | 69.1 | -- | 67.1 | 65.3 | 66.3 |
The only independent data point available at launch:
An independent evaluator ran Qwen 3.8-Max preview and Kimi K3 through a production-style software architecture task, analyzing 269 files across two real projects and proposing system designs, blind-reviewed and fact-checked. Result: Kimi K3 scored 83 out of 100, Qwen 3.8-Max preview scored 80. Notable: none of Qwen's 44 tool calls failed, a strong signal for agent reliability.
That is one test, one workload type. But it is a real one, not a preference vote.
Where Qwen3.8 Max Wins and Where It Does Not
Reading the benchmark table carefully reveals a model with specific strengths, not a universal crown.
Strong results:
On Terminal-Bench 2.1, Qwen 3.8-Max scores 86.6 against 84.6 for both Claude Opus 4.8 and Claude Fable 5. GPT-5.6 Sol leads the row at 88.8, so this is a second-place finish. Beating both Anthropic flagships on an agentic terminal benchmark is the headline result Alibaba wanted from this launch.
OSWorld-Verified lands at 86.1, above both Claude Fable 5 at roughly 85.0 and GPT-5.6 Sol at 83.2 on desktop OS interaction tasks.
PaperBench reaches 93.0, the highest reported score on scientific paper analysis, ahead of GPT-5.6 Sol at 90.5 and Claude Fable 5 at 88.8.
On IFBench, Qwen 3.8-Max posts 82.8. The nearest non-Qwen competitor in the table is GPT-5.6 Sol at 72.7, with Fable 5 at 63.5 and Opus 4.8 at 62.2. That is a 10 to 20 point gap, which is unusual on a benchmark this saturated. Instruction following looks like a durable Qwen strength across model generations.
Where it trails:
On SWE-bench Pro, Qwen3.8-Max lands at 67.7 against Fable 5's 80.0 and Opus 4.8's 69.2. It scores 73.5 on FrontierSWE against Fable 5's 88.8.
On the hardest software engineering tasks, Qwen3.8 Max sits mid-pack. It is ahead of GPT-5.6 Sol on SWE-bench Pro at 64.6, but a full 12 points behind Fable 5.
The pattern is clear. Qwen3.8 Max dominates terminal agents, document intelligence, instruction following, and computer use. It is not yet the top choice for pure SWE-level software engineering that requires the deepest code reasoning.
Pricing: Clear and Straightforward
The standard API is $2 per million input tokens, $6 per million output tokens, and $0.25 for cached input. The Token Plan subscription is the alternative for individual use.
This positions Qwen3.8 Max between Sonnet-class and Opus-class pricing from Anthropic. The implicit caching at $0.25 per million input tokens makes long-context agentic work meaningfully cheaper for workloads where the same large prompt gets reused across many calls.
Pricing comparison
How to Use Qwen3.8 Max via API
Qwen3.8-Max is OpenAI-spec compatible. Set the base URL and drop it into any existing client:
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ["QWEN_API_KEY"],
base_url="https://api.qwencloud.ai/v1",
)
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": "Analyze this codebase and identify the three highest-risk components."
}
],
extra_body={"reasoning_effort": "xhigh"},
)
print(response.choices[0].message.content)
With vision input:
import base64
from pathlib import Path
image_data = base64.b64encode(Path("screenshot.png").read_bytes()).decode()
response = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image_data}"}
},
{
"type": "text",
"text": "Identify the UI bugs visible in this screenshot."
}
]
}
],
)
print(response.choices[0].message.content)
Via Anthropic-compatible endpoint in Claude Code:
{
"model": "qwen3.8-max",
"base_url": "https://api.qwencloud.ai/anthropic",
"api_key": "YOUR_QWEN_API_KEY"
}
The official API guide lists low, medium, and xhigh reasoning effort. The official guide also provides an OpenAI Responses-compatible configuration for Codex and an Anthropic-compatible endpoint for Claude Code, plus Qwen Code, Qoder CLI, and OpenClaw integrations.
Qwen3.8 Max vs Kimi K3: The Direct Comparison
Both models launched within weeks of each other. Both target the same frontier agent use case. Here is what the evidence actually shows:
| Metric | Qwen3.8 Max | Kimi K3 |
|---|---|---|
| Total Parameters | 2.4T | 2.8T |
| Active Parameters | 95B | 32B |
| Context Window | 1M tokens | 256K tokens |
| Terminal-Bench 2.1 | 86.6% | 88.3% |
| SWE-bench Pro | 67.7% | 63.2% |
| OSWorld-Verified | 86.1% | 81.2% |
| API Pricing (Input) | $2.00/M | $3.00/M |
| API Pricing (Output) | $6.00/M | $15.00/M |
Kimi K3 activates far fewer parameters per token at 32B versus 95B, making it faster and cheaper per output token. Qwen3.8 Max leads on OSWorld and SWE-bench Pro. Kimi K3 leads on Terminal-Bench 2.1. Neither wins cleanly across all categories.
The sharpest summary of the difference: K3 shipped a full independent benchmark table at launch. Qwen3.8 Max shipped vendor scores. That can change with one announcement, but it matters for teams making production decisions today.
Who Should Use Qwen3.8 Max
The model's benchmark profile maps clearly to specific workloads:
Use Qwen3.8 Max for computer use agents, document parsing at scale, instruction-following pipelines where consistency across many tool calls is critical, and long-context research automation where PaperBench-class performance matters.
Alibaba's primary demonstration involves autonomous software development extending beyond ten days. While enterprises should treat these demonstrations as vendor claims until independently reproduced, they align with a growing interest in persistent coding agents that operate continuously rather than interactively.
Use a different model for the hardest SWE-level coding tasks. Qwen3.8-Max results support a narrow conclusion: it is competitive across coding, research, tool use, long context, documents, computer use, and video. They do not support a universal crown.
Conclusion
Qwen3.8 Max is the clearest signal yet that Alibaba is serious about the frontier. 2.4 trillion parameters, 95 billion active per token, a one-million-token context window, native multimodal input, and API pricing below competing flagships.
The benchmark story is honest: strong on terminal agents, computer use, instruction following, and document intelligence. Not yet the top choice for the hardest software engineering benchmarks. The generational improvement over Qwen3.7-Max is real, with DeepSWE 1.1 jumping from 21.6 to 56.6 and FrontierSWE from 40.7 to 73.5.
The one independent data point available, the 80/100 on a real-world architecture task versus Kimi K3's 83, and zero failed tool calls across 44 attempts, suggests a model that is reliable under agent load. That is the detail that matters more than any benchmark row for teams building production agents.
The open weights are coming. When they land on Hugging Face, the calculus changes for every team running self-hosted inference. Until then, the API is live, the pricing is clear, and the benchmarks are available to test against your own workloads.
Test it. Compare it. The era of trusting one lab's benchmark table without verification is over, and Qwen3.8 Max is honest enough in its results to reward that scrutiny.
FAQs
Q1. What makes Qwen3.8 Max different from other large language models?
Qwen3.8 Max uses a sparse Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters but activates only about 95 billion parameters per token. This delivers frontier-level performance while keeping inference costs closer to those of dense 100B-class models.
Q2. Is Qwen3.8 Max suitable for coding and AI agents?
Yes. Qwen3.8 Max performs strongly on terminal agents, computer-use tasks, document intelligence, and instruction following. While it is highly competitive for coding, some specialized software engineering benchmarks are still led by competing frontier models.
Q3. Does Qwen3.8 Max support multimodal inputs and long-context reasoning?
Yes. Qwen3.8 Max supports multimodal inputs including text and images, offers a context window of nearly one million tokens, and provides configurable reasoning levels for complex AI workflows and long-document analysis.
Simplify Your Data Annotation Workflow With Proven Strategies