Qwen3.8 Max (0902)

On September 2, Alibaba released Qwen3.8-Max (0902), an updated version of its flagship model. This is not a brand-new model generation. The 0902 release is a post-training update to Qwen3.8-Max, with most of the work focused on coding, agent workflows, and long-running tasks.
Advertisement 728 × 90
CompanyAlibaba
Context1M
Released2026-09
Updated2026-09-08

Qwen3.8 Max (0902) Overview

Qwen3.8-Max (0902) is the latest snapshot of Alibaba’s Qwen3.8-Max flagship model.

Its model ID is qwen3.8-max-0902, also listed as qwen3.8-max-2026-09-02.

The core specifications remain largely the same:

Up to 1 million tokens of context
Thinking and non-thinking modes
Text, image, and video input
Function calling
Structured outputs
Web search
Context caching

The biggest changes are in software engineering, agent-style work, and visual understanding.

You can think of 0902 less as a new engine and more as a major tune-up of the existing one.

The base model already had long context, multimodal input, reasoning, and tool use. This update is about making those features work better together on real jobs.

Qwen3.8 Max (0902) Pricing

PlanPriceDescription
API Input $1.65 / 1M tokens Charged by input usage
API Output $4.951 / 1M tokens Charged by generated output
Context Cache Discounted pricing Useful when reusing large prompts or long context

Qwen3.8 Max (0902) Key Features

1. Coding Is the Biggest Upgrade

Coding is where the 0902 update changes the most.

Alibaba added more post-training aimed at software engineering, especially tasks that take several steps or require the model to keep working over time.

On CodeArena’s frontend coding leaderboard, Qwen3.8-Max (0902) gained 22 points, reaching 1691 and taking the top spot.

Other reported benchmark gains include:

  • TerminalBench 3.0: 11.3 → 29.0
  • DeepSWE 1.1: 56.6 → 69.3
  • NL2Repo-Bench: 55.9 → 64.9
  • ProgramBench: 10.5 → 28.0
  • SWE-Marathon: 39.1 → 44.8
  • QwenSWE-Bench V2: 55.1 → 70.0

The jumps on TerminalBench and ProgramBench are particularly interesting.

Those tests are closer to real engineering work: using a terminal, dealing with unfamiliar environments, and figuring out how to reproduce or fix something instead of just writing one function.

That is a better fit for what 0902 is trying to be.

2. Better at Agent and Cowork Tasks

The second big area is Cowork.

Here, “cowork” means more than chatting with a user. The model may need to move between tools, documents, and several steps before a task is done.

Reported results include:

  • CoWorkBench: 74.8 → 76.1
  • JobBench: 53.4 → 64.0
  • WorkArena Elo: 1348 → 1468

The JobBench gain stands out.

These are the kinds of tasks where a model needs to read information, decide what to do next, use tools, and keep track of progress.

The first step is not the hard part.

The hard part is still being useful on step twelve.

That is where 0902 looks stronger than the original release.

3. Visual Understanding Also Improved

The update is not just about code.

Qwen3.8-Max (0902) also improves chart reasoning, document understanding, and multimodal perception.

Reported results include:

  • MMMU-Pro: 82.3 → 82.7
  • ERQA: 77.8 → 78.3
  • ClawEval-MM: 77.2 → 80.2
  • BabyVision + CI: 91.3 → 93.8

These gains are smaller than the coding improvements, but they matter for agent work.

If a model is going to operate software or handle office workflows, it needs to understand screenshots, charts, documents, and interfaces — not just text prompts.

That makes the visual upgrades more practical than they may look from the benchmark numbers alone.

4. The Main Step Forward Is Execution

Compared with the original Qwen3.8-Max, the 0902 version does not change the 1-million-token context window or introduce a new architecture.

The bigger change comes from post-training.

The original model already had long context, Thinking mode, multimodal input, and tool use.

0902 pushes those capabilities further into real work.

The clearest improvements are in:

  • Large codebase tasks
  • Terminal-based work
  • End-to-end software development
  • Long-running agent tasks
  • Multi-tool workflows

This is why coding and Cowork matter more than the headline model specs.

The update is less about making Qwen3.8-Max “smarter” in the abstract and more about making it better at finishing complicated work.

Summary

Strengths

  • Much stronger coding performance: Several software engineering and agent benchmarks improved sharply.
  • Better long-running task execution: More useful for work that takes many steps instead of one response.
  • Stronger Cowork workflows: Better at moving between tools, documents, and tasks.
  • 1-million-token context remains useful: A good fit for large codebases, long documents, and complex projects.
  • No major API price increase for the 0902 update: It remains in the same pricing tier as the original Qwen3.8-Max on Model Studio.

Limitations

  • It is still an update, not a full model generation change.
  • Not every area improved by the same amount: Coding gains are much larger than most vision gains.
  • Long jobs can still burn through tokens quickly: Especially with Thinking mode and very large context windows.
  • Independent testing is still limited: A good share of the early numbers come from Alibaba or reports based on its published tests.

Best For

  • Developers working on large software projects
  • Teams using coding agents or terminal automation
  • Companies dealing with long documents or large codebases
  • Users who want a Chinese flagship model for agent and workflow automation

Probably Not Worth It For

  • Casual chat and basic writing
  • Simple tasks that do not need long context
  • Workloads where very low latency matters most
  • Users who do not need agents, coding, or tool use

Comments (0)

Leave a comment

Advertisement 728 × 90

Alibaba Model Comparison

Model Context Pricing API Released Global Heat
Qwen3.8 Max (0902)
1M YES 2026-09
1M YES 2026-08
87/100
262K YES 2026-08
90/100
262K YES 2026-08
60/100
1M YES 2026-08
83/100

Similar Models

Qwen3-235B
92
Alibaba MoE model achieving high-quality reasoning at low cost with 22B active parameters.
Alibaba
Qwen3.8 2.4T A95B
90
What stands out to me isn’t how well it answers a single question—it’s whether it can keep a complex task going. In my own experience, hitting a wall usually isn’t about missing the basics. It’s about finding something that can pick up where you left off and move forward.
Alibaba
Qwen3.8 Max
87
Alibaba released Qwen3.8-Max on August 3. It has 2.4 trillion total parameters, activates around 95 billion parameters per inference step, and supports a 1 million token context window. The pitch is not just “better coding.” The model is supposed to handle the full process: break down requirements, write code, debug, iterate, and eventually deliver a working project.
Alibaba
Qwen3.7 Max
85
China's strongest AI of 2026, deep reasoning + autonomous execution, from conversation to getting things done.
Alibaba
Qwen3.8 Flash
83
On August 26, 2026, Alibaba’s Qwen team released Qwen3.8-Flash and open-sourced the model weights at the same time. It’s a multimodal MoE model built around one main idea: strong enough for real work, but cheap enough to use at scale.
Alibaba
Qwen3.7 Plus
82
Text + image input, text output. Upgraded vision-language capabilities, with full agentic strength in coding and tool use retained.
Alibaba
Qwen3.7 Flash
81
Alibaba’s Qwen3.7-Flash is a high-performance lightweight AI model optimized for multimodal understanding, AI Agents, coding, and fast, cost-efficient reasoning.
Alibaba
Qwen3.6 Flash
80
Alibaba’s Qwen3.6 Flash has been getting a fair amount of attention among developers. After using it for a few days, my impression is pretty straightforward: it is not trying to beat the Plus model on raw capability.
Alibaba

Related Tools

Related News