Alibaba’s Qwen 3.8-Max leans into long-horizon agents, betting execution beats raw scale

The author had Qwen 3.8-Max build a 3D large-model evolution universe website from scratch and iterated against GPT-5.6.

Qwen 3.8-Max building a 3D website in a live demo
Alibaba’s Qwen 3.8-Max was tested against GPT-5.6 on a 3D website build. (Source: LeiFengWang)

On ceiling, Qwen 3.8-Max showed far more delivered detail than GPT-5.6. On efficiency and cost, Max’s high completion came with low efficiency, and it ran out of quota mid-task because it could not judge the endpoint. For fast iteration and low-cost prototyping, GPT stays the first choice. For max-performance-without-cost-ceiling, Qwen 3.8-Max is an expensive but effective domestic alternative.

After 2.4 trillion parameters, why does Max stress agents

On 3 August, Alibaba released Qwen 3.8-Max. Specs still dazzle: 2.4 trillion total parameters, about 95 billion active, 1 million token context, across text, code and vision, with Max weights and a 27B small model planned. Benchmarks, scale, price and open source follow Qwen’s route.

But at this generation, those alone no longer explain differentiation. Against Kimi K3’s hybrid linear attention and Stable LatentMoE architecture narrative, Qwen 3.8-Max did not publish an equally weighty new base architecture. Its price stays competitive versus overseas flagships but not low enough to alone change model choice. Even if the 2.4T weights open, their main value leans to cloud vendors, enterprises and research, not local dev deploy. When top models already share high base ability, bigger, stronger, more open gets harder to claim as a flagship’s sole reason.

Qwen 3.8-Max technical blog chart on reinforcement learning environments
Qwen 3.8-Max leans on real-environment reinforcement learning to lift agent ability. (Source: LeiFengWang)

Qwen 3.8-Max gave more space to another direction: long-horizon agents. The official blog rushes past specs and benchmarks to dwell on 16-day autonomous development, 125 hours of research reproduction, hundreds of chip-optimisation rounds, and end-to-end execution projects in coding and office scenarios. They stress one ability: can the model call tools, take feedback, check results and keep editing in a real environment, pushing a complex task to completion.

That gives Max a clearer position. It need not carry popularity and local deploy like small models, but raises Qwen’s ceiling, then feeds that ability into workflows via Qwen Code, office agents and Alibaba Cloud. Whether it delivers is what the author tested.

Where does agent ability come from? Max puts training weight on real-environment RL

Qwen 3.8-Max did not define this generation with a new base architecture. The official report says it is built on the Qwen 3.5 architecture foundation, and explains the agent lift mainly through real-environment RL, reinforcement learning.

Real-environment RL does not mean drilling on a bigger question bank. It places the model in near-real work: read the project, call tools, run code, watch the page, take test results, then revise from feedback. The training goal is no longer a text closer to a standard answer, but a full work loop: plan, act, check, find problems, decide next.

That is the key difference between agent ability and ordinary generation. One generation gives an answer. An agent must read current state, choose tools and act on external feedback.

More from the original report

Qwen 3.8-Max parameter and context specification slide
Qwen 3.8-Max: 2.4 trillion total parameters, 95 billion active, 1 million token context. (Source: LeiFengWang)
Qwen 3.8-Max agent execution screenshot
Max emphasises end-to-end execution in coding and office scenarios. (Source: LeiFengWang)
Qwen 3.8-Max RL training curve
As RL training environments grow, composite scores rise from a 0.474 SFT baseline. (Source: LeiFengWang)
Qwen 3.8-Max online data balancing chart
Online data balancing keeps the model from over-fitting a few scenarios. (Source: LeiFengWang)
Qwen 3.8-Max training data distribution
Training data is dynamically balanced across task types and difficulty. (Source: LeiFengWang)
Qwen 3.8-Max 3D website build result
The author required a real in-browser 3D scene, not static mockups. (Source: LeiFengWang)

Editor’s note: This is an adapted translation of the original LeiFengWang report. It has been trimmed and restructured for readability for an international business audience.

Translated and adapted from LeiFengWang (https://www.leiphone.com/category/yanxishe/hFaRe5oZ7oF0U1XC.html).

Leave a comment