The author had Qwen 3.8-Max build a 3D large-model evolution universe website from scratch and iterated against GPT-5.6.

On ceiling, Qwen 3.8-Max showed far more delivered detail than GPT-5.6. On efficiency and cost, Max’s high completion came with low efficiency, and it ran out of quota mid-task because it could not judge the endpoint. For fast iteration and low-cost prototyping, GPT stays the first choice. For max-performance-without-cost-ceiling, Qwen 3.8-Max is an expensive but effective domestic alternative.
After 2.4 trillion parameters, why does Max stress agents
On 3 August, Alibaba released Qwen 3.8-Max. Specs still dazzle: 2.4 trillion total parameters, about 95 billion active, 1 million token context, across text, code and vision, with Max weights and a 27B small model planned. Benchmarks, scale, price and open source follow Qwen’s route.
But at this generation, those alone no longer explain differentiation. Against Kimi K3’s hybrid linear attention and Stable LatentMoE architecture narrative, Qwen 3.8-Max did not publish an equally weighty new base architecture. Its price stays competitive versus overseas flagships but not low enough to alone change model choice. Even if the 2.4T weights open, their main value leans to cloud vendors, enterprises and research, not local dev deploy. When top models already share high base ability, bigger, stronger, more open gets harder to claim as a flagship’s sole reason.

Qwen 3.8-Max gave more space to another direction: long-horizon agents. The official blog rushes past specs and benchmarks to dwell on 16-day autonomous development, 125 hours of research reproduction, hundreds of chip-optimisation rounds, and end-to-end execution projects in coding and office scenarios. They stress one ability: can the model call tools, take feedback, check results and keep editing in a real environment, pushing a complex task to completion.
That gives Max a clearer position. It need not carry popularity and local deploy like small models, but raises Qwen’s ceiling, then feeds that ability into workflows via Qwen Code, office agents and Alibaba Cloud. Whether it delivers is what the author tested.
Where does agent ability come from? Max puts training weight on real-environment RL
Qwen 3.8-Max did not define this generation with a new base architecture. The official report says it is built on the Qwen 3.5 architecture foundation, and explains the agent lift mainly through real-environment RL, reinforcement learning.
Real-environment RL does not mean drilling on a bigger question bank. It places the model in near-real work: read the project, call tools, run code, watch the page, take test results, then revise from feedback. The training goal is no longer a text closer to a standard answer, but a full work loop: plan, act, check, find problems, decide next.
That is the key difference between agent ability and ordinary generation. One generation gives an answer. An agent must read current state, choose tools and act on external feedback.
More from the original report






Editor’s note: This is an adapted translation of the original LeiFengWang report. It has been trimmed and restructured for readability for an international business audience.
Translated and adapted from LeiFengWang (https://www.leiphone.com/category/yanxishe/hFaRe5oZ7oF0U1XC.html).