Kimi and DeepSeek Showed Up in the Same Model Architecture

One model-architecture diagram featured both Kimi and DeepSeek at once. The GLM-5.3-Flash project config uses 34 layers of Kimi’s KDA linear attention and 11 layers of DeepSeek’s KPool-DSA sparse attention across its 45 layers.

Kimi and DeepSeek Showed Up in the Same Model Architecture
Kimi and DeepSeek Showed Up in the Same Model Architecture (Source: LeiPhone)

This year a clear shift appeared in large-model architectures: in many hybrid attention designs, the role once held by global attention is being taken over by sparse attention. To see why, start with attention itself. Broadly, common mechanisms fall into three: full or global attention reads the complete token history, most faithful but costliest; linear attention compresses history into a compact state to cut long-sequence cost; sparse attention keeps a fuller history but selects only important parts. Much efficient attention evolves along the latter two.

Classic Transformers used full attention, where compute and KV cache grow with context. Hybrid attention became popular as a compromise: efficient attention cuts cost, a little global attention keeps direct access to full history.

MiniMax finally kept full attention

Follow MiniMax’s architecture changes and you see the whole cycle: from betting on linear attention, to keeping a little global attention, back to global, then toward sparse. In late 2023, after talking with founder Yan Junjie, MiniMax bet most of its R&D on a then-unproven linear route. One researcher put his success odds near 99 per cent; Yan put his near half. Early 15-billion-parameter pure-linear tests approached Transformer quality, but scaling exposed a flaw: compressing history loses the ability to recall a precise detail from far back.

MiniMax-01 settled on a 7-to-1 hybrid, seven Lightning Attention layers then one SoftMax, pushing context to 4 million tokens after about 3,700 pretraining experiments. Yet the nagging question remained whether a small set of global layers was enough.

After scaling, hybrids exposed a new problem

At larger scale a second problem appeared: multi-hop reasoning. Mixed attention showed more capability loss than full attention, and the real worry was unknown loss that cannot be exhausted in advance. So M2 returned to full attention, trading higher cost for predictable behaviour.

Agents re-inflated the cost of full attention

By M3 in June 2026, context had become a growing work history, search results, web pages, code, tool returns, errors. Full attention’s cost climbed with every round. M3 adopted MiniMax Sparse Attention, and at 1 million tokens its per-token compute fell to about one twentieth of the previous generation.

Linear and sparse in one model

By 2026 the structure loosened further. MiniCPM-SALA combined Lightning and Sparse; Qwen3.8-Flash-Next kept the three-to-one shape but replaced the backstop attention with Qwen Sparse Attention. GLM-5.3-Flash then put 34 KDA layers and 11 KPool-DSA layers together. Linear maintains a long history cheaply but is weak at precise recall; sparse keeps token history yet reads only a subset. They solve different problems, which is why competing routes ended up in one architecture.

Technology crosses company borders

This recombination keeps happening because code, communities and researchers carry an innovation beyond its origin team. The FLA open-source linear-attention community, maintained since 2024, fed core contributors to Kimi’s KDA. By late 2025 one researcher proposed combining linear and sparse, naming Kimi’s KDA and DeepSeek’s DSA, and within a year the idea appeared in real models.

The labels get shorter-lived. What once read as MiniMax does linear, DeepSeek does sparse, Kimi does KDA, now blurs as technology is verified, absorbed and recombined. The lasting gap is not whether a lab has a technique, but when it adopts it, how it combines it, and what trade-offs it accepts across capability, cost and engineering risk.

Kimi and DeepSeek Showed Up in the Same Model Architecture
Kimi and DeepSeek Showed Up in the Same Model Architecture (Source: LeiPhone)
Kimi and DeepSeek Showed Up in the Same Model Architecture
Kimi and DeepSeek Showed Up in the Same Model Architecture (Source: LeiPhone)
Kimi and DeepSeek Showed Up in the Same Model Architecture
Kimi and DeepSeek Showed Up in the Same Model Architecture (Source: LeiPhone)

Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/yanxishe/S39sSkqntdgkDuJu.html.

Leave a comment