Arm adds AI to the GPU and leaves the NPU to partners in its mobile compute blueprint

What compute base does personal AI actually need? As phones start to understand intent, call tools and complete tasks continuously, the platform must support a full personal-AI experience, not one model inference. At Arm Everywhere China 2026, Arm released its latest CSS for Mobile 2, where AI plays two roles in the new phone compute subsystem: the CPU gains AI muscle to spread AI across more apps, and the Mali GPU integrates a neural-network accelerator natively for the first time, using AI to lift gaming.

Arm’s CPU is pushing on-device AI further. “Almost all flagship smartphones, Android or iOS, now use SME2,” said James McNiven, Arm’s vice president of edge-AI intelligent terminal compute. SME2 is the AI matrix-acceleration capability Arm added to the CPU. The new Arm C2 Ultra brings a 15 per cent single-thread gain, and the C2 CPU cluster with enhanced SME2 reaches up to 1.7 times the performance of the previous generation on specific AI models.

Arm CSS for Mobile 2 compute subsystem diagram showing CPU, GPU and partner NPU
Arm’s CSS for Mobile 2 keeps the NPU to partners while strengthening CPU and GPU AI. (Source: Leiphone)

The new Arm Mali GPU stands out. In the “Light Reborn” demo, Mali G2-Ultra NX with neural graphics reached up to four times the frame rate and energy efficiency of the previous-generation Mali G1-Ultra without AI-assisted rendering. AI begins to join picture reconstruction, letting phones gain gaming performance in a new way.

CPU accelerates AI adoption, AI accelerates GPU evolution. But why did this personal-AI platform not include Arm’s own NPU? McNiven explained: “In mobile, Arm’s consistent approach is to leave NPU-layer innovation to partners.” By leaving NPU differentiation to partners, Arm focuses on CPU, GPU and the software ecosystem, from KleidiAI to neural-graphics models, SDKs and game-engine partnerships, shortening the distance between processor capability and developer apps. A hardware-plus-software system and ecosystem is Arm’s key to staying a leader in the AI era.

SME2 becomes a flagship standard

The compute load of personal AI is far more complex than one inference. Booking an anniversary dinner, an agent does speech-to-text, searches favourite restaurants, calendar and photos, then generates instructions, browses the web and acts. if the restaurant is full, the whole flow reruns. The CPU runs light models, tool calls and task orchestration, and latency at each step shapes the experience.

C2 Ultra targets such loads with better single-thread, execution-engine and data-access performance: 15 per cent gains in single-thread and web browsing, 12 per cent in app launch and multi-thread. C2 Ultra takes burst loads. C2 Pro takes sustained efficiency. Partners tune the mix rather than cover every scene with one core.

The bigger AI gain comes from SME2: the C2 cluster with enhanced SME2 reaches 1.7 times on specific models. But on-device AI is also memory-bandwidth bound. Quantisation from INT8 to INT4 to INT2 cuts compute and data but can lose precision. quantization-aware training mitigates it. Arm favours low latency over peak: the C2 CPU cluster’s effective compute is about 5 to 6 TOPS. Private L2 and shared L3 caches cut DRAM accesses and the power they cost.

Mali GPU integrates an AI accelerator for the first time

If the CPU’s evolution lets more apps use AI, Mali G2-Ultra NX lets AI join graphics compute and change how the GPU outputs frames. Real-time ray tracing, complex geometry and dynamic global illumination are reaching phones while thermal, battery and bandwidth lag. With per-pixel rendering, developers traded quality, frame rate and power endlessly. Mali G2-Ultra NX’s native neural-network accelerator ends that trade-off.

Mali G2-Ultra NX neural graphics reconstruction demo showing AI-rebuilt frames
Mali G2-Ultra NX rebuilds most pixels with neural graphics rather than traditional rendering. (Source: Leiphone)

The GPU uses neural super sampling to render low resolution and let a neural network rebuild high resolution, neural super sampling plus denoising for ray-tracing noise, and neural frame-rate uplift to reconstruct intermediate frames from real ones. In the “Light Reborn” three-frame demo, only one quarter of pixels in frames one and three used traditional rendering. the middle frame was fully AI-rebuilt, so seven eighths of pixels were AI-reconstructed. This can add latency, so neural frame-rate uplift needs Android frame pacing and the engine. Mali G2-Ultra NX reaches up to four times the frame rate and energy of the previous generation and cuts DRAM traffic by up to 70 per cent. Without neural graphics, average game performance rose 20 per cent and same-frequency up to 14 per cent, the gap showing AI’s headroom.

Traditional graphics also improved: a new execution engine and third-generation ray-tracing unit cut repeated triangle data and DRAM bandwidth needs by about 13 per cent in ray-tracing benchmarks. “AI-enhanced graphics will become an industry standard,” McNiven said.

NPU to partners, system and software for personal AI

CSS for Mobile 2 has no Arm NPU, but personal AI still needs one. The CPU suits orchestration, tool calls and low-latency light loads. the GPU does graphics and general compute. the NPU suits sustained compute-intensive models. McNiven noted Arm keeps NPU innovation with partners and has the Ethos NPU for IoT. Phone-chip makers run their own NPU architectures optimised for their models, OS and first-party AI. Third-party apps must cover many brands and chips, so per-NPU adaptation is costly. CPU and GPU’s standardised programming fit a common base, with the system vendor scheduling the rest.

Arm leaves configuration space: the C2 cluster supports up to 14 cores mixing Ultra and Pro. Mali G2-Ultra NX allows configurable shader cores and NX units. Arm supplies system-optimised CPU, GPU and interconnect IP plus the CSS software system, and this year put the software stack and ecosystem first. KleidiAI connects SME2-optimised kernels to AI frameworks so developers call hardware without touching low-level instructions.

Arm Everywhere China 2026 repeated the software theme so new compute reaches apps, not just spec sheets. Moving from IP to CSS, Arm combines CPU, GPU and system IP to shorten chip development, while software connects up to models, frameworks and apps. As personal AI keeps shifting, Arm offers a configurable compute base and software that lowers development and migration cost.

Editor’s note: This is an adapted translation of the original Leiphone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/chipdesign/E3ZbypZiS817Q2G8.html.

Translated and adapted from Leiphone (https://www.leiphone.com/category/chipdesign/E3ZbypZiS817Q2G8.html).

Leave a comment