A Huawei veteran is betting China’s inference chips on the Agent era

Yuanchuan Micro, an LPU (language-processing-unit) startup, has closed a Pre-A financing round of hundreds of millions of yuan led by IDG Capital, with Futong Capital, Jiukun Ventures, Shangqi Capital, Shunxi Fund and Xuhui Sci-Tech Investment joining. It is the company’s fourth round in 2026. The money goes to LPU chip research, software ecosystem and product engineering, and Yuanchuan says it is already validating with several industrial customers and cloud providers while a next round is under way.

A Huawei veteran is betting China's inference chips on the Agent era
Yuanchuan Micro’s founder argues inference chips should be rebuilt around the Agent era (Source: LeiPhone).

A Huawei lifer’s bet

Founder and CEO Yang Bin spent more than two decades at Huawei, building the company’s processor team from scratch in the US and later leading wireless baseband algorithms and chip work, with long experience in large-scale hardware scheduling and data architecture. He sums up entrepreneurship as ‘the realisation of everything you accumulated’. The core team carries that weight: the CTO has more than thirty years across GPGPU, GPU and DSP architectures, and the chief software engineer has over twenty years in compilers, EDA and heterogeneous systems.

Reasoning backward from the AI endgame

Yuanchuan’s thesis is that the chip race is shifting from GPUs to inference-specific architectures. Yang argues a traditional GPU, built on the von Neumann model, does not fit today’s computation pattern, while full ASIC hardening is too rigid for models that keep evolving. The answer is a native LPU that rebalances what to harden against what to keep flexible. The trigger was the Agent era: by late 2025, reasoning models already handled about half of OpenRouter’s tokens, and agentic inference became the fastest-growing interaction style. An Agent chains planning, tool calls and verification, and any single inference delay compounds along the task chain, then ripples across multi-agent systems.

Yang is explicit that Yuanchuan is not copying Groq. Both chase low-latency, predictable inference, but Groq’s LPU was shaped by the CNN era’s models and needs. Yuanchuan’s question is what inference compute should be reorganised around once AI lives inside agents, on cloud, edge and device at once.

Read the original report (LeiPhone)

Translated and adapted from LeiPhone (leiphone.com).

Leave a comment