Huawei veteran’s LPU startup YuanChuan Micro closes a new nine-figure round

YuanChuan Micro LPU inference chip (Source: LeiPhone)
YuanChuan Micro’s LPU inference chip. Source: LeiPhone

Before Groq’s LPU became famous on the back of Nvidia’s spotlight, Yang Bin, founder and CEO of YuanChuan Micro, had already been sketching an inference chip of his own.

LeiPhone has learned exclusively that LPU startup YuanChuan Micro has closed a Pre-A round of several hundred million yuan. The round was led by IDG Capital, with Futong Capital, JiuKun Ventures, Shangqi Capital, Shunxi Fund and Xuhui Sci-Tech Investment joining. It is the company’s fourth round in 2026 alone.

YuanChuan Micro said the proceeds will go to LPU-plus chip R&D, software ecosystem building and product engineering. It is already running validation with several industrial customers and cloud providers, and its next round is already underway.

Public records show YuanChuan Micro was founded in September 2025. Its angel round this April, also several hundred million yuan, drew top VCs, government funds and supply-chain players.

From Huawei to a bet on the Agent era

The continued capital inflow reflects a new round of chip competition in the inference age, spreading from GPU to inference-specific architectures. On chip form alone, YuanChuan Micro looks like a “China Groq”: a spatiotemporal compiler, a hard-pipeline architecture, and high-bandwidth memory.

Yang disagrees with the simple comparison. “If our goal were only to imitate an LPU, the room to expand would be small,” he says. For him, the chip technology is the result of a market judgement, not the starting point.

That thinking is tied to his more than twenty years at Huawei, where he built Huawei’s processor team from scratch in the US, then led wireless baseband algorithms and chip work in China, owning large-scale hardware scheduling and data architecture.

“Starting a company is the cashing-in of everything you accumulated,” Yang sums up. The team’s base is decades of technical skill, industry experience and talent. CTO Dr Sun has over thirty years in GPGPU, GPU and DSP architectures. Chief software engineer Dr Will has more than twenty years in compilers, EDA and heterogeneous systems.

Why inference, and why now

Rather than debate “domestic substitution” or “cost performance,” this team of industry veterans asks what inference compute should be reorganised around as AI enters the Agent age.

In late 2025, as Yang and his CTO discussed opportunities in the AI chip track, they realised the value emerges after a model actually lands, not from its capability. DeepSeek R1 trained near-top-tier performance at far below expected cost. Manus opened the imagination of general agents. Jensen Huang made inference and agents the core narrative of GTC.

A 2025 study on a hundred-trillion-token dataset showed inference models already handling about half of OpenRouter’s tokens, with agentic inference the fastest-growing interaction. Yang notes an agent often runs dozens or hundreds of inferences per task, and any latency compounds down the chain, like a butterfly effect across the system.

So YuanChuan Micro settled on an LPU route: not full ASIC, not generic GPU, but a rethought balance of what to harden and what to keep flexible. The wager is on the fastest inference, not the cheapest.

Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/latest/index/id/4757.

Leave a comment