A 22-year Huawei veteran bets on LPU chips as Yuanchuan Micro closes a nine-figure round

Long before Groq’s LPU made headlines on the back of Nvidia’s backing, Yang Bin had already been sketching out an inference chip of his own. Lei Feng Network has learned that LPU startup Yuanchuan Micro has closed a new nine-figure yuan Pre-A round. The round was led by IDG Capital, with Futong Capital, Ninequant, Shangqi Capital, Shunxi Fund and Xuhui Science-and-Technology Investment joining.

Yuanchuan Micro LPU-class inference chip
Yuanchuan Micro’s LPU-class inference silicon. Source: Lei Feng Network

It is the company’s fourth financing round since the start of 2026. Yuanchuan Micro says the proceeds will go mainly into LPU-plus chip research, software ecosystem building and product engineering. The firm is already running product validation with several industrial customers and cloud providers.

Public records show Yuanchuan Micro was founded in September 2025. Its angel series earlier this year, also a nine-figure yuan round, drew top venture capital, government funds and supply-chain players, covering everything from wafer fabrication to end applications.

The steady capital inflow reflects a new round of chip competition in the inference era, one that is stretching from GPUs into purpose-built inference architectures. Backed this way, Yuanchuan Micro, with its spacetime compiler, hard-pipeline design and high-bandwidth memory, does look like a “Chinese Groq”.

Yang disagrees with the simple comparison. “If our only goal was to imitate an LPU, the room to expand would actually be small,” he says. In his view, the chip technology chosen is merely the natural result of a market judgement, not the starting point of the venture.

That way of thinking is tied to his more than two decades at Huawei. He built the processor team from scratch in the United States, then returned to lead wireless baseband algorithm and chip work, overseeing large-scale hardware scheduling and data architecture. “Starting a company is the cashing in of every resource you accumulated,” Yang sums up.

The core team’s decades of accumulated skill, industry experience and talent form the base of the startup. The CTO, Dr Sun, has studied GPGPU, GPU and DSP architectures for over thirty years. The lead software engineer, Dr Will, brings more than twenty years in compilers, EDA and heterogeneous systems.

Rather than debate “domestic substitution” or “cost performance”, this team of industry veterans asks a sharper question: as AI enters the Agent era, what should inference computing be reorganised around?

In the first half of 2025, Yang and his CTO already reasoned that a model’s true value only appears once it is deployed. DeepSeek R1 had trained near-top performance at a fraction of expected cost. Manus had just opened the imagination of general agents. Jensen Huang made inference and Agents the core narrative of GTC, with a full solution for the inference era.

As inference demand released, Yang turned to productivity apps such as vibe coding. “It is like home bandwidth going from 100 megabits to 1 gigabit. The charge does not rise linearly,” he says, familiar with the commercial logic from years in the industry. Customers will pay a premium for efficiency.

“A value-for-money pitch definitely has a market, but it gets compressed fast,” he says. So the first job is not to save the customer money, but to create higher value. “Do only inference, and do only the fastest inference” became Yang’s rule.

A 2025 study of one hundred trillion tokens showed inference models already handled about half of OpenRouter’s traffic, and agentic inference was the fastest-growing interaction. That amplifies the worth of low latency and high stability. Yang points out that an Agent often runs planning, tool calls and verification, which can mean dozens or hundreds of inferences. Any single delay compounds along the chain.

If several agents coordinate, one agent’s slow response propagates like a butterfly effect and drags the whole task. Yang also notes the human preference for stable rhythm in conversation: when an Agent outputs tokens at uneven pace, the user feels something mechanical.

With the Agent-era need clarified, rethinking inference chip design became the natural next step.

Editor’s note: This is an adapted translation of the original Lei Feng Network report. It has been trimmed and restructured for readability for an international business audience.

Leave a comment