NVIDIA’s B300 has become the hardest chip to buy in China, and the price is moving faster than anyone can plan around. Spot quotes for a single machine have climbed from roughly 4 million yuan, about 570,000 dollars, in the weeks after Lunar New Year to more than 12 million yuan, around 1.7 million dollars, per unit. The market moved from quoting a new price every week to quoting one every day.
Behind the shortage sits a quieter risk. Several deals this year involved sellers who claimed to have stock, then asked buyers for capital verification, company credentials and procurement details. The real aim, according to people close to the transactions, was to harvest the financial, purchasing and compute-deployment data of major Chinese enterprises, a clear information-security hazard.
The fight for real supply is intensifying. One major East China manufacturer is pushing a B300 purchase that could exceed 10,000 units, out-buying a previously aggressive internet giant. To protect its own orders, the company has asked some suppliers to prioritise it and to stop shipping to that rival at the same time.

Domestic inference chips are being pulled apart
While the B300 stays scarce, Chinese chipmakers are accelerating their own routes. One company that has long bet on the AFN architecture has taped out a chip aimed at FFN and mixture-of-experts workloads, while its Attention chip is still in design. Until that matures, it plans a heterogeneous inference setup pairing its chip with a GPU.
Another inference-chip startup chose to enter through the Prefill stage first. Its debut part targets performance above NVIDIA’s H100 and is scheduled for late next year. The first generation uses LPDDR memory and will later move to 3D stacking.
Big buyers are bypassing the middlemen
Compute relay stations emerged so domestic developers could reach overseas models such as Claude and GPT through an intermediate layer, usually cheaper than the official API. For large enterprises the saving is not worth the cost.
The objections are concrete. Business customers will not let data pass through a third party, so they trust only the original provider. Relay stations cannot meet the high concurrency that big firms demand on tokens per minute and requests per minute. Their billing is a black box, unauditable line by line, while enterprises require every charge to be reconcilable. And account bans are a live danger. ByteDance and Tencent have both been burned by reverse overuse of subscription accounts triggering vendor bans.
3D stacking becomes the new startup frontier
With inference demand rising and advanced-process capability constrained at home, 3D stacked chips are the fresh hotspot. Several startups have drawn capital fast. One saw its valuation climb from a few hundred million yuan to several billion yuan in six months.
Capital enthusiasm is not the same as industrial maturity. 3D chips carry real challenges across stacking process, yield, cost and supply chain. Beyond the funding race, what actually decides a company’s value is whether it can close the mass-production loop.
A storage giant chases the missing piece
A leading domestic memory maker tried to take control of a rival IP company’s RCD chip business at a valuation of tens of billions of yuan, to fill a gap in DDR5 server memory interface chips. RCD is the core interface chip in DDR5 RDIMM server memory, and only three suppliers dominate globally. The deal collapsed over valuation. The memory maker then turned to a western Chinese chip company to advance its RCD plans.
An AI infrastructure star loses its leaders
A North China AI infrastructure company known for its “AI factory” strategy has lost three core executives, including a chief technology officer who joined with fanfare from an international tech giant last year. Backed by local state-owned capital, the firm ran several regional compute-infrastructure projects and was seen as a key explorer of the domestic AI-factory model. As the smart-compute sector shifts from card-counting to operations, the departure of its core team raises an obvious question about direction.
Investors want tapeout before cheques
The new inference architectures are heating up, but investors are unsure where to place bets. Several say plainly they will not consider a startup that has not taped out. By contrast, the view on GPUs has converged. Many believe few new top-tier GPU companies can still be born, and that training will consolidate to one or two players. Inference is a larger market, but it too will not hold many winners.
The internet playbook no longer pays
The classic internet logic, acquire users first, grow scale, then count daily actives, does not transfer to the AI era. The difference is marginal cost. A new internet user adds almost no incremental cost. A new AI user consumes tokens, every single one a direct compute expense.
That is why the “give away 10 million free tokens” subsidy cannot work in business terms. Every token handed out is a real cash cost, and no platform can absorb it indefinitely. The free-to-grab-users-then-monetise model is dead. What survives is a model where inference cost per user is paid for by the value the user generates, not deferred to a mythical future.
Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience.