In an AI chip market that refreshes constantly, last generation’s GPU has not fallen in price the way legacy hardware did. ‘The H100 batch we deployed in 2023 now rents for more than it did at launch,’ Stephen Balaban, co-founder and chief technology officer of cloud provider Lambda, said on The MAD Podcast.
The old hope was that GPU compute would standardise and cheapen like traditional cloud, with performance up and price down. AI compute has not followed that path. What one GPU delivers now depends on the system around it, how it coordinates with others, and whether the expensive hardware stays utilised. Chip performance still matters, but it rarely explains the compute a full stack delivers.
Why compute will not commodity
Frontier tasks rarely fit one GPU. Weights, compute and data split across many GPUs and servers, and a training or inference run lives or dies on how fast those GPUs exchange data. Same model, same count, same spec does not mean same compute, because network, memory, storage and software differ. That is why the supernode became a key AI-infra form: Nvidia’s GB300 NVL72 wires 72 GPUs inside a rack over NVLink, and racks extend via InfiniBand or high-speed Ethernet. Lambda designs clusters to minimise network blocking.
The economics are heavy. Balaban estimates a gigawatt AI factory needs about US$2 billion to US$3 billion for power, US$10 billion to US$15 billion for the data-centre shell, and US$35 billion to US$45 billion for servers and compute, with the GPU the largest slice. In a GPU-hour’s cost, depreciation dominates. A roughly six-year accounting life is not the real economic life; an H100 that still takes jobs and finds renters can out-earn the market’s early guess.
Utilisation is the whole game
There are at least three clocks on a GPU: product iteration (is it newest), accounting depreciation (how cost is spread), and economic life (how long it still earns). At 50 per cent utilisation, the depreciation loaded onto each working hour roughly doubles, because an idle GPU keeps depreciating. A 10,000-GPU cluster also needs CPU servers, storage and several networks, and scheduling software decides whether costly physical kit becomes on-demand compute at all.
Neocloud is therefore not GPU rental. It is a vertically integrated service from land and power permits through data-centre build, HPC design, virtualisation and cloud software. The energy-to-token chain loses value at every weak link: power efficiency, GPU-to-GPU communication, model and software optimisation. The firm that keeps its GPUs warm, not the firm that bought the most of them, wins the cost war.



Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/latest/index/id/4764.