Peking University’s Zongqing Lu Bet His Start-up on the Latent Space, Not the Demo Reel

In June 2025 at the Beijing Academy of Artificial Intelligence conference, Zongqing Lu offered a claim that ran against the crowd: human video is the only data path that can scale up for embodied intelligence. While the industry fixated on robot hardware, simulation farms and data flywheels, Lu was already accumulating first-person human video, now past 500,000 hours.

Peking University's Zongqing Lu Bet His Start-up on the Late
Peking University’s Zongqing Lu Bet His Start-up on the Late (Source: LeiPhone)

Lu is a tenured associate professor at Peking University, a long-serving senior area chair at ICML, NeurIPS and ICLR, and the earliest systematic proponent of the latent-space route in China. In 2025 he founded BeingBeyond, pursuing a latent world-action model: instead of generating imagery, it predicts in the embedding space what the robot should do next and how the world will respond.

This route costs about 1 per cent of the video-generation approach to train and is fast enough for real-time control, but it has one glaring weakness. There is no picture to show. In a funding climate obsessed with demos, that takes conviction.

On 28 July, BeingBeyond released Being-H0.8, the first implicit tactile world-action model, bringing the touch modality into large-scale pretraining and unifying vision, touch, action and future-state change in one latent space.

Lu is blunt that the deterministic technology for an embodied foundation model has not yet appeared. GPT could be stacked because Transformer and next-token prediction came first, he argues. Embodied AI is not there yet. We do not even know what the deterministic technology is.

He is sceptical of the video-generation route. Its training cost is enormous, large models may need tens of thousands of accelerators and hundreds of millions of yuan, and Alibaba’s Wan team reportedly stalled at version 2.2. Worse, robots need real-time decisions, and by the time a frame renders, the moment has passed.

The latent route has two measurable edges: training cost near 1 per cent of the pixel route at equal data and parameters, and a better chance of learning physical causality because the supervision is not glued to pixels.

Lu also calls the popular data flywheel a false proposition for most companies. A start-up that only collects data from its own body overfits to one form factor and can never scale. The 500,000 hours matter because every hour was screened for diversity, not just stacked for length.

His real differentiator is pretraining, not post-training. A strong pretrained model can cut the real-machine data a downstream task needs by an order of magnitude, something post-training cannot buy. Being-H0.8 already adapts across bipeds, arms, grippers and dexterous hands.

He expects two to three years before the industry sees genuine commercial deployment, and a true general home robot only when a model can handle diverse tasks without heavy scene adaptation. The H0.8 label, he says, marks a node on the way to a finished H1.0.

Peking University's Zongqing Lu Bet His Start-up on the Late
Peking University’s Zongqing Lu Bet His Start-up on the Late (Source: LeiPhone)
Peking University's Zongqing Lu Bet His Start-up on the Late
Peking University’s Zongqing Lu Bet His Start-up on the Late (Source: LeiPhone)
Peking University's Zongqing Lu Bet His Start-up on the Late
Peking University’s Zongqing Lu Bet His Start-up on the Late (Source: LeiPhone)

Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/ai/cTpkd2xI7DZBdCRk.html.

Leave a comment