Embodied data enters the ‘out-of-box’ era: why robot training needs model-ready data

At WAIC 2026, one line landed hard: embodied intelligence’s data demand ‘has not come close to its ceiling.’ The bottleneck is not compute, it is data, and the industry has not even filled the pre-training gap.

Embodied data enters the 'out-of-box' era: why robot training needs model-ready data
At WAIC 2026, embodied-AI firms argued the real bottleneck is model-ready data, not compute (Source: LeiPhone)

What ‘model-ready’ means

Jianzhi Robotics calls the data it sells ‘Model-Ready Data’, data you can feed straight to a model without rework. It needs three traits at once: high precision, multimodality and diversity. Hand-tracking error of 2-3cm, the industry norm, means the model sees a finger a full digit off, so it grabs through cup walls. Without tactile data it never learns how hard to grip an egg.

The hard parts

Precision means reconstructing human motion at under 1cm error with sub-20ms sensor sync. Multimodality needs hardware-level time alignment across cameras, gloves and IMUs, but most teams bolt modules together and get 20-50ms drift. Diversity means sampling across households, roles and objects, not 900,000 hours of the same ‘pick up a cup’ loop in a white lab, which trains a robot that only works in a demo.

Scale is the final exam

Jianzhi runs an end-to-end stack: value-driven chain-of-thought labelling that records not just the action but its goal and logic, a Data Foundation Model for occlusion completion, sub-1cm joint tracking, and six-camera rigs capturing full-body motion including waist and legs. Its DPH metric, data-hours per skill across people and scenes, quantifies diversity so buyers know the data is not repetitive. The takeaway: robot training is becoming an industrial data pipeline, not a research chore.

Read the original report (LeiPhone)

Translated and adapted from LeiPhone (leiphone.com).

Leave a comment