Google ships robot models, Ant LingBo already works a pharmacy

On 30 July Google DeepMind released Gemini Robotics 2, a family spanning a vision-language-action model, a reasoning model and an on-device model, and for the first time delivering whole-body control of a humanoid. Google’s play is simple to state: own the robot brain layer. In the phone era it sold few handsets but reached most smart devices through Android. In the robot era it wants to replay that platform logic.

Then look at Ant LingBo, and the contrast is sharp. It is building a universal embodied-intelligence brain for different brands and body types, in plain terms the robot equivalent of Android. Between 7 and 10 July it released six results, LingBot-Vision, LingBot-Depth 2.0, LingBot-VLA 2.0, LingBot-Video, LingBot-World 2.0 and LingBot-VA 2.0, together a full-stack brain 2.0.

Humanoid robot operating in a simulated retail or pharmacy setting
Ant LingBo ran three different-brand robots on one shared brain inside a working pharmacy at WAIC.

Already on the shop floor

A week later that brain appeared in front of a pharmacy shelf at WAIC. Three robots from Leju, Xinghai Tu and LingBo’s own R-2 moved between shelves fetching medicine, all running the same LingBot-VLA 2.0. The author ordered a cold-medicine item at random; one Leju robot completed recognition, search, grasp and handover. Different bodies, one brain, the clearest picture of one-brain-many-bodies.

The old pattern was one body, one model: switch robots and re-collect data, switch tasks and retrain, enter a new scene and re-debug. The more robots sold, the larger the engineering team, and scale economies vanish into integration cost. LingBo folds different robots’ actions into a 55-dimension representation; LingBo-VLA 2.0 already adapts 17 brands and more than 20 body types across single-arm, dual-arm, biped and wheeled forms.

Robot navigating shelves in a pharmacy layout
One brain, many bodies: LingBo adapted its model to 17 brands and more than 20 robot types.

Why embodied-native matters

LingBo repeats one phrase, embodied-native, as its core direction. Recognising a cup is one ability; actually picking it up is another. A vision model can name the cup, but the robot must judge distance, approach angle, force and what happens after release. Intelligence in the physical world is a closed loop of perception, action and feedback learned on one causal chain from day one.

So LingBo built not a lone VLA but a stack: vision for boundaries, depth for true distance, VLA for intent to motion, a world-action model for constant correction, and video and world models for simulation and trial. Chief scientist Shen Yujun argues VLA and world models may not be the end state; future models grown specifically for the physical world will likely appear.

Open source to win the plug-in

To become the robot Android, step one is getting others to plug in. LingBo open-sourced its vision, VLA, video and world models, releasing weights, training code and technical reports for developers and body makers via Hugging Face and ModelScope. An efficient post-training version runs inference under 130 milliseconds on an RTX 4090, which brings the real deployment time and compute bill into range.

Leju showed the pull: its KUAVO 4 Pro reached LingBot-VLA readiness with just 150 demonstration samples and passed 95 real-robot task tests. A body maker can keep building the body and motion control, hand understanding and planning to the shared brain, and enter a new scene without restarting from the base model.

Hard gates remain. Can 20-plus body types become shorter integration cycles and lower delivery cost. Can the pharmacy move from showcase to volume with failure rate, human takeover and upkeep under control. And will the developers and partners drawn in by open source become paying customers. The model is promising, but the platform test is still ahead.

Editor’s note: translated and adapted for RobotBelt from OFweek Robotics. Read the original report here.

Leave a comment