In 2026 the large-model industry is living through a visceral shift. The Scaling Law that ruled the digital world is hitting the hard wall of physical law, spatial constraint and long-tail data as it crosses into the real world. The rift peaked on the final day of RSS 2026 in Sydney, where robotics academics’ reverence for physical causality met the AI industry’s faith in scale. On that fault line, Nvidia physical-AI expert Max Li unpacked Cosmos 3.

This is no prettier video generator. Cosmos 3 is the first world foundation model that, in a single Transformer, fully unifies text, vision, audio and action. If the past two years of embodied AI were the “assembled-PC era” of bolting vision, language and control models together, Cosmos 3 announces the “integrated-appliance era.” Li signalled that the second half of embodied AI leaves the data centre for the edge: by moving the AGI brain onto Jetson-class devices for real-time run, Nvidia ends the era of robots carrying servers on their backs.
Behind it is Nvidia’s ultimate ambition as compute ruler: not just to monopolise hardware, but to become the “Android” of embodied AI, defining a standardised brain-and-body protocol so fragmented robot forms can evolve on one logic base. Li described three core abilities of a world foundation model: understanding the world, simulating outcomes, and taking actions. Cosmos 3 ingests all modalities, from text and images to video, audio and action, and can generate any of them.
It also performs inverse dynamics, inferring the underlying action from multi-view observations, a powerful mining tool for the vast video corpora without paired action labels. Once the modelling of physical outcomes is deep enough, the model converts directly into a policy that acts in the real world. Cosmos Edge, a 4-billion-parameter MoE, runs real-time inference on Jetson without server-grade GPUs. The family spans Cosmos Nano at 16B parameters and Cosmos Super at 64B.
Architecturally, Cosmos 3 splits into a Reasoner (a VLM that understands the world) and a Generator (cross-modal generation) in a mixture-of-experts design, using causal attention for understanding and bidirectional attention for generation. It unifies action vectors across embodiments, from cars and single-arm robots to bipeds and humanoids, and aligns all modalities on a 24-frames-per-second time axis. Training used about 22 million pretrain samples for the reasoner and roughly 2.2 million samples for the generator, including about 10 million hours of video.
Crucially, Cosmos 3 is fully open source. Li said he replies to every GitHub issue himself, without any AI agent, and that more than 100 people contributed. The bet is that a standardised, open physical-AI base lets the whole robotics community build on one foundation instead of reinventing fragmented stacks.
Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience.