In August 2026, Dyna Robotics released Dyna-2, its new robot foundation model. Pretraining used over 1 million hours of first-person human video, roughly 170 years of continuous awake experience. Across seven benchmarks, the WAM architecture reached 1.55 times the success rate of the previous VLA models. High-precision manufacturing task success jumped from Dyna-1’s 20 per cent to 80 to 90 per cent. Scene deployment pass rate rose from 46 per cent to 87 per cent. Dyna-2 lets robots learn work by watching human video: 1 million hours of training, success from 20 per cent to 90 per cent.
Behind those numbers sits a detail most missed: Dyna once walked the VLA road too, then deliberately abandoned it.
Dyna-1 used VLA, vision-language-action, generating the next action directly from image and language command. It was the mainstream brain of embodied AI for two years. Figure’s Helix, AgiBot’s GO-2 and Galaxy’s G0.5 all walked it. But Dyna found VLA’s structural flaw in real deployment and switched to WAM, world-action model. This is not another model release. It is a story of abandoning the mainstream route for a better one.
VLA’s wall: what Dyna hit in the real world
Dyna-1, the world’s first dexterous-manipulation foundation model deployable in commerce, was not technically backward. It shipped in restaurants, laundries and gyms. But those deployments exposed VLA’s structural ceiling.
One: the lab-to-site success cliff. In the lab, Dyna-1 and Dyna-2 both near 100 per cent pass. At the customer site, an environment the robot never saw, Dyna-1’s pass rate plunged to 46 per cent. VLA shines in controlled settings and collapses in real-world complexity.
Two: fragile post-training data. Small post-training sets make Dyna-1 easily disturbed. VLA depends heavily on fine-tuning data, and real scenes never have enough high-quality data.
Three: no understanding of physical consequence. VLA solves what the robot should do next. It knows what, not what happens after. In real commerce this produces frequent did-but-wrong: action executed, physical result off target.
Dyna’s lesson from napkin-folding and other real tasks: for high-success, high-robustness scenarios, WAM clearly wins. Not theory. Combat experience.
What WAM adds: a layer of physical imagination
Dyna-2’s world-action model jointly predicts the next video frame and the next motor command. VLA is see, think, act. WAM is see, think, predict consequence, act: before generating motion, the robot simulates in its brain what the world becomes after this action.
One model, two outputs: Dyna-2 emits both future video, the predicted physical state after the action, and future action, the concrete motor command. The robot does not merely execute commands. It understands the causal law of the physical world.
Co-founder and chief scientist Jason Ma put it plainly: truly scarce is action data, but video is everywhere. Physical intuition need not be trained on robot arms for millions of hours. It can be learned directly from human video. Dyna-2 uses zero robot operation data in pretraining, only 1 million hours of first-person human video, then a tiny amount of robot demonstration in deployment fine-tuning.
How far ahead is WAM over VLA?
Under identical training and evaluation, WAM success is 1.55 times VLA. WAM quality score is 1.12 times VLA. Counting both pre and post training, Dyna-2 wins 65 per cent of the time, Dyna-1 only 29 per cent. Most of VLA’s wins came early in pretraining. WAM’s advantage widened as scale grew. The 1.55 times was fought out in real commercial deployment, not run in a lab.
What Dyna-2 really validated: three Scaling Laws
Its core contribution is using 1 million hours to validate three things.
First, the world-action model shows a Scaling Law, unsaturated even at 1 million hours. Dyna built nested subsets of 1,000, 10,000, 100,000 and 1 million hours. Every metric improved monotonically and fit a power law. Accuracy at 0.1 rose 51 per cent across the ladder. This is the first Scaling Law verified to 1 million hours on real-world manipulation data.
Second, the human-to-robot cross-embodiment Scaling Law is proven to exist. Zero robot data in pretraining, yet across 39 cross-embodiment tasks, offline action-prediction error fell monotonically with data, with a clear inflection between 10k and 100k hours. Human data scale directly predicts robot performance on new tasks.
Third, world modelling is necessary for cross-embodiment transfer. Action-only routes fail. Co-training future video frames is required for generalisation to keep growing with data. Even fixing 50,000 hours of action data and only expanding pure video, zero-shot robot prediction error fell sharply.
Dyna’s conclusion: video is itself an independent scaling axis. When data shifts from scarce teleoperation to infinite human video, robot foundation models gain a predictable expansion path for the first time.
What WAM means for the industry
Dyna-2 points to three deep shifts.
One, VLA and WAM are moving from replacement to fusion. Jason Ma said at WAIC 2026 the two capabilities are not fully in conflict and may merge. The consensus: the physical-AI base model is far from architecturally settled. VLA has not exited. WAM is accelerating, and the two are converging.
Two, data strategy is being rewritten. For two years the field leaned on teleoperation at 500 to 1,000 yuan an hour, costly and limited. Dyna-2 proved a different path: pretrain on human video, fine-tune with little robot data. One million hours of human video plus 13 minutes of robot demo taught a robot to unscrew a bottle cap. That ratio, 1 million hours versus 13 minutes, reveals the core direction. Dyna’s next plan expands human-video learning from 1 million to 10 million hours.
Three, Scaling Law finally has proof in robotics. LLMs have it as consensus. Robots did not, until now. Dyna-2 is the first verified to 1 million hours. Embodied brains can now improve by stacking data, like language models.
Dyna Robotics was founded September 2024, headquartered in Silicon Valley with a hardware R&D centre in Shanghai, registered as Dana Lingdong. The three co-founders, CEO Lindon Gao, York Yang and chief scientist Jason Ma, are all ethnic Chinese. Gao and Yang previously co-founded smart-cart company Caper AI, sold to a strategic buyer for US$350 million in 2021. Jason Ma earned his doctorate at Penn’s GRASP lab and worked at Google DeepMind, NVIDIA and Meta AI.
In September 2025 the company closed a US$120 million Series A with NVIDIA, Samsung, LG and Amazon, valuation above US$600 million. A one-year-old, Chinese-founded company won backing from three giants spanning chips, consumer electronics and e-commerce.
Dyna’s switch from VLA to WAM was not theory. It was survival wisdom forced out by real commercial scenes. Dyna-2 validated human video as a new scaling axis for physical AI, proved a robot Scaling Law, and confirmed the human-to-robot transfer law. VLA’s boundary is now mapped. WAM is the new direction. The next question: who reaches 10 million hours first.
Images






Editor’s note: This is an adapted translation of the original OFweek report. It has been trimmed and restructured for readability for an international business audience.