Dyna-2 abandons the VLA mainstream for a world-action model and lifts robot success from 20 to 90 per cent

In August 2026 Dyna Robotics released Dyna-2, a robot foundation model pretrained on over 1 million hours of first-person human video, about 170 years of continuous waking experience. Across seven benchmarks the WAM architecture hit 1.55 times the success rate of the prior VLA model. High-precision manufacturing success rose from Dyna-1’s 20 per cent to 80 to 90 per cent. Scene-deployment pass rate rose from 46 to 87 per cent. The detail most miss: Dyna once walked the VLA road and abandoned it on purpose.

Dyna-1 used VLA, vision-language-action, the most popular robot brain of the past two years, used by Figure’s Helix, AgiBot’s GO-2 and Galbot’s G0.5. But real deployment exposed VLA’s structural ceiling, so Dyna switched to WAM, world-action model.

VLA’s wall in the field

Dyna-1 was the first commercial dexterous base model, deployed in restaurants, laundries and gyms. Those very deployments exposed the ceiling. Problem one: the lab-to-field cliff. In the lab both Dyna-1 and Dyna-2 near 100 per cent, but in a customer site neither had seen, Dyna-1 fell to 46 per cent; VLA generalises poorly off the controlled set. Problem two: fragile post-training data. Small fine-tune sets leave VLA easily disrupted. Problem three: no sense of physical consequence. VLA answers the next move, knows what but not what happens after, so robots often “did it but wrong”.

Dyna’s lesson from napkin folding and real sites: for high-success, high-robustness tasks, WAM wins clearly. Not theory, field truth.

Chart of Dyna-2 benchmark success rate versus prior VLA model
Dyna-2’s WAM tops the prior VLA model by 1.55 times on shared benchmarks. (Source: OFweek Robot)

WAM adds a layer of physical imagination

WAM jointly predicts the next video frame and the next motor command. VLA is see, think, act. WAM is see, think, predict consequence, act: before moving, the robot simulates in its own brain what the world becomes. One model, two outputs, future video and future action, so it understands physical cause, not just instructions.

Co-founder and chief scientist Jason Ma put it plainly: action data is the scarce part, but video is everywhere. Physical intuition need not be earned over millions of hours on a manipulator, it can be learned from human video. Dyna-2 uses no robot data in pretrain, only 1 million hours of human first-person video, then fine-tunes with a sliver of robot demos.

Head to head, same training and protocol: WAM success is 1.55 times VLA, quality 1.12 times, and in combined pre and post training Dyna-2 wins 65 per cent versus Dyna-1’s 29 per cent. Most VLA wins came early in pretrain; WAM’s edge grows with scale. The 1.55 times was not run in a lab, it was fought in real sites.

Scaling-law chart showing WAM success rising with pretraining hours
Dyna-2 validates a world-action scaling law still rising at 1 million hours. (Source: OFweek Robot)

Three scaling laws

Dyna-2’s core gift is not a new model but proof, with 1 million hours, of three things. One, world-action models show a scaling law not yet saturated at 1 million hours. Two, video data scales the same way robot data was meant to. Three, physical-intuition pretrain plus tiny robot fine-tune closes the field gap VLA could not.

Editor’s note: This is an adapted translation of the original OFweek Robot report. It has been trimmed and restructured for readability for an international business audience.

Leave a comment