Fei-Fei Li and Yilun Du’s Rare Joint Paper: Stop Building Robots a Brain, Borrow the Video Model’s

One of the most star-studded “fusion” papers of the year in video generation and embodied intelligence has just appeared. A few days ago, a paper titled Masked Visual Actions for Unified World Modeling landed on arXiv. It has not yet spread through the academic circuit, but it is about to blow up. The author list … Read more

Fei-Fei Li and Yilun Du just co-signed a paper that skips the robot brain

The missing translation interface The paper’s core point is that today’s video models understand physics but cannot issue commands. A model fed the start of a kettle tilting will continue the scene plausibly, water pours into the cup, yet it has never done anything, and faced with a robotic arm told to move a cup … Read more