Fei-Fei Li and Yilun Du’s Rare Joint Paper: Stop Building Robots a Brain, Borrow the Video Model’s

One of the most star-studded “fusion” papers of the year in video generation and embodied intelligence has just appeared. A few days ago, a paper titled Masked Visual Actions for Unified World Modeling landed on arXiv. It has not yet spread through the academic circuit, but it is about to blow up. The author list … Read more