Why Spatial Intelligence Became AI’s Route Into the Physical World

AI is moving from generating a picture that looks real to understanding space, predicting change and executing action. At ECCV 2026 that trend was impossible to miss.

Why Spatial Intelligence Became AI's Route Into the Physical
Why Spatial Intelligence Became AI’s Route Into the Physical (Source: LeiPhone)

The programme chair reported 10,473 submissions with 2,834 accepted. By topic, image and video synthesis led with over 1,400 papers, vision-language-reasoning was second, and 3D fell to third. But 3D’s lower rank is not a retreat. Problems long centred on 3D, geometry, physics consistency, dynamic modelling and interactive generation are spreading into video, robotics, embodied intelligence and world models.

Of 93 workshops, about 10 touched 3D directly, 11 covered spatial intelligence, 9 pointed at physical AI and 5 took world models as their theme. Awards reinforced the signal: the Koenderink Test-of-Time award went to Fei-Fei Li’s team, while best-paper honours returned to 3D surface representation and geometric perception.

3D vision: from static geometry to dynamic 4D

The best paper, Heat Kernel Textures, from Imperial College, re-asks how to represent a 3D surface’s texture, using an anisotropic heat kernel in geodesic space to avoid the seams and stretching of traditional UV mapping. Two honourable mentions landed on 3D reconstruction and geometric perception, including Meta’s LSRM for high-fidelity object-centric reconstruction and Stony Brook’s Poppy for polarization-based surface-normal estimation that plugs into existing RGB models without retraining.

Spatiotemporal memory: let the model keep knowing one world

Qunhe Technology with Zhejiang University proposed WalkerBench, an interactive spatial-intelligence benchmark where an agent navigates long ranges using only first-person RGB, stripping GPS and street names. Mainstream vision-language models got lost on long trajectories; the team’s Spatial-IDE, with an explicit topological memory and goal-directed perception, lifted nine agents’ average score by 104.97 per cent and was deployed on a Unitree G1 for kilometre-scale street navigation. Tencent Hunyuan, Tsinghua and NTU proposed Spatial-TTT, turning part of the model into fast weights updated during inference to hold long-video spatial memory.

State prediction: from representing change to predicting it

Yann LeCun, in a keynote, argued a model should not predict pixels but future state in abstract representation, then plan by the outcomes of actions; a robot needs a goal state, not a target image. Jamie Shotton of Wayve traced embodied AI from Kinect to autonomous driving as its testbed. Wu Jiajun, Tan Ping, Angela Dai, Li Yunzhu, Gao Yang and Liang Xiaodan pushed world models, 3D reconstruction and physical consistency across workshops.

DreamWorld feeds 3D geometric priors into video generation for cross-view stability; Interaction-Aware 4D Gaussian Splatting handles hand-object dynamics; a physics-grounded benchmark asks whether multi-agent world-model changes obey real physics.

Embodied interaction: from spatial understanding to executable action

Kristen Grauman urged AI to move from understanding to enabling, becoming an AI guide that teaches real-world skills. NVIDIA’s Baiyi Li shared MultiGen, which adds sound to simulation video so robots judge physical state by audio, and DexImit, which turns monocular human video into bimanual robot data. RoboTracer, from Beihang, Peking, BAAI and CAS, converts language into a 3D trajectory with depth and distance, robot-agnostic. FEEL synced first-person video with about 3 million frames of force signals; SPEAR turns Unreal Engine into a programmable embodied simulator.

Toward Physical AI

Vincent Sitzmann showed video models can hold long scene consistency without a 3D map; Qixing Huang argued explicit 3D still defines multi-view consistency and reveals what a model learned. The route is unsettled, yet spatial intelligence as the base capability for Physical AI grows only clearer.

Why Spatial Intelligence Became AI's Route Into the Physical
Why Spatial Intelligence Became AI’s Route Into the Physical (Source: LeiPhone)
Why Spatial Intelligence Became AI's Route Into the Physical
Why Spatial Intelligence Became AI’s Route Into the Physical (Source: LeiPhone)
Why Spatial Intelligence Became AI's Route Into the Physical
Why Spatial Intelligence Became AI’s Route Into the Physical (Source: LeiPhone)

Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/academic/hxpVB01n8UMq6NDS.html.

Leave a comment