XPeng bolts a ‘time axis’ onto physical AI, rewriting how self-driving reads the world

A self-driving system can see a car, a pedestrian, a red light. What it struggles with is time. XPeng says its second-generation VLA model fixes exactly that, trading a static 3D view of the world for a moving 4D one.

XPeng vehicle perceiving the road environment
XPeng’s assisted-driving system reading a complex road scene. (Source: CheDongXi)

From 3D space to 4D spacetime

At a recent briefing, Liu Xianming, head of XPeng’s general intelligence centre, argued the industry over-indexed on spatial understanding: knowing where things are now. Physical AI, he said, also has to know how a situation developed and what comes next. The upgrade, shipped as XOS 6.3.0, builds “time” into the model through three capabilities: remember the past, predict the future, respond faster.

The Infini-VLA long-sequence architecture remembers the previous 30 seconds as one continuous event chain rather than disconnected frames. The X-Foresight world model predicts the next six seconds of road-user behaviour from position, speed, motion trend and road structure, with internal tests reaching up to 21 seconds. Streaming inference lets the model perceive, think and act at once, lifting end-to-end response by 300 per cent.

Together the three modules close a spatiotemporal decision loop: Infini-VLA owns the past, streaming inference the present, X-Foresight the future.

Why it matters for the driving stack

The harder problem is the compute bill. Holding 30 seconds of history and a six-second forecast on top of road structure, traffic participants, right-of-way and navigation intent pushes against the physical limit of in-car chips. XPeng’s answer is a tighter software and hardware co-design rather than bigger models for their own sake.

For European observers, the signal is that China’s smart-driving teams are competing on temporal reasoning, not just sensor counts. The car that understands time is the car that handles like a human, and that is where the next leg of the autonomy race is being run.

Image gallery

XPeng second-generation VLA architecture diagram
The three-module spacetime decision loop. (Source: CheDongXi)
XPeng car front sensor view
Front-facing perception of an XPeng vehicle. (Source: CheDongXi)
XPeng VLA remembering 30 seconds diagram
Infini-VLA holds the prior 30 seconds as one event chain. (Source: CheDongXi)
XPeng predicting future six seconds
X-Foresight forecasts up to 21 seconds of behaviour. (Source: CheDongXi)
XPeng streaming inference diagram
Streaming inference cuts end-to-end latency by 300 per cent. (Source: CheDongXi)
XPeng 3D spatial understanding graphic
From 3D space toward 4D spacetime. (Source: CheDongXi)
XPeng vehicle on test road
XPeng test vehicle in real traffic. (Source: CheDongXi)

Editor’s note: This is an adapted translation of the original CheDongXi report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://chedongxi.com/p/375182.html.

Translated and adapted from CheDongXi (https://chedongxi.com/p/375182.html).

Leave a comment