Xspark AI’s Ding Wenbo: touch cannot replace vision, but robots need their own ‘spinal cord’

Xspark AI’s Ding Wenbo: touch cannot replace vision, but robots need their own ‘spinal cord’

Embodied intelligence has entered its “tactile era,” and Ding Wenbo proposes a “spinal cord” architecture to fix the physical-interaction gap in robots. Ding is hard to pin down with one word: communications PhD, materials postdoc, machine-tactility scholar. In 2025 he added another title, co-founder and chief scientist of Xspark AI (WuJie ZhiHang). A Tsinghua professor who spent years on sensors did not build the company as a “sensor company.” Better sensors, he argues, do not automatically make a smarter robot. Robot touch may need its own model.

Robotic hand performing a dexterous manipulation task with tactile sensing
Dexterous manipulation is where tactile feedback earns its place. (Source: Leiphone)

The case starts with the limits of pure vision. The dominant embodied path still follows the scaling logic of large models, more video, more robot trajectories, bigger VLAs, lifting generalisation along a see-understand-act chain. Ding does not oppose it, calling VLA “a pretty solid logic” that will leave its mark. But a robot ultimately reaches out and touches the world. Vision already solves perhaps 99 per cent of perception, so touch need not compete with it. It fills in what happens after contact: is the cup starting to slip, is the grip too tight, did the arm hit someone. Once a robot enters the physical world, the question is not only whether it saw, but whether it can sense, correct and avoid danger in time.

The direct move, stuffing touch into the existing model as a new modality, does not satisfy him. By first principles, touch carries far less information than vision yet demands far higher feedback efficiency. Fuse it straight into the vision model and you get either redundancy or a few critical signals drowned out. Hence his call for “tactile-native intelligence.” It need not match a vision model’s semantic richness. It should be lighter, faster, closer to the body, like the human spinal reflex that acts before the brain identifies a fall.

Diagram of layered robot intelligence from slow brain to spinal reflex
Xspark’s three-layer architecture: slow brain, fast brain and spinal cord. (Source: Leiphone)

The hard parts remain. Hardware consistency across tactile sensors is poor, even within one batch. Touch lacks the internet’s inherited stock of images and video, so data shifts when the person, sensor or robot changes. The deeper puzzle is how to represent and embed touch, whether force, temperature, shear and vibration can converge into a hardware-agnostic common representation. Ding frames the next two to five years around three “crosses”: cross-sensor, cross-modal, cross-body.

Xspark’s answer is a three-layer architecture, slow brain VLM, fast brain VTLA plus TWAM, and spinal cord VTA. The VLM handles semantics, high-level cognition and task planning. Touch first enters the fast brain alongside vision, language and action for long-horizon and dexterous tasks, then drops to the spinal layer for the fastest local reaction on collision, slip and safety. The same touch thus plays two roles: in the fast brain it helps the robot do better, in the spinal cord it guarantees the robot reacts in time.

Researcher demonstrating a tactile sensing robotics prototype
Building trusted physical intelligence is the business Xspark says it actually wants. (Source: Leiphone)

Ding is not claiming a mature “tactile large model” exists. The realistic near-term path is pragmatic: perfect one multimodal tactile sensor, then fuse it; focus pre-training on general capability, introduce touch in dexterous-operation post-training and motion correction; and only long term let touch grow its own model and logic. What he is doing is finding the answer to a sharper question than whether robots need touch.

Editor’s note: This is an adapted translation of the original Leiphone interview. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/robot/RMYKt5OCvHa4yodm.html.

Leave a comment