Touch is becoming a base capability of the next generation of embodied intelligence. Over two years the field poured resources into bodies, large models, dexterous hands and vision. As those advanced, the bottleneck gathered at contact: a robot must not only find an object but sense force, friction, slip and deformation while gripping, plugging and assembling.
On 10 August Daimon Robotics released its tactile-anchored world model Daimon-TWM. One day later it closed a several-hundred-million-yuan strategic round, its second nine-figure cheque in two months. PaXini launched a foot multi-axis tactile sensor, PX-FOOTRIX. In July 1X put tactile skin on its new NEO hand. Boston Dynamics Atlas and Figure 03 list touch as base config. For leading firms, the question is settled: robots need touch.

The move worth watching is Daimon linking visuo-tactile sensors, native tactile data, a VTLA and Daimon-TWM into one route, so the model learns to understand, predict and use touch.
From sensing contact to predicting consequence
At WAIC the booth showed two tasks. One packed fragile fruit, controlling force to lift, place and close a box. The other tidied a pencil case, unzipping, placing items and closing, with a deforming bag and changing zipper resistance. The author dropped a marker on the table; the robot paused, re-judged the space and adjusted. That pause matters: in real plants and homes, position, hardness and friction are not fixed, and the robot must read mid-contact change.
Take the most common factory task, USB plug and unplug. Once the plug nears the port, the key contact is hidden; a camera cannot tell aligned from jammed. A human feels it. A standard VLA model often emits a motion block and runs it, so slip, bump and jam accumulate. Daimon-TWM puts touch into the understand-predict-control chain: read contact state, infer consequence, correct at 100 Hz. If the plug jams, pull back; if the object slips, change force.
Daimon’s UniTacVLA paper tested eight tasks, adjust, wipe, insert and assemble. With no disturbance the model averaged 64 per cent success, with disturbance about 53 per cent, against a comparison model at 26 and 6 per cent. Ablation showed adding touch alone helped little; understanding contact state, predicting future touch and correcting at high frequency raised success clearly.
The result points a road: old embodied models learned from video how a human acts; a tactile world model learns why a human slows here, changes angle there, and stops pushing when resistance appears. Robots move from copying motion to understanding its physical consequence, and Daimon-TWM belongs on that axis.

No common tactile language yet, but models pick a path
Robot touch has no single route. Piezoresistive and capacitive suit large areas, Hall and magnetic arrays measure multi-axis force, visuo-tactile reads deformation through a camera for contact, force and texture. The author’s read: robots will not bet on one route. But for fit with large models, visuo-tactile is best today because its images and deformation fields are data existing AI knows. Meta’s Sparsh used over 460,000 visuo-tactile images; Daimon walks the same road.
The hard part is letting one model read touch image, deformation, shear, depth and force. Daimon’s “unified tactile token” tackles this: turn each tactile modality into a token the model understands, and let touch run through perception, prediction and control, not sit as an isolated input.


Editor’s note: This is an adapted translation of the original OFweek Robot report. It has been trimmed and restructured for readability for an international business audience.