InTime Robotics, Shanghai Jiao Tong and Fudan open-sourced HandEdit, a dexterous-hand editing benchmark

InTime Robotics, Shanghai Jiao Tong and Fudan open-sourced HandEdit, a dexterous-hand editing benchmark

Training a robot hand used to mean owning the robot, and that bottleneck is exactly why progress stalled. InTime Robotics, with Shanghai Jiao Tong University and Fudan University, has now open-sourced HandEdit, a dataset and benchmark that turns ordinary first-person videos of human hands into paired robot-hand images.

HandEdit dataset showing a human first-person view and the matched robot-hand render
HandEdit maps a human first-person action onto a target robot-hand configuration while keeping the scene intact. (Source: OFweek)

The release is large. HandEdit ships more than 200 million edited image samples and 300,000-plus video clips, drawn from five egocentric datasets (EgoDex, ARCTIC, OakInk2, HOI4D, HO-Cap). It covers 26 robot configurations, 13 standalone dexterous hands and 13 hand-arm builds, across 600 scenes, 1,100 objects and 400 task types. InTime’s own RH56DFX and RH5DG2 hands are among the supported models, alongside arms from JAKA, KUKA, Panda, RM, UR and xArm.

The method is substitution, not capture. HandEdit removes the human hand from a clip and drops in a target dexterous hand that matches a specified URDF, while preserving the object, the contact and the background. A pipeline using SAM3 for segmentation, ProPainter for inpainting and MANO-to-URDF retargeting renders the matched robot foreground back into the repaired scene, then screens out bad samples.

InTime RH56DFX dexterous hand operating a coffee machine in the HandEdit demo
InTime’s RH56DFX hand is one of the benchmark’s supported configurations. (Source: OFweek)

The benchmark grades results on three layers: general similarity (PSNR, SSIM, LPIPS, FID), vision-language judgement of semantic consistency and perceived quality, and embodied-task metrics that check the hand was fully removed, the structure is correct and the interaction survived. Eleven image editors were tested, including GPT-Image, Nano Banana, FLUX, Seedream, Qwen-Image-Edit, HunyuanImage, FireRed and OmniGen2.

The finding is uncomfortable for the image-generation crowd: a pretty picture is not a correct robot. Some models scored high on language judgement while failing structure and interaction. HandEdit’s point is that for robot data, finger geometry, robot identity and contact must each be checked on its own.

Editor’s note: This is an adapted translation of the original OFweek report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://robot.ofweek.com/2026-09/ART-898890-8120-30701741.html.

Translated and adapted from OFweek (link).

Leave a comment