Imagine a task: open the cupboard, find the blocks hidden inside, then build them into a bridge. For a person this is trivial, a few actions flowing together that a nursery-school child completes without effort. For a robot it cuts across completely different abilities: finding the blocks needs visual understanding and language following; opening the door needs complex interaction with the environment; stacking needs geometric planning and precise manipulation while guessing how the world will change.

Worse, those abilities belong to different model families: VLA (vision-language-action), WAM (world-action), RL policies and TAMP (task-and-motion planning). Each is strong in its lane and isolated from the others, trained differently, fed different input formats, living in different state spaces. In robotics today, capability is not the problem. Coordination is.
The mirage of one model for everything
This summer saw a cluster of harness-themed works all pointing at one realisation: a bigger model will not solve everything for now, and robots need a layer of executive organisation. In July, Tsinghua professor Yu Chao’s team published Harness VLA, importing the harness layer widely used in digital intelligence into embodied AI so a frozen VLA focuses on contact-rich manipulation while a harness layer learns when and how to call it, and how to reset or retry after failure. On the LIBERO-Pro perturbation benchmark that lifted success from 50 per cent to 82.4 per cent.
RoboHarness, from Huawei’s Noah’s Ark Lab, addresses a different layer: what happens when a task exceeds the boundary of any single model. Both are called harness and both use a coding agent as the organising layer, but they solve different dimensions. One makes a specialist more stable; the other makes a team of specialists fight together. As robot tasks grow more complex, the second problem is becoming impossible to dodge.
How the conductor decides who plays
RoboHarness wraps independently built controllers (VLA, WAM, RL, TAMP) as uniformly schedulable agentic skills, with a coding agent making high-level decisions. A crucial premise: it does not retrain or alter any underlying policy. Each keeps its original implementation; RoboHarness only adds an organisation layer on top.
Choosing the right policy is harder than it looks. A coding agent cannot reliably tell from a raw image which policy fits, because that decision hides quantitative questions a language model cannot answer, such as how similar the current scene is to a policy’s training distribution. The system adds three helper skills. Understanding skills turn raw input into quantifiable signals (embedding similarity to each policy’s training data, pose stability, lighting and image quality). Memory skills retrieve past execution experience and, through a Memory Bridge, reconcile state-distribution mismatches between policies. Evolution skills update policy metadata and orchestration logic from online feedback.
The relay is won at the handover
Two independently trained policies usually live in separate data worlds with different input formats and different experience distributions. Where TAMP stops is very likely a state VLA has never been trained to start from. The Memory Bridge, which handles stable handoffs, is the single most important module in RoboHarness. It retrieves the Top-K similar trajectories a target policy once executed successfully, models the state range that target policy is comfortable inheriting, then samples and scores candidate bridge states and generates a transition trajectory.
In an ablation, removing the Memory Bridge barely changed task progress but drove full success from 86.0 per cent down to 60.4 per cent. Countless tasks reached the final step and failed at the handoff.
Two lineages, one quiet convergence
RoboHarness’s authors observe two long academic lineages in North American robotics: the US west coast (Stanford, Berkeley) leans toward learning and scaling, while the northeast and Canada (MIT, Toronto, MILA) emphasise hierarchy, symbolic reasoning and structured planning. RoboHarness extends the latter. In industry the same split shows: Zhipu leans toward scaling, while Sudo Tech and Magic Atom lean toward reusable skills and hierarchy; abroad, Physical Intelligence’s pi series represents scaling, while Gemini Robotics continues the high-level-reasoning-plus-low-level-skills tradition.
Recently the two routes are converging. Physical Intelligence pushed controllability and skill composition to the centre in pi0.7, while Nvidia keeps shipping more hierarchical robot-agentic systems. Foundation-model scaling and agentic hierarchy are evolving from two separate paths into complementary capacities inside one robot system.
RoboHarness has real limits: it can only schedule policies that already exist. But its core argument stands. The competition in embodied intelligence is quietly extending into system-level design, from hunting the strongest model to organising the most fitting capability.
Editor’s note: This is an adapted translation of the original Leiphone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/ai/vE6z6buPMecczPrc.html.