Caltech and Stanford’s HomeBody lets a GPT-class model tidy your kitchen

Astra learns to act, not just think

On 28 September, the roboticswire outlet reported that researchers at the California Institute of Technology and Stanford University have introduced HomeBody, a system that gives GPT Astra a persistent spatial memory and the ability to perform compound humanoid actions for long-horizon tasks. In a typical kitchen scene, HomeBody lets a Unitree G1, guided by GPT Astra, clean a room and retrieve remembered objects from vague requests. The authors say the system needs no training for a specific environment and no extra policy learning to do it.

Autonomous control of a humanoid usually uses a three-part architecture: a System 2 vision-language model handles visual information and instructions, a trained System 1 vision-language-action model generates control commands, and a System 0 controller coordinates the actual motions. As frontier models like Astra grow more capable, the team argues, a trained VLA is no longer required between high-level reasoning and robot skills. System 2 may directly manage a library of reusable motor skills.

A plug-and-play memory layer

HomeBody replaces the connection between the systems with a plug-and-play vision-language model that can call a composable skill library. Cleaning a kitchen means judging what to keep, what to discard and where each item belongs. Every time the robot moves, its view changes, so it must rely on memory and action feedback to track what is done and what remains. In one video, HomeBody groups coffee bags onto the kitchen counter and discards a specified milk carton, coordinating multiple trips, grabs and placements. In another, the medicine is initially out of view; the team says HomeBody uses stored keyframes to locate the drawer, takes out the medicine, hands it over and discards the drink carton. The videos show the robot opening a drawer with its right hand and discarding a carton with its left, two-hand coordination on a complex task, though the footage is heavily sped up and efficiency still needs work.

Real2Sim: building the twin from the robot’s own view

The team also showed how it built the system. First it gives the humanoid background on its role and lets it explore an unseen environment. HomeBody’s spatial data uses iPhone 0.5x lens video, D435i camera observations, SLAM lidar scans, joint poses and Astra-chosen waypoints. Because the robot records the room from its own viewpoint as it moves and interacts, the spatial context is built on what the robot actually observes, so even objects leaving its first-person view stay in context. Next, HomeBody uses Astra as a Real2Sim agent to build a digital-twin environment in Isaac Sim from the robot’s self-collected data, mapping observations to a spatial model so Astra can reason about positions outside its first-person view. Finally, instructions are given in the virtual world, for example tidy the kitchen, and HomeBody uses its spatial context to choose actions and goals without scripting at the action level.

One vague sentence drives the whole task

Architecturally, the robot first understands the environment through a semantic map, its own pose and keyframes, remembering latent context. When a user says bring the medicine and tidy up, GPT Astra breaks the vague sentence into steps: fetch the medicine, hand it to the person, throw away the carton. At the throw-carton step, Astra picks the grasp action from its skill library and further decomposes it into selecting the target, computing the grasp pose, planning the arm trajectory and finally driving the arm to reach, then feeding back the result before the next step. Throughout, no precise human command is needed; one sentence lets the robot see the environment, split the task, choose the action and do it, illustrating the current embodied-intelligence route of a large model as the brain and a skill library as the cerebellum.

The team showed the current skill library’s richness still needs improvement. The skills run on a Razer Blade laptop with an RTX 4090 GPU, handling both perception and motion planning, while GPT Astra runs remotely, sending skill requests and goals and receiving results. This lightweight deployment lets HomeBody run locally while still reaching frontier models over the network. But the team admits Astra’s inference latency causes pauses between skills, likely why the videos are sped up, and that even on an RTX 4090, adding more complex perception or skills may need more compute.

HomeBody again shows the potential of large models in embodied applications: no environment-specific training, no action scripting, one vague instruction drives a compound task. From controlling an arm to controlling a whole humanoid, it demonstrates what the brain model can do. But directly using a large model to drive a robot still has a long way to go.

Editor’s note: This is an adapted translation of the original Zhidx report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.zhidx.com/p/597689.html.

Leave a comment