On 22 August at the Ice Ribbon arena in Beijing, as the second World Humanoid Robot Games opened, CCTV cut to a real tennis court. A humanoid stood on the baseline. A ball came at more than 50 kilometres per hour. In milliseconds it judged the landing point, moved its feet, turned its body and swung. The ball brushed the net cord and landed in.

Serves, forehands, backhands, baseline rallies and net volleys held up throughout. In the doubles warm-up it covered for its human partner and switched tactics. In one high-speed exchange it made a desperate retrieval, fell hard, pushed itself up and kept playing. The humanoid that completed the world’s first fully autonomous tennis rally is from Galbot.
Ten years ago AlphaGo showed that AI could think. Galbot set out to show that AI can do. The gap between thinking and doing took a decade, and that is why this counts as a marker: it does not just move a single benchmark, it puts a humanoid on a real court where thinking and acting must be online at the same time. Galbot named it the AstraTennis moment.
Its strokes look nothing like the robots of a few years ago. The hitting is human-like, with feet arriving first, body driving the arm and upper and lower body working together. More unusual is generality. It handles serves, forehands and backhands, uses baseline movement, net volleys and placement control, plays singles and doubles, and adjusts tactics to the format. It coordinates with a human partner and makes decisions on the spot, and after falling in high-speed play it gets up and continues. Judging a player also means judging the opponent, and across the net stood Zheng Jie, a former Grand Slam doubles champion, so the robot had to execute while working out how to win the point.

The answer to playing and thinking at once is Galbot’s embodied foundation model, AstraBrain. Its defining feature is welding the pieces a humanoid needs into one model. The brain understands and decides, judging the ball, choosing the shot, organising tactics and coordinating with a doubles partner. The cerebellum handles motor control, keeping whole-body balance at speed, producing explosive swings and human-like movement. A pons translates the brain’s decisions into motor commands. Historically these sat in separate modules and the interfaces dropped signals. AstraBrain makes them one, which is the basis for calling it the world’s first end-to-end brain-cerebellum-nerve-control integrated whole-body and whole-hand model.
On court the value is visible. A ball arrives and the brain judges in milliseconds whether it is deep or short and whether to attack or defend, while the cerebellum is already executing a slide left, a crouch and a backswing, and at contact the brain updates its read, deciding the next shot must press the line. That real-time loop of thinking while playing only closes in an integrated architecture.
In June, during CVPR 2026, the team wrote about AstraBrain-WBC 0.5, which pushed zero-shot generalisation to 92.58 per cent in the lab with inference latency under 1.5 milliseconds. AstraTennis is the first outing of the same system in real opposition. A strong cerebellum is still only half a machine. Whether brain and cerebellum can coordinate under real pressure is the true coming-of-age test for embodied intelligence.
Galbot founder and chief technology officer Wang He calls tennis the ultimate exam for a humanoid, since whoever answers it proves they have solved both cerebellum and brain. Showing good motor control and real-time decisions on court requires heavy practice off it. Humans repeat drills and copy professionals from video. Robots traditionally learn motion by teleoperation, but in tennis at speed a human cannot react in time, so teleoperated data is useless, and recording a full match with a motion capture system is prohibitively expensive. The old road was blocked, so Galbot took another: AstraBrain Latent, the world’s first whole-body real-time intelligent planning and control algorithm for tennis opposition. The name is deliberate, since latent means hidden, and the algorithm’s job is to mine the underlying patterns of how to play from imperfect, incomplete human motion data.
The thinking is counterintuitive. Robot makers assume motion data should be as professional as possible, ideally complete moves from expert teleoperators captured frame by frame. Galbot’s team did the opposite, skipping perfect data to collect ordinary people’s forehands, backhands, side slides and crossover steps, then using algorithms to combine, correct and generalise those fragments into a full tennis skill set. The reason is simple: perfect data is too expensive and too hard to obtain, and imperfect data already carries the core priors of human movement, how weight shifts, how the arm swings, how the foot pushes off. The rest is left to the algorithm. The idea matters beyond tennis because it means the threshold for robots learning motor skills can drop sharply. No need to wait for elite athletes or buy costly motion capture, since real skill can be distilled from rough data.

To keep motion consistent at high speed, the team added a latent action barrier, drawing a boundary in latent space that keeps the robot’s movement inside a human-like style while adjusting stance and swing to the incoming ball. Random perturbations during training taught it to self-correct. By that stage it plays live balls and adapts on the spot.
Tennis is hard enough that data alone is not enough. Galbot also built its own virtual tennis world on top of a self-developed, hundred-billion-scale embodied dataset called Galaxy Workshop. Training runs in two steps: massive rehearsal in simulation where the robot plays, reflects and spars with virtual opponents of different levels, then limited calibration on the real machine. The most important step is multi-agent game play. Rather than imitating humans move by move, Galbot lets robots play each other, with two or more agents competing in simulation, each rally forcing an adjustment. Over time abilities never explicitly taught emerge on their own, the skill emergence that matters most in the foundation-model era, where enough data produces a qualitative jump and the robot learns to generalise.
Crucially, skills learned in the virtual world transfer to a real court, which is why the AstraTennis moment happened on a global broadcast rather than in a lab video. That completes the chain: AstraBrain thinks and moves, AstraBrain Latent distils skill from incomplete data, and Galaxy Workshop supplies training scenarios and opponents.

The most striking image of the match was the fall. In a high-speed exchange the humanoid missed a desperate ball and hit the ground hard. No staff ran in, no one called a halt, and it pushed itself up and readied for the next ball. That detail is more striking than any parameter, because it shows the motor control is strong enough to recover from real-world accidents. Zooming out, the path Galbot opened, where brain decisions, cerebellum execution and data emergence sit in one system, offers a reusable paradigm for embodied intelligence. Previously each new action meant new data collection and retraining. A general motion-control base means dance, inspection, rescue and chores can share one body operating system with marginal cost falling as scale rises.
The capability is already running elsewhere. In smart pharmacies robots pick targets from tens of thousands of products and hand them to couriers, and on industrial lines the Galbot S1 can lift 50-kilogram payloads with two arms. Tennis is just the latest piece lit up. From the digital world to the physical one, from a board game move to a tennis swing, AI took a full decade to step out of the screen and onto the floor of human life.
Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience.