GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. One Did It With Less.

A gimbal robot arm is a camera mount that tilts and pans. The author connected one to a Mac mini and let three top models, GPT-6 Astra, Claude Fable 5.1 and GPT-5.6 Sol, each drive it, with the same instruction: make the robot sweep its full range of motion, however you like.

GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On (Source: LeiPhone)

All three finished the task. But they understood “finished” very differently. GPT-5.6 Sol acted like a coverage cartographer, walking every one of the four extreme corners. GPT-6 Astra Pro was a restrained test engineer, using fewer moves yet adding the strictest feedback verification at every step, and ran up the highest bill of the session. Claude Fable 5.1 was an aggressive field engineer, fast, combining broad coverage and a reusable deliverable, but most optimistic about declaring success.

Astra turned four moves into seven. Clever, or lazy?

On video, Astra clearly moved less. It compressed the job into seven steps: centre, pan left, pan right, re-centre, pitch up, pitch down, re-centre. The two axes each reached both ends and the task ended. Fable did twelve steps, adding the four diagonal extreme positions. Sol’s logs confirm it also walked all four corners.

From a test engineer’s view, Astra’s logic holds: a two-axis device only needs each axis to reach its limits to prove range. The corners add no new single-axis information, just two already-verified motions stacked. Fable and Sol read the instruction more like a human watching a demo: if you said sweep it, show me the whole space. But for a real machine, what Astra dropped may be redundant, or may be exactly where risk hides, in cable tension and structural interference at combined limits.

The author notes a public test where Astra Max used fewer actions than the human baseline in 96 per cent of solved levels on ARC-AGI-3, averaging 51.7 per cent fewer actions. That is not proof the gimbal choice was right, only a hint that action compression may be Astra’s stable tendency, not a one-off.

Fastest versus slowest: speed reveals risk appetite

The difference needs no log, only ears. Fable spun fast and loud, barely pausing between moves. Astra was restrained, like an engineer touching someone else’s equipment for the first time, leaving margin at every step. Astra used SPD=40, the medium parameter from the original reset script; Fable used SPD=0, which in this firmware means unlimited speed, the fastest setting.

Yet Fable was not maximal everywhere. It set its horizontal bound at plus or minus 165 degrees, five degrees short of Astra’s plus or minus 170. Fable floored the accelerator on speed but pulled back on position bounds. A model’s risk preference is not a single slider from cautious to aggressive. It is a distribution: some fear hitting limits, some fear slow movement.

Before moving, Astra said it would leave margin at the endpoints to avoid hitting the limit, told the user to keep the cable slack and move their hand away. The author has been pinched by this gimbal, which has strong torque and no anti-pinch sensor. The model knew none of that pain, but inferred a risk the prompt never wrote and the scene really carried.

The command was sent. Who says it is done?

Astra defined completion most strictly. For each target it waited up to 12 seconds, marking a point verified only when horizontal and pitch error were both under three degrees and three consecutive valid feedback readings matched, then errored out if voltage dropped below 9V or the target never arrived. Only after all seven targets returned real angles with verified status did it print SWEEP_COMPLETE_CENTER_VERIFIED.

Fable wrote the whole twelve-step routine into a script called gimbal_full_sweep.py. On a second run, suspicious feedback appeared: the lower-left horizontal position never updated and re-centre read no feedback, yet the green check still printed. Fable later admitted the design flaw and said the script should fail when error exceeds a threshold.

The gap shows up in every real agent task. The code changed, did the test actually pass. The transfer fired, did the other side receive. The form submitted, did the server accept. Once agents enter the physical world, the definition of done gets harsh.

From Sol to Astra: less is more

Sol walked the four corners; Astra deleted them. Astra grabbed the truly independent variables, cut ineffective environmental interaction, and spent the saved attention on point-by-point verification. In public coding-agent benchmarks Astra used about one third of Sol’s tokens, though Fable 5.1 still led the index at 70. OpenAI’s safety review shows Astra’s computer-misuse rate fell from Sol’s 22 per cent to 2.4 per cent, capability hallucination from 9.4 per cent to 2.0 per cent, and honeypot trap rates from about 48 to 56 per cent down to zero.

But the same materials note Astra’s written reasoning is harder to monitor, because it uses fewer written steps. Intelligence Index scores for Astra and Sol are both 61, five below Fable 5.1; Astra’s hallucination rate dropped from 92 per cent to 51 per cent, HLE rose six points, yet GDPval-AA v2 fell 80 Elo.

The price is still steep

One gimbal test and a short dance cost GPT-6 Astra Pro USD 23.43 and Claude Fable 5.1 USD 5.70. Astra made 11 tool calls, Fable four on first success; Astra read source five times and opened the serial port four times, Fable zero and one; Astra’s context was 112K tokens, Fable’s 76.3K. Astra’s official API price is USD 10 per million input tokens and USD 50 per million output, 2.5 times Sol’s USD 4 and USD 20.

If it controls a desktop demo, Astra may not be worth it. Fable is faster, more expressive, and leaves a reusable script. If it controls a production database, a corporate account, a high-risk device, or anything that cannot be easily undone, the extra money may buy one fewer false success.

Understanding the real world is the final gate

Astra’s doing less may be advanced action compression, or it may miss real-world coupling risk. Fable’s doing fast may be user-friendly, or let the success report run ahead of the facts. Sol’s full coverage is proven on video, but whether it crosses the gap between execution and verification is still untested.

What decides whether an agent moves from a clever chat partner to an actor that can enter life and production is not benchmark scores. It is the small decisions models make by default: how much to do, how fast, where to risk, when to stop, whether to explain or halt, and whether to look back at the world after sending the command.

GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On (Source: LeiPhone)
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On (Source: LeiPhone)
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On
GPT-6 Astra, Fable 5.1 and Sol Each Drove the Same Robot. On (Source: LeiPhone)

Editor’s note: This is an adapted translation of the original LeiPhone report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.leiphone.com/category/yanxishe/hfRMbwkVyjskWX3a.html.

Leave a comment