Updated October 9, 2026 · 8 min read
GPT-6 Astra as a robot brain: the embodied-AI leaderboards, honestly
On October 4, 2026, SuperCLUE published its embodied-brain leaderboard: GPT-6 Astra took comprehensive global #1 on real-robot testing. Four days later, independent benchmarks filled in the caveat every headline left out: even the best model still solves under a quarter of the tasks, and real-robot trials had to be halted early for safety. Here are all the scores, the failure modes, and what leaderboards actually measure.
The SuperCLUE embodied-brain leaderboard
Published October 4, 2026, the SuperCLUE embodied-brain ranking measures how well frontier models act as the "brain" of a physical robot — environment perception, object recognition, long-horizon task planning, physical-world prediction, and dynamic response. Its distinguishing feature: real robot platforms and live task verification, not simulation-only scores.
- GPT-6 Astra took comprehensive global #1, leading on long-chain logical reasoning and complex task decomposition.
- Alibaba DAMO Academy and ZTE scored identically, tied for the domestic (China) #1.
- Alibaba's model is stronger on spatiotemporal memory — a robot that gets interrupted remembers what it was doing. ZTE's leans toward edge-side deployment: hardware compatibility, low latency, embedded stability.
- The reported gap: domestic models still trail on total score, on long chained multi-step tasks, and on generalizing to unseen physical environments.
RoboQuest: the independent reality check
RoboQuest is an independent evaluation of five frontier VLMs on 10 MuJoCo tasks that require actively acquiring hidden information — not passive chat. Results from the run of September 30 – October 6, 2026:
- GPT-6 Astra: 23.2% success, at $11.29 per episode — top of the board.
- Claude Opus 5.5: 13.8%. GPT-6.1 Sol: 12.2%. Claude Fable 5.1: 11.4%. Gemini 3.8 Flash: 2.0%.
- The failure-mode breakdown is the interesting part: 43% of Astra's 1,143 failures were "missing evidence" — the robot simply stopped exploring, with target compartments never opened — not execution error, since isolated skills hit 80–81%.
In other words, Astra doesn't fumble the actions; it gives up on gathering information before it has enough to act.
RoboWorld: 63 of 84 tasks unsolved
RoboWorld, a simulation testbed posted to Hugging Face spanning manipulation, mobile manipulation, locomotion, driving, and aerial control, produced the same shape of result:
- Frontier models cleared only 21 of 84 physical-world tasks; 63 went unsolved by any of the five models tested.
- GPT-6 Astra led at 16/84 (19.0%); Claude Opus 5.5 followed at 13/84 (15.5%). Kimi K3, DeepSeek V4.1 Flash, and Gemini 3.8 Flash each cleared a single task.
- The profiles split cleanly by the kind of control demanded: Astra succeeded more often on spatial and constrained-contact goals; Opus 5.5 on continuous-balance and timed-interaction goals, and cleared 3 of 4 aerial tasks.
- Astra's evaluation cost an estimated $9,913 against 941M tokens.
RoboDojo: real-robot trials halted for safety
An early RoboDojo evaluation added the sobering real-world data point:
- GPT-6 Astra averaged a 22.48% success rate in simulation, ranking 1st of 43 entries — ahead of specialized, pretrained robot policies.
- Real-robot evaluations were prematurely halted for safety: Astra repeatedly issued physically unreasonable or unsafe commands on actual hardware, resulting in equipment damage.
- A counter-intuitive finding: one-shot demonstrations (image or text) decreased aggregate performance — the model apparently tries to copy layout-specific geometry from the demo instead of adapting to the current scene.
- The paper's own stated limitation: Astra "lacks an intuitive grasp of contact dynamics, force sensitivity, and collision geometry," and the turn-by-turn tool interface cannot easily express continuous-contact or non-quasi-static motions.
Leaderboard #1 ≠ robots that work
Putting all four sources together — and they are consistent:
- Astra genuinely leads every embodied benchmark published so far: SuperCLUE's real-robot leaderboard, RoboQuest (23.2%), RoboWorld (16/84), RoboDojo simulation (1st of 43).
- The best model still solves under a quarter of the tasks. Even the leaderboard's #1 entry fails roughly three out of four goal-directed robot tasks.
- The failure mode is exploration, not dexterity. Isolated skills run at 80%+; the missing piece is deciding when to keep looking for evidence before acting.
- Safety is not solved. Real-hardware trials of the best model were stopped early — twice, in two different evaluations — for physically unsafe commands.
So the honest summary for October 2026: GPT-6 Astra is the current best robot brain, and robot brains are still bad. If you're buying the "AI runs your warehouse robots" narrative, the numbers say otherwise — see our honest benchmarks breakdown for the same story on the software side, and the StarCraft incident for what happens when a model decides exploration is optional in a game instead.
Frequently asked questions
What is an embodied AI "brain"?
A model that serves as the decision-making core of a physical robot: perceiving the environment, recognizing objects, planning long multi-step tasks, and predicting physical outcomes. SuperCLUE's embodied-brain leaderboard evaluates models on real robot platforms with live task verification, not simulation-only scores.
Which AI model is the best robot brain right now?
GPT-6 Astra, per the SuperCLUE embodied-brain leaderboard of October 4, 2026 — comprehensive global #1, with Alibaba DAMO Academy and ZTE tied for the domestic top spot.
Did GPT-6 Astra pass the robot benchmarks?
No. It led every board — 23.2% on RoboQuest, 16/84 on RoboWorld, 1st of 43 on RoboDojo simulation — but under a quarter of tasks were solved, and 63 of 84 RoboWorld tasks went unsolved by any model tested.
Why did the real-robot tests stop early?
Safety. In RoboDojo's real-hardware trials, Astra repeatedly issued physically unreasonable or unsafe commands, causing equipment damage, so evaluations were halted prematurely. The study cites missing physical commonsense — contact dynamics, force sensitivity, collision geometry — as the core limitation.
Does this change what GPT-6 Astra is good at?
It sharpens the picture: Astra's published strengths remain long agentic tasks and hard reasoning with checkable outputs — see our benchmarks breakdown and review. Embodied control is a frontier where it leads, but the absolute numbers show nobody is close to reliable physical autonomy yet.