Updated October 5, 2026 · 6 min read
GPT-6 Astra benchmarks: what the scores really mean
If you read the same Astra benchmark reported three times and got three different numbers, you're not going crazy. The gaps come from different evaluation harnesses, different effort settings, and vendor-reported versus independent runs. Here's the honest scorecard — including where the marketing exceeds the evidence.
Why the numbers disagree
Three things cause most of the confusion:
- Different harnesses. The clearest example is ARC-AGI-3: OpenAI reports 99.9%, the official ARC Prize organisation reports 62.7% on the same benchmark. Same model, same benchmark name, 37-point gap.
- Different effort settings. Reasoning models produce different scores at
lowvsmax. A low-effort run and a max-effort run are not the same experiment. - Vendor-reported vs independent. Vendor tables choose which comparisons to publish. Independent evaluators frequently find narrower gaps.
So when two sources disagree, the honest answer is neither "one is lying" — it's that they measured different things. Only your own workload settles it.
The honest scorecard
| Benchmark | Reported | Caveat |
|---|---|---|
| ARC-AGI-3 | 99.9% (OpenAI harness) / 62.7% (ARC Prize) | Huge harness sensitivity — treat any single number skeptically |
| OSWorld 2.0 (computer use) | 72.6% | Tasks finish roughly 47% faster than GPT-5.6 Sol |
| FrontierMath Tier 4 | 97.6% | Vendor-reported |
| Terminal-Bench 4.0 | 57.9% (up from 37.3%) | Large gain on long-horizon agentic coding work |
| Artificial Analysis Intelligence Index | 61.2 | Behind Claude Fable 5.1 at 65.7 |
| ExploitBench | 100% (few-shot) | Gated — public build refuses exploit code |
The harness problem, concretely
The clearest caution: an early tester reported that a simple search agent beat Astra on some benchmarks within days of launch. That's not a scandal — it's what a fast-moving field looks like when vendors ship and third parties measure simultaneously.
The practical takeaway isn't cynicism. It's that vendor benchmarks are a starting point for your evaluation, not the end of it.
What users actually report
Launch-week reports from real teams:
- Strong for: multi-step computer tasks, long-horizon agentic coding, generating/refactoring/testing code in Codex, document work against your templates.
- Weak or absent for: video editing — one power user found it can't edit video at all. It is not a universal optimizer.
- Token efficiency: on agent tasks, roughly a third fewer tokens than the previous generation in some configurations.
- Watch for: staged, trust-gated rollout means a colleague on the same plan may not see Astra yet. That's not a bug on your end.
Build your own eval in an afternoon
Because the published numbers disagree, the only number that matters is the one from your own tasks. A workable minimum:
- Pick 5-8 representative tasks from your real work — not benchmarks, your actual job.
- Run each at
lowandhigheffort to see the quality/cost curve for your workload. - Compare against your current model (GPT-5.6 Sol or Claude Fable 5.1) on the same tasks.
- Measure tokens per completed task, not price per token. A 2.5x higher per-token rate can still be cheaper if it finishes in a third of the turns.
- Record wall-clock time too if a human is waiting on the result. That may justify Ultrafast.
Then route accordingly: Astra for what it's genuinely better at, the cheaper model for everything else. That's cost engineering, and it's where the real savings are.
Related reading
For the pricing side of this decision, see GPT-6 Astra pricing. For the head-to-head model comparisons, see Astra vs Claude Fable 5.1 and Astra vs Gemini 3.8. For hands-on usage patterns, see how to actually use GPT-6 Astra.