Updated October 5, 2026 · 6 min read

GPT-6 Astra benchmarks: what the scores really mean

If you read the same Astra benchmark reported three times and got three different numbers, you're not going crazy. The gaps come from different evaluation harnesses, different effort settings, and vendor-reported versus independent runs. Here's the honest scorecard — including where the marketing exceeds the evidence.

Why the numbers disagree

Three things cause most of the confusion:

  1. Different harnesses. The clearest example is ARC-AGI-3: OpenAI reports 99.9%, the official ARC Prize organisation reports 62.7% on the same benchmark. Same model, same benchmark name, 37-point gap.
  2. Different effort settings. Reasoning models produce different scores at low vs max. A low-effort run and a max-effort run are not the same experiment.
  3. Vendor-reported vs independent. Vendor tables choose which comparisons to publish. Independent evaluators frequently find narrower gaps.

So when two sources disagree, the honest answer is neither "one is lying" — it's that they measured different things. Only your own workload settles it.

The honest scorecard

BenchmarkReportedCaveat
ARC-AGI-399.9% (OpenAI harness) / 62.7% (ARC Prize)Huge harness sensitivity — treat any single number skeptically
OSWorld 2.0 (computer use)72.6%Tasks finish roughly 47% faster than GPT-5.6 Sol
FrontierMath Tier 497.6%Vendor-reported
Terminal-Bench 4.057.9% (up from 37.3%)Large gain on long-horizon agentic coding work
Artificial Analysis Intelligence Index61.2Behind Claude Fable 5.1 at 65.7
ExploitBench100% (few-shot)Gated — public build refuses exploit code
The two rows worth pausing on: Terminal-Bench 4.0 jumping from 37.3% to 57.9% is the strongest evidence for real capability gain on agentic coding. But losing to Claude Fable 5.1 on the Artificial Analysis index shows the "clearest model ever" framing doesn't survive contact with independent measurement.

The harness problem, concretely

The clearest caution: an early tester reported that a simple search agent beat Astra on some benchmarks within days of launch. That's not a scandal — it's what a fast-moving field looks like when vendors ship and third parties measure simultaneously.

The practical takeaway isn't cynicism. It's that vendor benchmarks are a starting point for your evaluation, not the end of it.

What users actually report

Launch-week reports from real teams:

Build your own eval in an afternoon

Because the published numbers disagree, the only number that matters is the one from your own tasks. A workable minimum:

  1. Pick 5-8 representative tasks from your real work — not benchmarks, your actual job.
  2. Run each at low and high effort to see the quality/cost curve for your workload.
  3. Compare against your current model (GPT-5.6 Sol or Claude Fable 5.1) on the same tasks.
  4. Measure tokens per completed task, not price per token. A 2.5x higher per-token rate can still be cheaper if it finishes in a third of the turns.
  5. Record wall-clock time too if a human is waiting on the result. That may justify Ultrafast.

Then route accordingly: Astra for what it's genuinely better at, the cheaper model for everything else. That's cost engineering, and it's where the real savings are.

Related reading

For the pricing side of this decision, see GPT-6 Astra pricing. For the head-to-head model comparisons, see Astra vs Claude Fable 5.1 and Astra vs Gemini 3.8. For hands-on usage patterns, see how to actually use GPT-6 Astra.