Updated September 26, 2026 · 6 min read

GPT-6 Astra benchmark results, explained

Astra posts eye-watering numbers — 99.9% on ARC-AGI-3, 97.6% on FrontierMath. But benchmarks measure narrow things, and knowing which score maps to which real-world job is the difference between a smart purchase and an expensive mistake.

The full scoreboard

BenchmarkWhat it measuresAstraComparison
OSWorld 2.0Real computer-use tasks72.6%Sol 65.7%, in 47% less time
Terminal-Bench 4.0Agentic coding in a terminal57.9%Sol 37.3%
FrontierMath Tier 4 v2Research-level mathematics97.6%Sol 83.0%, Fable 5.1 87.8%
ARC-AGI-3Abstract reasoning / novel problems99.9%—
ExploitBenchFinding & exploiting vulnerabilities100%First model to max it
BenchCAD3D object reconstruction from renders95.9%Sol 83.3%, Fable 84.3%
AA Intelligence IndexBroad general intelligence61Sol 61, Fable 5.1 66

Three numbers that actually matter

1. OSWorld 2.0 — 72.6%

This is the one to watch. It measures whether a model can genuinely operate a computer: navigating apps, completing workflows. Astra beats Sol's score while taking ~47% less time per task. If you care about AI doing work autonomously, this is your metric.

2. ExploitBench — 100%

Astra is the first model to max it out, which is why OpenAI rates it "Critical" in its Preparedness Framework and gates exploit development behind a vetted defensive-security program. Impressive — and a reminder that capability and risk travel together.

3. AA Intelligence Index — 61

The reality check. On broad general intelligence, Astra ties the previous Sol generation and trails Claude Fable 5.1 by five points. The frontier moved forward in specific directions — math, tool use, computer operation — not across the board.

What the benchmarks don't tell you

How to use benchmarks wisely

The honest method: take 3–5 tasks you actually do, run them on Astra and on your current model, and compare. Thirty minutes of your own testing beats any leaderboard. Our review does exactly this for common professional work.