Updated October 6, 2026 · 6 min read
GPT-6 Astra computer use: what it can do on a real desktop
Astra is the first OpenAI flagship built for direct computer operation — not just answering questions about work, but doing it. It navigates real applications, completes multi-step professional tasks, and produces finished artifacts. Here is what it can do, how well it scores, and the safety gating that shapes who gets access.
What "computer use" means for Astra
Most models answer. Astra operates. OpenAI describes it as a model that can manage software and professional workflows with minimal intervention: navigating graphical interfaces, chaining tool calls, and finishing jobs that used to need a human at the keyboard.
The difference is architectural, not just a new prompt. Astra is specifically optimized for computer use and professional-grade accuracy, built on years of reinforcement learning work aimed at agentic execution. In the Codex harness it goes further — instead of compressing long sessions into summaries, it keeps searchable notes across context windows, which is what makes hours-long debugging sessions practical.
Real tasks OpenAI demonstrates
OpenAI's documentation shows Astra completing complete professional tasks end-to-end, including:
- PCB layout in KiCad — electronic design work, not just code generation
- Building 3D objects in Blender — operating creative software through its interface
- Managing financial models in Excel — the classic knowledge-work environment
That breadth matters. A coding agent that only writes code is a tool. An agent that can drive KiCad, Blender, and Excel is closer to a junior colleague with a very wide skill set.
How well does it actually perform?
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | What it measures |
|---|---|---|---|
| OSWorld 2.0 | 72.6% | — | Real desktop tasks; Astra completes them in ~40 minutes, 47% less time per task |
| AutomationBench | 41.4% | 18.1% | Professional workflow automation — more than 2x its predecessor |
| FrontierMath Tier 4 | 97.6% | — | Research-level mathematics |
| ARC-AGI-3 | 99.9% | — | Abstract reasoning, with OpenAI's provider adapter harness |
Two caveats worth keeping in mind: the OSWorld figure is OpenAI's own reporting, and the ARC-AGI-3 score depends on a specific harness (the provider adapter that preserves reasoning state). Scores measured differently will not match. For a skeptical read of every headline number, see our honest benchmarks breakdown.
Professional artifacts, not generic text
A distinct Astra capability is artifact generation: output trained to be polished and template-adherent. That means documents, spreadsheets, and presentations that follow the formatting conventions of the template you give it — a deliverable you can send, rather than raw text you have to reformat.
For agentic workflows this compounds: the agent finishes the work and produces the file your workflow expects, in one run.
Who can use it
| Channel | Status |
|---|---|
| ChatGPT Plus / Pro / Business / Enterprise | Rolling out; Astra appeared in the model picker for staged access after launch |
| Enterprise workspaces | Off by default — an administrator must enable Astra per workspace; workspace access does not grant API access |
| OpenAI API | Available via gpt-6-astra |
| Amazon Bedrock | Live, including an "UltraFast" mode up to 300 tokens per second |
If your engineers can see Astra in ChatGPT but your API calls return model-not-found, that is the enterprise gotcha: workspace access and API access are two independent switches. See our guide to finding Astra in ChatGPT.
The safety gating behind computer use
A model that can operate a computer can also attack one. OpenAI's system card states Astra is the first model to reach the Critical level of cybersecurity capability under the Preparedness Framework — including discovering and using previously unknown zero-day vulnerabilities in internal testing, two of which were in the process of being disclosed to maintainers.
The consequences for users:
- Exploit-creation capabilities are gated under the Daybreak program
- Stricter isolation and checkpoint encryption on OpenAI's side
- Full trajectory monitoring including chain of thought, with misalignment monitoring on all tool-using inference — a long agent run may pause for review
Frequently asked questions
Does computer use cost extra?
Astra's API rates are the same regardless of task type: $10 input / $50 output per million tokens on Standard. The Ultrafast tier (6x price) is what accelerates interactive computer-use sessions where every step waits on the previous one.
Can Astra use my computer right now?
Computer use runs inside OpenAI's harnesses — ChatGPT desktop, Codex, and the API — not arbitrary software you install. Access follows the Astra rollout: limited organizations first, then broader ChatGPT plans.
Is Astra the best model for computer use?
Its 72.6% on OSWorld 2.0 leads the reported field, and AutomationBench more than doubles GPT-5.6 Sol. But these are vendor-reported scores on new benchmarks; independent replication is still catching up. See how it compares in our hands-on review.
What happens if the agent makes a mistake on my machine?
Treat it like granting access to a new hire: start with sandboxed or low-stakes environments, keep destructive actions behind confirmation, and review trajectories. OpenAI's own monitoring may pause a run for review regardless of your settings.
How is this different from GPT-5.6 Sol's agent mode?
Three ways: it operates end-to-end in real applications (KiCad, Blender, Excel), it keeps searchable notes across context windows for long sessions, and it produces template-adherent artifacts. The benchmark gap — 41.4% vs 18.1% on AutomationBench — reflects that gap in capability.