Updated October 5, 2026 · 6 min read

GPT-6 Astra Ultrafast: 300 tokens/second at 6x the price

Ultrafast is not a new model — it is GPT-6 Astra served dramatically faster, announced on September 29, 2026 at OpenAI's DevDay. In Codex it generates up to 300 tokens per second, roughly 8x Standard. On the API it runs up to 6x faster and costs exactly 6x as much. Here is what actually changed, what it costs, and who should pay for it.

What Ultrafast actually is

Ultrafast is a service tier, not a model. The model ID stays gpt-6-astra. You turn it on by setting service_tier: "ultrafast", and OpenAI serves the same weights with much lower generation latency.

At launch, Astra is the only model listed in the Ultrafast pricing table. A GPT-6.1 Sol Ultrafast variant was promised for the days after DevDay, which would matter — Sol costs a fifth of Astra Standard, so Ultrafast on the cheaper model is far easier to justify for volume work.

The 8x and 6x figures measure different things. OpenAI's API documentation says "up to 8x faster than Standard"; its launch thread gives the API a separate ceiling of up to 6x. Codex is quoted at up to 8x, or 300 tokens per second. All of these are token-generation claims — not end-to-end agent task completion time.

Pricing: exactly 6x Standard, no exceptions

Rate (per million tokens)Astra StandardAstra Ultrafast
Input, up to 272K prompt$10.00$60.00
Cached input$1.00$6.00
Cache writes$12.50$75.00
Output$50.00$300.00
Input, above 272K prompt$20.00$120.00
Cached input (long)$2.00$12.00
Cache writes (long)$25.00$150.00
Output (long)$75.00$450.00

The multiplier is uniform — six times on every line — so the shape of a workload's cost does not change on Ultrafast, it just scales. A long agent run that costs $1 on Standard costs $6 on Ultrafast and finishes in a fraction of the wall-clock time.

Note the long-context rule, inherited from Standard: once input passes 272K tokens, the whole request is billed at the higher rate, not just the overflow. See our context window guide for where that cliff sits.

Which plans actually get Ultrafast

PlanUltrafast access
Pro 500 ($500/mo)Included — 25x the Plus usage allowance
Enterprise / EduEligible plans; off by default on Enterprise until an owner enables it
Pro 100 / Pro 200Not available at launch, even with purchased credits
Plus, Business, Go, FreeNot available
OpenAI APIAll API users, at low rate limits

Inside ChatGPT Work and Codex, Ultrafast draws your included usage at 8x the Standard rate, and purchased credits at 6x. OpenAI is explicit that these multipliers describe billing, not how much faster the model runs — don't read the 8x as a speed claim.

Pro 500 includes 25 times the Plus allowance. Spent entirely in Ultrafast at 8x, that goes about as far as a bit over 3x the Plus allowance used at Standard speed. Individual early buyers reported a two-minute Ultrafast session consuming 2% of a weekly allowance, and one user exhausting the weekly allowance in under 30 minutes. Those are reports, not published quotas.

How to enable Ultrafast in the API

Set the model to gpt-6-astra and service_tier to ultrafast on a Responses request:

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6-astra",
    input="Explain why the sky is blue in one sentence.",
    service_tier="ultrafast",
)

For agentic work with many tool calls, OpenAI recommends a persistent WebSocket connection and passing previous_response_id between turns. Network overhead on a fresh connection per call can erase the latency gain entirely.

Default rate limits for Ultrafast are their own, separate from Standard:

Usage tierUltrafast tokens per minute
Tiers 1–3500,000
Tier 41,000,000
Tier 55,000,000

Ultrafast supports global processing and US data residency only — no EU or other regional endpoints. On Enterprise, the workspace needs a credit-based or USD usage-based agreement; older rate-limit-billed Enterprise plans are not supported. Higher limits can be arranged through an OpenAI account team.

Is 6x the price worth it?

Ultrafast pays for itself when a human is blocked waiting on generation:

In sequential agentic loops the effect compounds: hundreds of model calls finish the whole chain several times sooner, which can matter more than per-token cost.

It makes little sense for batch and background work — exactly the segment OpenAI addressed the same day with GPT-6.1 Sol at one-fifth of Astra's Standard price. The pricing ladder now runs from $2 per million input tokens for Sol volume work, through Standard Astra for frontier quality, up to $60 for Ultrafast Astra when speed is the product.

Worth doing the math yourself: at sustained maximum generation speed, 8x faster at 6x cost can spend roughly 48x as much on output per wall-clock second. That worst case assumes you are generating flat out with no tool waits. Time spent waiting on tools changes the real rate dramatically.

Frequently asked questions

How fast is Ultrafast in real terms?

Up to 8x Standard token generation in Codex — around 300 tokens per second — and up to 6x Standard throughput on the API. One early tester running Astra at high effort saw lower displayed token rates because reasoning tokens count toward generation too. These figures are not benchmarked end-to-end task times.

Do I need the Pro 500 plan to use Ultrafast?

For the API, no — any API customer can set service_tier: "ultrafast" at low rate limits. The $500 Pro 500 subscription is only the route to interactive Ultrafast in ChatGPT Work and Codex.

Is Ultrafast a different model?

No. Same model, same gpt-6-astra ID, same capabilities — faster generation, higher bill. Context window stays at 1,050,000 tokens. See what GPT-6 Astra is.

Can I use Ultrafast outside the US?

Through the API, yes — Ultrafast runs on global processing as well as US data residency. EU and other non-US regional inference endpoints are not supported for Ultrafast.

How does this change my existing Astra bill?

It doesn't, unless you opt in. Standard rates are unchanged: $10 input / $50 output per million tokens. Full breakdown in our GPT-6 Astra pricing guide.