Updated October 5, 2026 · 5 min read
GPT-6 Astra's context window: 1.05M tokens, and the 272K cliff
GPT-6 Astra's context window is 1,050,000 tokens with up to 128,000 tokens of output. Sounds like you can shove anything in. There's a catch: once your input passes 272,000 tokens, the entire request is billed at roughly double — not just the overflow. Here's what that means in practice.
The actual numbers
| Spec | GPT-6 Astra |
|---|---|
| Context window | 1,050,000 tokens |
| Maximum input | 922,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Modalities | Text and image in, text out (no audio or video) |
| Fine-tuning | Not supported |
These figures are identical to GPT-5.6 Sol. The context window is the headline — but it's the least interesting number for planning, because you will rarely approach it.
The 272K billing cliff
Requests are priced in two bands. Crossing the threshold moves the whole request up a tier, not just the part above it:
| Rate (per million tokens) | Up to 272K input | Above 272K input |
|---|---|---|
| Input | $10.00 | $20.00 |
| Cached input | $1.00 | $2.00 |
| Cache writes | $12.50 | $25.00 |
| Output | $50.00 | $75.00 |
So input is 2x, output is 1.5x, and the surcharge applies retroactively to the full request. A request of 273K input tokens that would have cost $2.73 now costs $5.46 — the extra thousand tokens doubled the entire bill. On Ultrafast the same cliff applies at $60→$120 input and $300→$450 output.
What actually fits in practice
Rough token-to-material conversions, useful for sizing a job before you send it:
| Material | Approximate tokens |
|---|---|
| A typical source-code file | 1,000–10,000 |
| A large codebase, 100K lines | ~1,000,000 (fills the window) |
| A long technical document or contract | 20,000–50,000 |
| A book, full text | ~100,000–200,000 |
| Hundreds of agent execution records | Varies widely |
The honest framing: 1.05M is enough to hold an entire repository or a large document set in one shot, which removes the need for retrieval pipelines in a lot of cases. Whether you should is a cost question, and the next section answers it.
Planning around the cliff
- Stay under 272K unless the long-context case is the point. If a chunked or retrieval approach works, it is dramatically cheaper at Standard rates.
- Use prompt caching before you use a bigger window. Cached input runs at $1 per million against $10 uncached — a 90% discount. On cache writes at $12.50, the break-even sits around a 78.3% cache hit rate. The same holds on Ultrafast, where cache reads are $6 against $60.
- On long-horizon agent work, the gains are real. Terminal-Bench 4.0 goes from 37.3% to 57.9% versus the previous generation. On single-shot knowledge questions the improvement is marginal, and you are paying 2.5x per token for it — route accordingly rather than switching wholesale.
- Measure your own token count before the cliff, not after. A token-count estimate against your real prompt is the cheapest guard against a surprise bill.