GPT-6 Astra Ultrafast billing starts 5 October 2026 at $60/M input and $300/M output tokens
OpenAI's Ultrafast service tier for GPT-6 Astra exits its grace period on 5 October 2026. Here's what you'll pay and what to expect.
If you enabled the Ultrafast service tier for GPT-6 Astra during its grace period and haven’t checked your billing settings, now is the time. From 5 October 2026, OpenAI has started charging for Ultrafast usage, and the numbers are meaningfully higher than standard rates.
What changed on 5 October 2026
OpenAI previewed the Ultrafast service tier on 13 August 2026, giving API customers a window to test it without incurring charges. That grace period is now closed. Any Ultrafast request made on or after 5 October 2026 will appear on your next invoice.
The pricing is straightforward: Ultrafast costs exactly six times the standard GPT-6 Astra rate across every billing line.
| Token type | Standard | Ultrafast |
|---|---|---|
| Input (uncached) | $5.00 / 1M | $60.00 / 1M |
| Cached input (read) | $0.50 / 1M | $6.00 / 1M |
| Cache writes | $12.50 / 1M | $75.00 / 1M |
| Output | $25.00 / 1M | $300.00 / 1M |
These rates apply to short-context requests, defined as prompts at or under 272,000 input tokens. Longer prompts are billed at a further uplift on top of the standard rate, and that uplift applies to Ultrafast too.
How to use it
There is no separate model string or model card for Ultrafast. You access it by passing service_tier: "ultrafast" on any request to gpt-6-astra in the Responses API. It is a per-request choice, not a plan or subscription add-on on the API side, so you can mix Standard and Ultrafast calls within the same application.
{
"model": "gpt-6-astra",
"service_tier": "ultrafast",
"input": [...]
}
Rate limits for Ultrafast are 500,000 tokens per minute on usage tiers 1 through 3, rising to 1,000,000 on tier 4 and 5,000,000 on tier 5.
One regional constraint worth noting: Ultrafast supports US data residency and global processing only. EU and other non-US regional inference residency endpoints are not supported. If your architecture requires EU data residency, Ultrafast is not an option at this point.
What you are paying for
The speed uplift is substantial. For GPT-6 Astra specifically, OpenAI’s DevDay 2026 announcement on 29 September described Ultrafast as delivering up to 8x faster token generation in Codex (around 300 tokens per second) and up to 6x faster in the API. For comparison, an earlier Ultrafast preview on GPT-5.6 Sol was clocked at up to 14x faster than standard processing with up to 750 output tokens per second. The GPT-6 Astra figures are lower, which likely reflects the model’s greater size and reasoning depth.
For most API workloads, faster token generation matters in real-time applications: voice interfaces, coding assistants where a developer is waiting on a response, and interactive agentic tasks where latency compounds across multiple steps. If your use case is batch processing or background work, Standard is almost certainly the right choice.
What does this mean for you?
The cost difference between Standard and Ultrafast is significant enough to warrant a deliberate decision for each use case. To put it concretely: a call with 20,000 input tokens and a 2,000-token output costs roughly $0.30 on Standard. The identical call on Ultrafast costs $1.80. At scale, that six-times multiplier accumulates quickly.
For API developers, the immediate action is to audit which of your integrations had service_tier: "ultrafast" enabled during the grace period. If any production pipelines were running Ultrafast experimentally, check whether the speed gain justifies the cost at your volumes. Because the tier is set per request, you can be surgical about it, applying Ultrafast only to latency-sensitive endpoints and routing everything else to Standard.
For ChatGPT subscribers, access to Ultrafast in ChatGPT Work and Codex requires the Pro 500 plan at $500 per month. Credits purchased on Pro 100 or Pro 200 do not unlock Ultrafast. Enterprise and Education workspace admins can enable Ultrafast for their users through workspace permissions, with usage drawing from workspace credits at the same six-times rate.
A note on long-context requests
If you are working with GPT-6 Astra’s full context window (the model supports up to 922,000 input tokens), be aware that requests exceeding 272,000 input tokens are billed at a further premium on top of the standard rate, applied to the entire request rather than just the tokens above the threshold. That uplift stacks with the Ultrafast multiplier, so very large-context Ultrafast calls can become expensive quickly. OpenAI’s usage dashboard breaks down your spend across uncached tokens, cache reads, and cache writes separately, which makes it easier to spot where the cost is coming from.
If your region requires a data residency endpoint, note also that models released on or after 5 March 2026 (which includes GPT-6 Astra) carry a 10% uplift on regional endpoints. That applies on top of whichever service tier you select.
The short version
Ultrafast is a well-defined, useful capability at a well-defined price. The six-times cost over Standard is not a surprise, it was published when the preview launched. What changes from 5 October 2026 is that the grace period is over. Review your integrations, confirm which service tiers are active in production, and make sure the applications running Ultrafast are the ones where speed genuinely justifies the premium.