OpenAI launches Ultrafast service tier for GPT-6 Astra in the Responses API
OpenAI's Ultrafast tier delivers up to 8× faster token generation for GPT-6 Astra. Available now via service_tier: 'ultrafast'. Billing starts 5 October 2026.
OpenAI announced Ultrafast at DevDay 2026 on 29 September 2026, and it is exactly what it sounds like: a new, faster service tier for the Responses API, aimed at developers building latency-sensitive applications on GPT-6 Astra. Billing does not start until 5 October 2026, so you have a short window to test it at no cost.
What Ultrafast actually is
Ultrafast is a paid service tier for the Responses API, sitting above Standard in terms of both speed and price. OpenAI describes it as up to 8× faster than Standard for GPT-6 Astra, with throughput of around 300 tokens per second. The direct API figure is closer to 6× in some contexts, but either way the difference is substantial if your use case is sensitive to output latency.
To use it, you do not need a new model string. You pass service_tier: "ultrafast" on any request to gpt-6-astra:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-astra",
input="Explain why the sky is blue in one sentence.",
service_tier="ultrafast",
)
print(response.output_text)
That is the entire configuration change. The tier is broadly available to all API customers on GPT-6 Astra from 29 September 2026, with preview access also available for GPT-5.6 Sol.
What it costs
Ultrafast is priced at six times the Standard rate across all pricing dimensions. For prompts of 272,000 input tokens or fewer:
| Standard | Ultrafast | |
|---|---|---|
| Input | $10 / 1M tokens | $60 / 1M tokens |
| Cached input | $1 / 1M tokens | $6 / 1M tokens |
| Cache writes | $12.50 / 1M tokens | $75 / 1M tokens |
| Output | $50 / 1M tokens | $300 / 1M tokens |
For long-context prompts above 272,000 tokens, Ultrafast charges $120 per million input tokens and $450 per million output tokens.
Those numbers are not subtle. The output cost of $300 per million tokens is significant, and you would want a clear reason to justify it. For agentic workflows that make rapid, sequential tool calls where every second of latency compounds, it will make sense. For batch processing or anything that runs in the background, Standard is almost certainly the right choice.
OpenAI recommends using a persistent WebSocket connection for agents making many tool calls in quick succession, because without one, network overhead can eat into the latency gains Ultrafast is designed to deliver. It works over plain HTTP too, but you may not see the full benefit.
Rate limits
Default rate limits for Ultrafast are tiered by your API usage level:
- Tiers 1, 2, and 3: 500,000 tokens per minute
- Tier 4: 1,000,000 tokens per minute
- Tier 5: 5,000,000 tokens per minute
If you need higher limits, OpenAI suggests contacting your account team directly.
What this means for you
If you are building latency-critical applications, Ultrafast is worth testing before billing begins on 5 October 2026. Real-time voice interfaces, interactive coding assistants, and agentic pipelines that loop through tool calls quickly are the obvious candidates. Dropping in service_tier: "ultrafast" takes about thirty seconds, and you have the billing grace period to measure whether the speed gain justifies the cost at your volume.
If you are processing large batches or running non-interactive workloads, this tier is not for you. The 6× price multiplier only makes sense when your users or your system is actively waiting on each response.
If you are based in Europe, there is an important limitation to be aware of. Ultrafast supports US data residency and global processing only. EU regional inference residency is not supported. You can still send Ultrafast requests through global processing endpoints, but you cannot keep inference within an EU regional endpoint. If data residency within Europe is a compliance requirement for your organisation, you will need to stick with Standard for now.
If you are an Enterprise or Edu customer, your workspace owner needs to enable Ultrafast access explicitly. Legacy rate-limit-based Enterprise plans are not supported. Check with your admin before assuming the tier is available to your team.
ChatGPT subscribers should note that Ultrafast in the chat interface is reserved for ChatGPT Pro 500 subscribers ($500 per month). Plus, Pro 100, Pro 200, and Business plans do not include it at launch.
Also available on Amazon Bedrock
For teams running on AWS, OpenAI and Amazon announced GPT-6 Astra Ultrafast on Amazon Bedrock on 30 September 2026. On Bedrock, Ultrafast is available through US geographic and Global CRIS endpoints, and through Bedrock Managed Agents in us-east-1. Standard and Priority service tiers on Bedrock are not supported for this model.
The short version
Ultrafast is a straightforward addition to the Responses API: one parameter, a meaningful speed improvement, and a price that reflects it. Add service_tier: "ultrafast" to your requests, test your latency-sensitive workflows before 5 October 2026, and decide whether the cost is justified for your specific use case. For high-frequency agentic systems where output speed genuinely matters, it will be. For everything else, Standard remains the sensible default.