What happened
On 8 October 2026 OpenAI announced that Ultrafast mode is rolling out for GPT-6.1 Sol in the OpenAI API, in Codex and in ChatGPT Work. The announcement describes it as “near-Astra intelligence at up to 8x faster speeds than Sol Standard”. It does not give a benchmark or a method behind the 8x figure, so treat it as OpenAI’s own claim.
Ultrafast is a service tier, not a new model. In the API it is switched on per request with service_tier set to ultrafast, and it also works with GPT-6 Astra.
Key details
- API price for GPT-6.1 Sol in Ultrafast mode: $12 per million input tokens and $60 per million output tokens, according to OpenAI’s announcement. The pricing page lists Sol Standard at $2 and $10, so Ultrafast costs six times as much per token. OpenAI’s post puts the Sol Ultrafast price at 1.2 times the cost of Astra; on the pricing page that matches Astra at its Standard rate ($10 input), while Astra in Ultrafast mode costs $60 input.
- Cached input: the pricing page lists $0.60 per million tokens for Sol in Ultrafast mode, against $0.10 at Standard.
- Astra in Ultrafast mode: the pricing page lists GPT-6 Astra at $60 input and $300 output per million tokens in this tier.
- Codex and ChatGPT Work: access is included on Pro 500, on eligible usage-based Enterprise plans and on credit-based Edu plans. On Enterprise, administrators must enable it.
- Regions: Ultrafast for GPT-6.1 Sol is available in all supported regions, including US and EU data residency. OpenAI also added EU data residency for GPT-6.1 Sol Fast and GPT-6 Luna Fast.
- Rate limits: Ultrafast has its own limits, separate from Standard and Fast. For Sol the default is 1,000,000 tokens per minute on the Build tier, 4,000,000 on Launch and 40,000,000 on Grow.
- Connection: OpenAI’s documentation strongly recommends WebSockets, especially for agents that make many quick tool calls. Without a persistent connection, network overhead can reduce the latency gains. HTTP is supported as an alternative.
- What OpenAI says it is for: debugging outages, agents that navigate apps, and live experiences where response time matters. The documentation says to use it “when speed justifies the higher cost”.
Why it matters
For work where a person or a live system waits on every response, a model that answers several times faster can change what is practical, such as interactive coding help or an agent that clicks through an app. The price is the trade-off. Six times the Standard rate means Ultrafast only makes sense where the saved time is worth more than the added cost, and a long agent run can add up quickly.
The 8x speed is a ceiling (“up to”), stated by OpenAI without a published method. We found no independent measurement at the time of writing. Test it on your own prompts and tool calls, with a persistent connection, before you budget around it.
Disclosure: Claude, made by Anthropic, was one of the AI tools used to research and draft this article. Anthropic is a competitor of OpenAI. Speed and capability claims are OpenAI’s own.
Sources
- Ultrafast is rolling out today for GPT-6.1 Sol in the API, Codex, and ChatGPT Work Primary source , OpenAI Developer Community (OpenAI announcement)
- Ultrafast mode Primary source , OpenAI
- API pricing Primary source , OpenAI