OpenAI’s Astra Ultrafast announcement promised faster code generation. Its implementation guide for Ultrafast mode addresses a practical question for developers: how do you request that speed, and what might keep an application from benefiting?
The API change is small: request gpt-6-astra with service_tier set to ultrafast. The engineering decision takes more work. OpenAI recommends persistent WebSockets for agents that make frequent tool calls, warns that default Ultrafast rate limits are low, and directs developers to a separate pricing schedule. Each could affect whether the faster tier is practical for a real workload.
Request Ultrafast With a Service Tier, Not a Different Model
OpenAI’s guide identifies gpt-6-astra as the model and ultrafast as the service tier. In a Responses API request, the configuration looks like this:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"service_tier": "ultrafast",
"input": "Write a short Python function that removes duplicates while preserving order."
}'
Ultrafast is requested through a field on the API call, so adapting existing code does not call for silently substituting a different model identifier. Developers can make a request over HTTP, as above, then decide whether repeated interactions warrant the persistent WebSocket approach OpenAI recommends.
OpenAI says Ultrafast is available to all API developers, though default limits may provide little capacity. Its September 30 developer announcement also describes access through Codex and ChatGPT Work. Those product paths have different billing and eligibility rules from an API request.
Faster Generation Does Not Guarantee a Faster Agent
OpenAI claims Ultrafast delivers up to six times faster token generation in the API. That is a claim about model output speed under favorable conditions, not a promise that a complete application task will finish six times sooner.
An agent might generate a tool call, wait for a search or database result, read it, and then generate a follow-up. The total time includes connection setup, network travel, tool execution and pauses between model responses. Faster token generation affects only part of that sequence. When model responses are short, the other delays can become more noticeable.
OpenAI’s guide recommends persistent WebSockets for agentic applications with frequent tool calls. Keeping a connection open can avoid repeated HTTP connection overhead across a multi-step exchange. The guide also describes an HTTP alternative; WebSockets are not required to request Ultrafast. An application making occasional independent calls may have little reason to change its transport before measuring the benefit.

For a tool-heavy agent, test the complete loop. Record the time to the first useful output, each tool round trip and completion of the user’s task. Compare those results with the same workflow on the existing tier to see how much of the generation-speed gain reaches the user.
Default Rate Limits May Be the First Scaling Constraint
OpenAI says every API user can access Ultrafast at low default rate limits. Its guide provides rate-limit tables and directs developers who need higher limits to their account teams. A successful test request therefore does not establish that the tier can sustain a production agent’s traffic.
One user action can prompt an agent to issue several requests. Concurrent users add demand, and a burst can hit a limit even when average traffic appears manageable. Developers should check the limits applicable to their account against expected peak usage, not only a quiet prototype.
Test what happens when capacity is exhausted, too. Users may end up waiting through retries or failed requests during busy periods. If the workflow needs higher limits, account-team approval is a deployment dependency to address before launch.
OpenAI’s published speed figure is an up to claim. Connection choice and rate limits can help explain a smaller improvement in practice without contradicting the narrower token-generation claim.
The API Price Needs a Per-Task Calculation
Ultrafast has separate API token pricing, including higher rates for GPT-6 Astra when long-context pricing applies. Developers should consult the current pricing table for the input and output rates that match their workload before projecting costs. A subscription price cannot stand in for that calculation.
For an agent, cost per completed task may be more useful than cost per request. A task can send the same instructions or accumulated conversation through multiple steps, call tools, and generate several responses before reaching an answer. Faster output does not, by itself, reduce the number of billable tokens.
Run representative tasks on the existing configuration and on Ultrafast. Count the tokens used, price each run under the relevant schedule, and measure completion time. If the application uses long contexts, test those separately instead of assuming short-prompt economics still apply.
Selective use may make sense. An interactive coding task or time-sensitive agent step might justify a higher price for a measurable delay reduction. Work that runs without a user waiting for each response has a weaker case. The decision depends on the application’s traffic, not OpenAI’s maximum speed figure alone.
Codex’s 300-Token Claim Comes With a Different Access Path
For Codex, OpenAI claims up to eight times faster token generation, reaching up to 300 tokens per second with Astra Ultrafast. That ceiling belongs to the Codex claim; the announcement describes the API separately as delivering up to six times faster generation. Neither figure establishes a sustained rate for every prompt or a proportional reduction in coding-task completion time.
OpenAI says Astra Ultrafast is available in Codex and ChatGPT Work through the new Pro 500 plan and eligible Enterprise access. Developers can request it through the API. Business Insider reported that Pro 500 costs $500 per month and includes OpenAI’s highest plan usage allowance.
The monthly subscription provides access in the named products. It is not the price of Ultrafast API tokens and does not indicate that API usage is covered by the plan. An individual using Codex, an Enterprise workspace and a team building on the Responses API face different purchasing decisions.
OpenAI said Ultrafast for GPT-6.1 Sol was coming soon in the announcement. The configuration in this guide is for GPT-6 Astra; developers should not assume the same availability applies to Sol.
Final Thoughts
The guide gives builders a request setting to try and practical conditions to check before scaling: transport for repeated agent calls, available capacity and token pricing.
A representative task run end to end is the strongest adoption test. Ultrafast has a concrete purpose if it cuts the delay users experience enough to justify its cost and the required capacity is available. If tool waits, connection overhead or rate limits dominate, a higher token-generation ceiling alone will not solve the problem.
Frequently Asked Questions
4 questions
1How Do I Enable Astra Ultrafast in the OpenAI API?
Set model to gpt-6-astra and service_tier to ultrafast in a Responses API request. OpenAI’s guide shows Ultrafast over WebSockets and an HTTP alternative, so you can make an HTTP request before changing your connection strategy. Check your account’s rate limits before relying on the tier for production traffic.
2Does Astra Ultrafast Require WebSockets?
No. OpenAI documents an HTTP alternative but recommends persistent WebSockets for agentic applications with frequent tool calls. Keeping a connection open can reduce repeated connection overhead. For occasional independent requests, measure the complete task before deciding whether WebSockets would help.
3How Fast Is Astra Ultrafast in Codex and the API?
OpenAI claims up to eight times faster token generation in Codex, with an up-to-300-token-per-second ceiling, and up to six times faster generation in the API. These are vendor-reported generation figures, not guaranteed end-to-end task speeds. Tool execution, connection overhead, rate limits and output length can affect the improvement a user experiences.
4Is Pro 500 Required to Use Astra Ultrafast?
No. OpenAI says API developers can request Astra Ultrafast through API billing, subject to low default rate limits. For Codex and ChatGPT Work, OpenAI identifies Pro 500 and eligible Enterprise access. Business Insider reported that Pro 500 costs $500 per month; Ultrafast API tokens have separate pricing.
Sources
- implementation guide for Ultrafast modedevelopers.openai.com
- September 30 developer announcementcommunity.openai.com
- separate API token pricingdevelopers.openai.com
- Business Insider reportedbusinessinsider.com




