GPT-6 Astra Ultrafast: 300 Tokens Per Second at Six Times the Price
OpenAI's Ultrafast tier runs GPT-6 Astra at about 300 tokens per second in Codex and 6x API throughput, priced at $60 input and $300 output per million.
By Dillip Chowdary • Sep 30, 2026 • Source: OpenAI
OpenAI introduced Ultrafast at DevDay 2026, a premium speed tier that accelerates GPT-6 Astra to roughly 300 tokens per second in Codex — up to eight times its standard generation speed — and up to six times the throughput in the API. The speed costs real money: Ultrafast is priced at six times standard rates, which for GPT-6 Astra means $60 per million input tokens and $300 per million output tokens, against the standard $10 and $50.
This piece covers what Ultrafast actually is, where it is available and for whom, how the pricing math works against standard Astra, and which workloads justify a 6x multiplier. It is written for engineering teams already paying for frontier-model inference and wondering whether latency is worth this much.
Ultrafast: what OpenAI shipped
Ultrafast is not a new model — it is GPT-6 Astra served at drastically higher generation speed, offered as a paid tier for workloads where waiting on tokens is the bottleneck. In Codex, OpenAI quotes up to 8x faster token generation, around 300 tokens per second; in the API the ceiling is up to 6x standard throughput. The recap positioned it alongside availability in ChatGPT, Codex, and the API rather than as an API-only option.
Access is gated by tier: Ultrafast is available in the API now, and in ChatGPT Work and Codex for Enterprise customers and subscribers to the new Pro 500 plan — the $500-per-month subscription OpenAI launched the same day with roughly 25 times the usage allowance of Plus. A GPT-6.1 Sol Ultrafast variant is promised in the coming days, which will bring the speed tier to the much cheaper new model.
The pricing math: standard vs Ultrafast

| GPT-6 Astra, per 1M tokens | Standard | Ultrafast |
|---|---|---|
| Input | $10 | $60 |
| Output | $50 | $300 |
| Codex generation speed | baseline | up to 8x (~300 tok/s) |
| API throughput | baseline | up to 6x |
The multiplier is uniform — six times both input and output — so the relative cost of a workload does not change shape, it just scales. A long agent run that costs $1 on standard Astra costs $6 on Ultrafast and finishes in a fraction of the wall-clock time. Note the asymmetry between the marketing numbers: the 8x speedup figure is quoted for Codex generation, while API throughput tops out around 6x, so measure your own path before budgeting on the larger number.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
Where the 6x premium pays for itself
The tier makes sense where a human is actively blocked on generation: an engineer steering a coding agent, a support workflow with a customer on the line, or interactive computer-use sessions where every step waits on the previous one. In agentic loops the effect compounds — an agent that takes hundreds of sequential model calls finishes the whole chain several times sooner, which can matter more than per-token cost for time-sensitive work.
It makes far less sense for batch and background workloads, which is precisely the segment OpenAI addressed the same day with GPT-6.1 Sol at one-fifth of Astra's standard price. The pricing ladder now runs from $2-per-million-input Sol for volume work, through standard Astra for frontier quality, to $60-per-million Ultrafast Astra when speed is the product. Picking the right rung per workload is now a real cost-engineering decision.
Who is affected: Enterprise, Pro 500 and API users
API developers can adopt Ultrafast immediately with no plan change — it is a pricing decision, not a subscription one. Inside ChatGPT Work and Codex, though, the tier doubles as a reason to upgrade: Enterprise seats and the $500 Pro 500 plan are the only ways to get Ultrafast speeds interactively. Combined with the quiet reduction of the $200 Pro plan's allowances announced the same day, the subscription lineup now clearly steers heavy individual users toward Pro 500.
For teams running Codex heavily, the practical effect is that agent throughput becomes a purchasable quantity. A team can decide that its coding agents should run at 300 tokens per second during incident response or release crunch, and pay for exactly that period via API usage rather than re-platforming.
What to watch after GPT-6 Astra Ultrafast
The near-term item is GPT-6.1 Sol Ultrafast, due in the coming days — Sol's $2/$10 base prices would put its Ultrafast rate at a far more approachable level if the same 6x multiplier holds, though OpenAI has not yet published that number. Also unstated so far: latency SLAs, rate limits at Ultrafast speeds, and whether the tier extends to the Agents API's new computer-use workloads, where sequential steps make speed most valuable.
Before committing spend, benchmark your actual pipeline: the 8x figure applies to Codex token generation, and real-world gains depend on how much of your wall-clock time is generation versus tool calls, retrieval, and orchestration. If generation is under half your loop time, Ultrafast's ceiling on end-to-end speedup is correspondingly lower — measure first, then buy speed where it actually shows up.
Developer Action Items
- ☐ Diff the official changelog for OpenAI / ChatGPT / Codex 6.1 before you bump — APIs, defaults, and removed flags only.
- ☐ Install through the vendor's documented channel in staging; keep a one-command rollback and time-box the canary.
- ☐ Grep your repo for old flag names, lockfile pins, and plugin versions that the notes mark as breaking.
- ☐ Prefer the first patch cut over the day-zero tag unless you have a reason to be on the leading edge.
- ☐ If OpenAI did not name a region, plan, or SKU, screenshot the official availability line before you promise it to users.
Author
Dillip Chowdary
Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.
Related on Tech Bytes
OpenAI's $500 Pro 500 Plan Arrives as $200 Pro Limits Get Halved
Read →
Codex Security Cloud Scans GitHub Repos and Drafts Fixes Remotely
Read →
Codex Code Review Lands in ChatGPT Desktop With Background Scans
Read →
OpenAI Decisions API Returns Model Classifications in 150 Milliseconds
Read →
Today's Tech Pulse briefing
Full briefing →
Advertisement