Keep going
Keep learning with Rundown Pro
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime
Guides OpenAI published sep 30, 2026
Start with Standard for work that can wait. Try Fast when shorter pauses help you finish sooner. Consider Astra Ultrafast when the time saved justifies the extra usage. These are our starting recommendations.
OpenAI introduced Pro 500 on September 29, 2026, with Astra Ultrafast access in ChatGPT Work and Codex. Eligible Enterprise and Edu workspaces can also get access. The API has separate access and billing rules.
This guide compares access, speed, and cost, then walks you through a decision matrix. You can choose a mode without building an app or downloading a template.
These options control how quickly OpenAI processes work with a supported model. Your model choice and reasoning setting remain separate decisions. OpenAI describes the speed controls as a way to increase speed while preserving model intelligence.
For example, choosing GPT-6 Astra with Ultrafast keeps Astra as the model. You should still check its answers and any actions it takes.
| Mode | What it offers | Our suggested starting use |
|---|---|---|
| Standard | The model's baseline processing speed and usage rate. | Reports, routine tasks, and work you can leave running. |
| Fast | OpenAI advertises up to 2.5× faster speeds in the API. Gains vary by model and product. | Writing, analysis, and coding while you actively review each response. |
| Ultrafast | OpenAI advertises up to 8× faster API speeds. Its Codex guide specifies an Astra token-generation comparison. | Time-sensitive Astra work that involves repeated exchanges. |
These are OpenAI's maximum speed claims. Total task time also includes tools, network delays, and review. Keep the model and product attached to each claim; the numbers do not establish the gain for every task.
A token is a small unit of information the model processes. Faster token generation can shorten the wait for a response, though other parts of the task may still take time.
A useful comparison starts with a task and a clear definition of a correct result.
In Intro to AI Evals, Nate Grahek teaches how to use real inputs, expected answers, and repeatable checks to compare models and workflows. The course covers quality, speed, and cost using spreadsheet and Claude Code examples.
Watch Intro to AI Evals with University Pro →
University Pro and OpenAI subscriptions are separate. The course teaches evaluation methods; it does not provide Ultrafast access.
Checked against OpenAI documentation on September 30, 2026. Speed figures are vendor claims, and the time and cost examples are illustrative. We have not run paid benchmarks or verified speed options in an authenticated account.
Start with your product and plan. The Ultrafast controls described here belong to ChatGPT Work and Codex. OpenAI documents a separate API service tier for developers.
| Your setup | Decision |
|---|---|
| ChatGPT Plus, Pro 100, Pro 200, or Business | Compare Standard and Fast where available. These plans lack Ultrafast at launch, even with extra credits. |
| ChatGPT Pro 500 | You can select Astra Ultrafast in Work and Codex. The plan costs $500 per month. |
| Eligible Enterprise or Edu workspace | Check admin permissions and the workspace's billing agreement. Enterprise owners must enable Ultrafast access; legacy rate-limit-based Enterprise plans are unsupported. |
| An app using the OpenAI API | Astra Ultrafast is available to all API users under its rate and processing-region limits. A ChatGPT Pro 500 subscription is unnecessary. |
For an eligible ChatGPT account: open Work or Codex, select GPT-6 Astra, then choose Ultrafast in the model picker. Check your selected workspace when an expected option is missing.
OpenAI lists these default Astra Ultrafast token limits:
| API usage tier | Tokens per minute |
|---|---|
| Tiers 1–3 | 500,000 |
| Tier 4 | 1,000,000 |
| Tier 5 | 5,000,000 |
Treat these as capacity limits for the API. ChatGPT allowances follow separate rules.
The API supports global processing and US data residency. EU and other non-US regional processing endpoints are unsupported. Workspaces that require inference outside the US also lack access; the user's physical location alone does not decide eligibility.
For a planned API rollout, check the limits against your expected traffic before choosing a speed.
Think about the task you need to finish. Use this table as a starting rule.
| Your situation | Start with | Why |
|---|---|---|
| You can leave a report running and review it later. | Standard | A shorter wait may add little value. |
| You are editing a proposal and checking each revision as it arrives. | Fast | Shorter pauses may make the editing session more useful. |
| You need Astra for a difficult task and will ask several follow-up questions under a tight deadline. | Ultrafast, when available | Faster responses may save useful time across repeated exchanges. |
| You have many routine tasks and a fixed budget. | Standard | Keep more of the budget available for completed work. |
| The answer keeps missing facts or instructions. | Review the prompt, sources, and model first. | Correctness needs its own check before you spend more on speed. |
A difficult task can still run at Standard speed. Choose the model and reasoning level that produce an acceptable answer, then decide how much you value a faster response.
Some tasks spend much of their time generating an answer. Others wait for websites, file operations, or outside services. OpenAI's guidance treats model speed, output length, request count, and tool design as separate ways to reduce delays.
| Most of the wait comes from… | Decision |
|---|---|
| Generating a long response or making repeated model calls. | Compare Fast and Ultrafast when the saved time matters. |
| Repeated network exchanges between an API app and OpenAI. | Review connection reuse as well as the speed setting. |
| A slow website, database, or connected app. | Check that service before paying for a faster model response. |
| More detail than you need. | Ask for a shorter result before increasing speed. |
| Reviewing and correcting weak output. | Improve the source material and review rules first. |
| A task you will not read until tomorrow. | Keep Standard as the default. |
Here is a hypothetical example. A task takes 60 seconds: 40 seconds waiting on tools and 20 seconds generating text. If text generation becomes eight times faster, that part falls to 2.5 seconds. The full task still takes 42.5 seconds, assuming everything else stays unchanged.
OpenAI strongly recommends a persistent WebSocket connection for Ultrafast tasks with frequent tool calls. HTTP requests also work.
A WebSocket keeps the connection open across steps. OpenAI's WebSocket guide describes continuing with new input and the previous response ID. This reduces per-turn continuation overhead; a slow external tool still needs its own check.
This is guidance for your own API integration. In Work or Codex, choose an available speed through the product's controls.
Use the time to a correct, usable result when judging the upgrade.
OpenAI distinguishes included subscription allowance, purchased credits, and API charges. Your plan uses its included allowance first, then available credits for eligible usage.
For the same supported model, OpenAI lists these multipliers relative to Standard:
| Usage type | Standard | Fast | Astra Ultrafast |
|---|---|---|---|
| Included subscription allowance consumed | 1× | 2.5× | 8× |
| Purchased-credit charge | 1× | 2× | 6× |
These figures describe usage and charges. They do not predict how much sooner a task will finish. Enterprise usage-based billing follows the workspace's agreement.
Under an equal-workload assumption, an allowance that covers 80 Standard runs would cover about 32 Fast runs or 10 Ultrafast runs. Real task sizes vary, so treat this as a ratio example rather than a promise about your plan.
Decision: Choose a faster mode when the time it saves justifies the extra usage. Keep Standard for work where you can wait.
For GPT-6 Astra, current short-context API text rates are shown below. Every price is in US dollars per million tokens, for prompts with up to 272,000 input tokens.
| Mode | Input | Cached input | Cache writes | Output |
|---|---|---|---|---|
| Standard | $10 | $1 | $12.50 | $50 |
| Fast | $20 | $2 | $25 | $100 |
| Ultrafast | $60 | $6 | $75 | $300 |
Keep the model fixed when comparing these rates. Cached input reuses saved prompt content; cache writes create that saved content. Longer prompts, tools, and regional processing can change the bill.
For an illustration with 20,000 ordinary uncached input tokens and 5,000 total billed output tokens, the token charges would be $0.45 on Standard, $0.90 on Fast, and $2.70 on Ultrafast. This calculation excludes cache-write fees, tools, and other charges.
API charges are separate from ChatGPT subscriptions. Codex with your own API key uses API pricing, so the ChatGPT allowance and purchased-credit multipliers above do not apply to that route.
You can use different speeds for different tasks. We recommend a clear default so each new feature does not silently raise your usage.
| What you have learned | Your choice |
|---|---|
| Standard finishes soon enough and meets your quality checks. | Keep Standard. |
| Standard's delays interrupt your work, and Fast gives useful time back. | Use Fast for those interactive tasks. |
| You need Astra, have Ultrafast access, and the shorter waits justify the cost. | Reserve Ultrafast for that work. |
| Tools, network overhead, or review dominate the task's duration. | Address those delays before spending more on model speed. |
| You have no evidence yet that faster processing changes the outcome. | Keep Standard until you can compare. |
For a team, write the rule in plain language: “Standard for routine work. Fast for active editing and analysis. Ultrafast for approved, time-sensitive Astra tasks.”
Keep your review standards the same at every speed. For a fair comparison, hold the model, reasoning level, task, source files, tools, and cache conditions steady. Record the full time, errors, and cost of each attempt.
Keep going
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime
The API guide advertises up to 8× faster speeds. The Codex guide defines its comparison as Astra token generation. A full task also includes tool calls, network delays, and review.
OpenAI advertises up to 2.5× faster speeds in the API. Product and model details still matter: its Codex guide lists a 1.5× speed increase for GPT-5.6 and GPT-5.5. Keep those claims separate from Fast's 2.5× included-allowance rate.
OpenAI's August 13, 2026 preview used GPT-5.6 Sol and reported up to 14× faster processing and 750 output tokens per second. The September 29 Astra rollout uses a different model and comparison. Keep the model name attached to any speed claim.
The current API guide still lists GPT-5.6 Sol Ultrafast as preview access. Broad API availability applies to Astra.
On Pro 100 and Pro 200, extra credits do not unlock Ultrafast at launch. Among personal Pro plans, access requires Pro 500. Eligible Enterprise and Edu workspaces have separate rules.
In the API, OpenAI renamed Priority processing Fast mode on July 30, 2026. The API accepts both fast and priority as service-tier values for supported models.
Choose speed to reduce waiting. To address an incorrect answer, review the sources, instructions, model, and reasoning level. Keep facts and output checks in place even when responses arrive quickly.
It deserves consideration when you already need Astra, spend meaningful time waiting for it, and can justify the extra usage. For occasional tasks or work you can review later, we would start with Standard. Fast is the next option to consider when delays become costly.