OpenAI starts rolling out Ultrafast, the fastest mode for GPT-6.1 Sol, in the API, Codex, and ChatGPT Work
IA en un minuto newsroom · Editor: Jon Elgezabal

In 30 seconds
OpenAI is gradually switching on Ultrafast, an option for GPT-6.1 Sol that puts response speed first in exchange for a higher price. Any developer can already use it in the API, at rates six times the normal ones. In Codex and ChatGPT Work it is limited to the $500 Pro tier and certain business and education accounts, where it burns through the allowance much sooner.
OpenAI has started rolling out Ultrafast for GPT-6.1 Sol, one of its AI models. Ultrafast is the fastest mode in its API, which developers use to connect the models to their own applications, and the company pitches it for latency-sensitive work, especially agents that make many tool calls in quick succession. According to OpenAI, with GPT-6.1 Sol it responds up to 8 times faster than the standard mode, with intelligence close to that of GPT-6 Astra, its other large model, for which the documentation also lists this mode.
The rollout was announced on October 8 on OpenAI's developer forum, and the official documentation already reflects it. In the API, it is open to all users and, according to the pricing page, costs $12 per million input tokens and $60 per million output tokens (tokens are the chunks of text used to measure usage), six times what GPT-6.1 Sol costs in the standard mode, which stays at $2 and $10. With long contexts, the price rises to $24 and $90, also six times the standard rate. It has its own rate limits, separate from those of the standard and Fast modes, and lets data be processed in the United States or in Europe.
It is also coming to Codex, its AI coding tool, and to ChatGPT Work, the version of ChatGPT built for work, which share usage, pricing and limits. There, it is only available on the $500 Pro plan and on some plans for companies (Enterprise plans that pay by usage or with credits) and schools (Edu plans with credits); other self-serve plans do not have access at launch. It uses up the usage included in the subscription 8 times faster than the standard mode, and purchased credits and Enterprise pay-as-you-go usage 6 times faster. On Pro, it draws on the included usage first and then on credits once that runs out. In companies, it is turned off by default and workspace owners have to turn it on, for selected users or for everyone.
Why it matters · analysis and opinion
Speed becomes something bought separately, just like model capability. For a team building agents that chain many calls together, the wait between one step and the next is what shows most, and that is where this mode can change the experience. For most uses, however, the extra cost outweighs the seconds saved, so the sensible move is to measure it on a real task before turning it on by default. In Codex, keeping it to the most expensive plan leaves almost every professional out for now, and in companies the decision rests with whoever administers the account, who would do well to set per-user spending limits before opening it up. Being able to keep data in Europe also makes it viable for those who have that requirement.
Official source: OpenAI · Written with the help of AI: how we make the news


