OpenAI has unveiled a preview of Ultrafast, a new service tier for running GPT-5.6 Sol up to 14 times faster than Standard processing. The feature will launch first in the OpenAI API and uses Cerebras technology to reach up to 750 output tokens per second.
What changes when an advanced AI stops taking several seconds and can respond almost at the pace of a conversation? According to OpenAI, that speed can bring more capable models to tasks where every moment affects the outcome.
More speed without choosing a less capable model
Until now, companies that needed real-time responses often turned to smaller or specialized models. Ultrafast aims to change that decision: keep the intelligence of GPT-5.6 Sol while drastically reducing response time.
This does not mean that every AI process will become instant or that capacity limits will disappear. It is an option focused on interactive products and workflows, where waiting can affect an operation, a purchase, or a conversation with a customer.
Ultrafast’s central idea is to produce more useful work per second, not just respond faster.
Where it can make an impact
OpenAI is testing this mode with companies in software development, commerce, financial research, customer support, and other interactive applications. The identified scenarios include:
- Incident response: analyze logs, traces, recent code changes, and engineering reports while an outage is still happening.
- Financial research and security: review market signals, evaluate transactions, and detect suspicious activity while conditions change.
- Customer support and voice: solve complex problems during a conversation, even when the response requires checking several systems.
- E-commerce: answer questions about products, check inventory, personalize recommendations, and solve issues during checkout.
- Research and experimentation: turn processes that once required an overnight run into interactive sessions with multiple iterations throughout the day.
The difference may seem small in a single response. But in a workflow with hundreds of queries, checks, and decisions, a few seconds less can completely transform the experience.
An example: investigating a system outage
OpenAI explains that its own teams are testing Ultrafast for incident response. When an alert appears, engineers need to gather information quickly while the data is still changing.
With GPT-5.6 Sol in Ultrafast mode, they can read logs, analyze traces, summarize conversations, identify the next checks, and help prepare or validate a solution. The model reduces the time between observing a signal, testing a hypothesis, and deciding what to do next.
That said, AI does not replace the technical team’s responsibility. Engineers still make the important decisions and control the implementation of any changes.
Research with shorter cycles
Another use case is related to research. A team can search for information, query data, organize results, and summarize findings across several connected tools.
In a traditional workflow, a group might launch experiments overnight and review the results the next day. With greater speed, the same team can test an idea, analyze what happened, adjust its approach, and run another test during the workday.
What is the real benefit? It is not simply getting an answer sooner. It is being able to maintain the pace of thought and reduce the pauses that often interrupt an investigation.
A partnership with Cerebras
Ultrafast represents the next step in the collaboration between OpenAI and Cerebras to deliver very low-latency inference. In simple terms, Cerebras provides the infrastructure needed for GPT-5.6 Sol to generate responses at speeds of up to 750 tokens per second.
A token can be a word, part of a word, or a symbol, depending on the language and how the text is processed. That is why the figure does not correspond exactly to 750 words per second, but it does help put the announced speed into perspective.
During the preview phase, OpenAI will work with an initial group of customers to identify which tasks benefit most from this improvement. Access will gradually expand as capacity increases and results are analyzed in real-world environments.
The bet is clear: if speed no longer forces you to sacrifice intelligence, AI can become part of much more time-sensitive moments, from a support conversation to a response to a critical failure. The challenge will be proving that this speed also translates into accuracy, safety, and better-informed decisions.
