OpenAI has released a new model called Ultrafast, which plays a higher model response speed. According to the company, this model allows GPT 5.6 Sol to run 14 times the standard processing speed, with output speeds up to 750 tokens per second, which is now open to a small number of customers in the form of previews.
This update aims at the business needs for real-time AI processing. OpenAI indicates that, in the past, if a response is to be closer to real time, there is usually a need to change to a smaller model or to select a system that is optimized for a given task. Ultrafast is directed towards increasing the effective workload that can be performed within a unit time, without relying solely on small models.
Output Speed Up
According to OpenAI, the core selling point for Ultrafast is the significant compression generation of waiting time.
- Maximum output speed of 750 tokens per second
- Processing speed is about 14 times the standard mode
- The current version is still in the preview
TechCrunch mentioned that other competitors such as Anthropic had previously introduced accelerated models. In Claude, for example, the product also provides a fast-track model, but OpenAI has a higher speed indicator this time.
Business-oriented flows
OpenAI directs Ultrafast applications to a wide range of business tasks, focusing on incident response, customer services and support, financial market analysis, and electronics operations. Such scenarios are usually more sensitive to delays, and model response times have a direct impact on manual collaboration efficiency and automated process performance.
From a product orientation point of view, Ultrafast is not a stand-alone new model, but a high-speed run around GPT 5.6 Sol. For business clients, this means that while retaining stronger model capabilities, attempts are made to embed AI more directly into real-time business processes.
Supported by Cerebras
OpenAI states that Ultrafast is supported by it in cooperation with the chip company Cerebras. The first preview eligibility is currently open to only a small number of clients and will be progressively expanded as the computing capacity increases.
This also shows that large-scale model competition is extending from a equation of parameters and capabilities to response speed, deployment efficiency and bottom-level computational synergy. For AI products oriented to the business market, speed is becoming a key indicator alongside model quality.
