OpenAI has launched a preview of Ultrafast mode for its GPT-5.6 Sol model, delivering up to 750 output tokens per second. This new tier combines the speed of smaller models with the full reasoning capabilities of the flagship system to increase useful work per second. The company cites incident response as a primary use case, allowing engineers to analyse logs and code changes in real time during an outage. Similar applications include evaluating market signals in finance, resolving complex customer support queries instantly, and personalising e-commerce recommendations before a purchase is abandoned. Research teams can also run experiments interactively rather than waiting for overnight batch jobs to complete.
The move signals a shift in how OpenAI monetises inference speed, creating a third pricing tier alongside its existing Fast Mode. This structure mirrors cloud providers like AWS, which charge more for higher performance levels. If speed becomes a bottleneck across industries, this model allows OpenAI to capture a direct share of the revenue gains generated by faster processing. The Ultrafast tier is likely pricier than the current Fast Mode, which already offers up to 2.5x speed at roughly double the standard API cost.
* Ultrafast mode is currently available as a preview
* GPT-5.6 Sol is the underlying model for the new tier
* The pricing structure adds a third, faster option to the API




