OpenAI launched a limited preview of Ultrafast, a new processing mode for its GPT-5.6 Sol model that generates up to 750 output tokens per second, 14 times faster than the model's standard mode, the company said Thursday.
The speed comes from Cerebras, whose wafer-scale chips keep model weights in on-chip SRAM instead of shuttling them between memory layers, according to Cerebras. Each wafer-sized chip holds 44 gigabytes of SRAM, the company said.
Cerebras said Ultrafast mode runs 11 times faster than Claude Fable 5 and five times faster than Claude Opus 4.8 running in its own fast mode. On the Humanity's Last Exam benchmark, Sol Ultrafast finished 2,500 PhD-level questions in 11 hours 11 minutes, versus 78 hours 27 minutes for Claude Fable 5, Cerebras said. On GDP-Val, a benchmark built around economically valuable tasks, the company said Ultrafast delivered a 5.6-times speedup with no drop in quality.
"With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate," OpenAI's Rohan Varma said, according to Cerebras. OpenAI said the mode targets tasks that need real-time response, including incident response, customer service, financial market analysis and e-commerce workflows, without forcing developers to fall back on a smaller, less capable model. Access is limited to a small group of customers for now, with OpenAI saying it will expand as capacity grows.
For builders shipping agents that react to live data, support queues, trading signals, monitoring alerts, sub-second responses at frontier-model quality change what's worth attempting without standing up a dedicated inference cluster. The catch is access: this is a preview for a small customer list, not something most developers can test yet.