OpenAI has opened a preview of Ultrafast, an API tier that runs GPT-5.6 Sol at up to 750 output tokens per second, roughly 14 times the speed of standard processing. The model is not a smaller one. Cerebras, which announced on August 13 that its hardware powers the tier, says it is the same GPT-5.6 Sol running on a different machine.
That machine is Cerebras’s Wafer-Scale Engine, a processor the size of a dinner plate carrying 44 gigabytes of on-chip SRAM. Enough memory sits beside the compute that model weights stay resident on the chip, which removes the trip to external memory that sets the pace of GPU inference.
Cerebras published two comparisons. On Humanity’s Last Exam, a 2,500-question set at graduate level, GPT-5.6 Sol on Ultrafast finished in 11 hours and 11 minutes, against 78 hours and 27 minutes for Claude Fable 5, with what Cerebras described as comparable accuracy. On GDP-Val, which covers economically valuable knowledge work, it reported a 5.6 times end-to-end speedup over standard processing and no drop in quality. The tests were run in July, with GPT-5.6 Sol driven through Codex. Both sets of figures are Cerebras’s own.
“By combining GPT-5.6 Sol with Cerebras’ inference technology, we’re exploring what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency,” said Sachin Katti, OpenAI’s vice president of compute strategy and GPT infrastructure. Andrew Feldman, Cerebras’s chief executive and co-founder, said speed and intelligence are no longer mutually exclusive.
Access is limited to selected customers while capacity grows. OpenAI has not announced pricing for the tier or a date for general availability.
Sol’s own price did not move in the last round of cuts. On July 30 OpenAI took 80 percent off GPT-5.6 Luna and 20 percent off Terra, crediting rewritten GPU kernels and better speculative decoding, and left Sol at $5 and $30 per million tokens.
The floor under it keeps shifting. DeepSeek launched V4-Pro and split its API into peak and off-peak billing, with increases reported as high as 1,100 percent. Google released Gemini 3.7 Flash on August 13 at an introductory 75 cents per million input tokens, a rate that runs to December 31 and then doubles. Anthropic’s Sonnet 5 has been selling at an introductory rate that expires on August 31.
Ultrafast also puts a frontier workload on silicon that is neither Nvidia’s nor OpenAI’s own. OpenAI has committed to $750 billion in compute spending through 2030, and Nvidia has been assembling financing to keep its customers buying, including a $500 billion facility secured against the chips themselves. Anthropic has gone further upstream, talking to Samsung about a custom accelerator of its own.
Cerebras is taking signups for wider access.
Sources: Cerebras, GlobeNewswire, OpenAI
–
By the Control Plane Editorial Team