Accelerating GPT-5.6 Sol Ultrafast
AI builders have long faced a tradeoff between speed and intelligence: larger models incur higher computational and data movement costs, slowing responses. Cerebras and OpenAI are addressing this with Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. It is initially available to a select group of customers, with access expanding over time.
On Ultrafast, GPT-5.6 Sol delivers up to 750 output tokens per second with no quality compromise. Compared with output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast runs 11x faster than Claude Fable 5 and 5x faster than Opus 4.8 on Fast mode. In Cerebras's head-to-head test on Humanity's Last Exam (HLE)—a benchmark of 2,500 questions typically answerable only by PhDs in fields like chemistry, economics, and literature—Ultrafast answered all questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes to arrive at the same conclusions, meaning Ultrafast achieved comparable accuracy nearly 7x faster, completing frontier human knowledge in a single working day. Benchmarking used GPT 5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10, and Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15.
On GDP-Val, a benchmark for economically valuable knowledge work, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, accelerating tasks like legal briefs, financial models, and engineering reports. Cerebras performed this benchmark on July 31, 2026, using GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium reasoning within Codex.
Faster inference changes what's possible for high-stakes work, allowing agents to operate on the critical path of time-sensitive problems. As Rohan Varma, Product at OpenAI, put it: "With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We're excited to see how workflows and applications are transformed by Ultrafast inference." Ultrafast represents a persistent edge for organizations seeking both frontier intelligence and minimal latency.