OpenAI's new Ultrafast mode for GPT-5.6 Sol generates 750 output tokens per second, a 14-fold jump that removes the latency barrier that has kept frontier models out of real-time enterprise workflows.
OpenAI's Ultrafast mode for GPT-5.6 Sol generates 750 output tokens per second, a 14-fold speedup that lets enterprises run its most capable model in real-time applications where latency previously forced them onto smaller, specialized systems. The feature, announced Aug. 13, is in limited preview for select API customers and marks the first time OpenAI has paired its flagship model with a non-GPU inference architecture at scale.
"Speed and intelligence are no longer mutually exclusive in AI development," Andrew Feldman, chief executive of Cerebras Systems, said.
The mode is powered by Cerebras' Wafer-Scale Engine, which keeps all 44GB of SRAM on-chip and bypasses the memory-bandwidth bottleneck that slows large-model inference on GPU-based systems. OpenAI's Sachin Katti, vice president of compute strategy and GPT-infrastructure, said early customer testing will reveal where the speed creates the most value. OpenAI engineers described the effect in a preview video: Jeff Liu, a member of the technical staff, said the pace felt like "cheating," while Sayan Sisodiya said a code-base refactor "costs almost nothing and happens almost instantly."
The speed-intelligence combination could expand the addressable market for frontier AI into time-sensitive sectors such as financial market analysis, automated incident response, and high-volume customer service, where every second of delay carries a measurable cost. OpenAI cited benchmark results showing GPT-5.6 Sol completes "Humanity's Last Exam" nearly seven times faster than Anthropic's Claude Fable 5, and delivers a 5.6x speedup on economically valuable tasks measured by the GDP-Val benchmark.
The Hardware Shift Behind the Speed
Cerebras' architectural advantage is structural rather than incremental. Traditional GPU systems shuttle model weights between on-chip memory and off-chip storage, a constraint that caps inference speed for large models. The Wafer-Scale Engine eliminates that transfer by holding the entire model in on-chip SRAM, which is why the 14x gain comes without a reduction in model quality. The trade-off is operational: OpenAI's ability to scale Ultrafast depends on consistent access to Cerebras' specialized hardware, and the company said it will expand availability gradually as infrastructure capacity grows.
The partnership signals a potential shift in the AI hardware landscape, where specialized inference architectures are challenging the GPU-centric model that Nvidia has dominated. Anthropic already offers a fast mode for Claude, and Google launched Gemini 3.7 Flash this week for coding and autonomous workflows, underscoring that low-latency inference has become a competitive battleground across the industry.
What It Means for Investors
OpenAI remains a private company, though it has filed confidential IPO papers with US regulators, a development global investors are tracking. The company raised $100 billion in 2026 at a post-money valuation of $850 billion, according to StartupHub.ai data, and holds a leadership score of 84 out of 100 on the platform's ranking. For listed names, the ripple effects run through the semiconductor and cloud complex: Cerebras trades on Nasdaq under CBRS, while Nvidia's data-center franchise faces a new benchmark for inference speed that could pressure its pricing power in the low-latency segment. The next monitorable is the timeline for a wider public rollout of Ultrafast and any update on OpenAI's path toward a listing.
This article is for informational purposes only and does not constitute investment advice.