OpenAI's New Ultrafast Mode Runs GPT-5.6 Sol on Cerebras Chips, Not Nvidia
OpenAI just previewed a mode that pushes GPT-5.6 Sol to 750 tokens a second, and the chips doing the work aren't Nvidia's.
On August 13, OpenAI unveiled a limited preview of Ultrafast mode for GPT-5.6 Sol, its flagship model. It doesn't run on the Nvidia GPUs that power the rest of OpenAI's stack. It runs on Cerebras' wafer-scale inference chips. The result: up to 750 output tokens per second, roughly 14 times faster than GPT-5.6 Sol's standard processing speed of about 53 tokens per second, according to OpenAI's own announcement. Same model, same intelligence, just delivered at a pace that starts to feel like a live conversation instead of a typing indicator.
That's the headline. Here's why it matters more than a routine speed bump.
Every major AI lab has spent the last two years optimising for one thing: how smart the model is. Benchmarks, reasoning scores, context windows. Ultrafast is OpenAI saying, in public, that raw speed is now a selling point on its own, worth building a separate product tier around. Coding agents that wait on a response lose their edge. Voice assistants that lag feel broken. A model that thinks well but answers slowly loses to one that's merely good and instant. OpenAI is betting the market has started to notice the difference.
Ultrafast is preview-only for now, offered to a select group of customers with access expanding as capacity grows. OpenAI hasn't said what it will cost. Its existing Fast tier, which runs at roughly double the speed of standard processing, already charges a premium: $10 per million input tokens and $60 per million output tokens, against $5 and $30 for GPT-5.6 Sol at standard speed. If Ultrafast follows that pattern, real-time speed is going to cost real money.
The company behind the chips
Cerebras has spent a decade as the company everyone mentioned as an Nvidia alternative and almost nobody actually deployed at scale. That changed in January, when OpenAI signed a multiyear deal for up to 750 megawatts of Cerebras inference capacity running through 2028, a contract Reuters reported was worth more than $10 billion. Cerebras runs on wafer-scale chips, single silicon wafers roughly the size of a dinner plate, built specifically to move data faster between memory and compute than a rack of GPUs can manage. The company has claimed inference speeds up to 15 times faster than GPU-based setups.
The number kept climbing. When Cerebras reported its first quarter as a public company on June 23, it valued the OpenAI agreement at more than $20 billion, a single contract. That's roughly 23 times the midpoint of its full-year 2026 revenue guidance. Cerebras had gone public two months earlier in the largest semiconductor IPO on record, raising $6.4 billion. Its Q1 core revenue came in at $193.4 million, up 92% year over year. The company still posted a net loss, though, and its shares fell after the earnings release on investor worries about margins.
What it means for Nvidia
None of that spending mattered as a validation story until Ultrafast shipped. A $20 billion contract is a promise. A frontier lab actually routing its flagship model's traffic through your chips, in production, with a public feature built around the speed gain, is proof the promise is real. Cerebras has now got what Nvidia has had for years and no challenger has managed to take away: a marquee customer running its most important product on your silicon, by choice, not as a hedge.
It also puts pressure back on Nvidia, whose dominance in AI inference has rested partly on the assumption that nobody serious would bet a flagship product on unproven alternative hardware. OpenAI just did that, at least for one tier of one model. Whether Ultrafast stays a niche premium option or becomes how most GPT-5.6 Sol traffic gets served depends on pricing OpenAI hasn't announced and capacity Cerebras still has to build out through 2028.
For now, the fastest way to talk to GPT-5.6 Sol runs through Oklahoma and Texas data centers built by a company that IPO'd four months ago. That's not a small thing to say about the AI industry in August 2026.
Also read: AI Agent Approval Fatigue Is Quietly Undermining Startup Safety • California Approves Waymo's Biggest Robotaxi Expansion Across 18 Counties • How AI Agent Token Budgets and Rate Limits Actually Work in Production