Remember when everyone was streaming tokens at reading speed and calling it “real-time AI”? You’d fire off a prompt, watch the words trickle out like a nervous typist, and convince yourself this was fine because, hey, at least you could see it working. That era might be ending. On August 13, 2026, OpenAI previewed Ultrafast mode, a new service tier that runs GPT-5.6 Sol at up to 14x the speed of standard processing — up to 750 tokens per second, powered by Cerebras hardware.
I review AI toolkits for a living, and my honest first reaction was equal parts excitement and skepticism. So let me walk you through what we actually know, what it might mean for builders, and where I’m reserving judgment until I get hands-on access.
What Ultrafast Actually Is
The pitch is simple. GPT-5.6 Sol is the most capable model in the GPT-5.6 family, and Ultrafast is a service tier that runs it dramatically faster than the standard offering. OpenAI says up to 14x faster, topping out around 750 output tokens per second. The speed comes from running the model on Cerebras hardware rather than conventional infrastructure.
Right now, it’s a limited preview. It’s launching first in the API to a select group of customers, which means most of us — myself included — are reading about it rather than testing it. That matters, and I’ll come back to it.
Why Speed Is the Feature Now
For the past couple of years, the story in AI has mostly been about capability. Smarter models, better reasoning, fewer hallucinations. Speed was the thing you traded away to get quality. Want the top-tier model? Fine, but you wait for it.
Ultrafast flips that trade-off, at least on paper. You’re not getting a smaller, faster, dumber model. You’re getting the flagship — Sol — at speeds that make entirely different product categories viable. OpenAI’s stated aim is enhancing real-time AI applications, and that framing makes sense. At 750 tokens per second, the model finishes responding faster than most people finish reading the first sentence.
Think about what that enables:
- Voice interfaces where latency is the difference between a conversation and an awkward walkie-talkie exchange
- Agent workflows that chain dozens of model calls, where per-call delays compound into minutes of dead time
- Live coding assistants that can regenerate entire files before you’ve reached for your coffee
- Customer-facing products where users simply won’t tolerate a spinner
If you’ve ever built an agent pipeline, you know the pain. Ten sequential model calls at standard speed is a coffee break. Ten calls at 14x speed starts to feel like software.
The Cerebras Angle
The hardware detail is worth pausing on. This isn’t OpenAI squeezing more out of the same infrastructure — it’s a different chip vendor entirely. Cerebras has been building specialized AI hardware for years, and getting the flagship OpenAI model running on it at these speeds is a notable endorsement. For the broader ecosystem, it signals that inference hardware competition is heating up, and that’s good news for anyone who pays API bills.
My Honest Reservations
Now the reviewer part of my brain kicks in, because “up to 14x” is doing some load-bearing work in that announcement.
“Up to” is marketing’s favorite phrase for a reason. What’s the typical speed, not the peak? Does quality hold at full velocity, or are there quiet compromises? What does pricing look like? None of that is public yet, and until it is, this is a preview in every sense — including the sense where you can’t verify the claims yourself.
The limited-access rollout also means the early signal will come from select customers, who tend to be large partners with incentives to say nice things. I’d rather see what happens when a thousand indie developers throw messy real-world workloads at it.
Should You Care Yet?
Yes — but with your wallet still in your pocket. If you’re building anything latency-sensitive, Ultrafast is worth tracking closely, because if the numbers hold up in practice, it changes what you can ship. If you’re running batch workloads where speed doesn’t matter, this changes nothing for you today.
My plan is simple. The moment access opens up, I’m putting it through real workloads — agent chains, long generations, sustained throughput — and reporting what I find, marketing numbers be damned. Fourteen times faster is a big claim. I genuinely hope it survives contact with reality, because a flagship model at 750 tokens per second isn’t an incremental upgrade. It’s a different kind of tool.
Until then, I’ll be refreshing my inbox for that preview invite.
đź•’ Published: