Eleven days. That’s the median gap between major AI model releases in 2026, down from 37.5 days in 2023. Three and a half years ago you had over a month to get to know a model before the next one showed up. Now you get a week and a half, and half of that is spent reading the changelog.
CNBC put a name to what a lot of us have been feeling: model fatigue. AI labs are burning out on their own release cycles. Engineering teams are struggling. Investors are struggling. Customers are struggling. And some labs are now floating the idea of slowing down, which is a remarkable admission from an industry that spent three years treating speed as the entire strategy.
I review AI tools for a living. Let me tell you what an eleven-day cycle actually does to that job.
Testing takes longer than shipping
When I evaluate a model for this site, I’m not asking whether it can write a haiku. I’m running it against the same set of tasks I’ve been using for a year, checking whether it holds up under long context, seeing where it fails on tool calls, figuring out if the failure modes are the kind you can design around or the kind that quietly poison your output three steps downstream. That’s not a one-afternoon job. Done properly, it’s a week or two of real work, plus time spent living with the thing in an actual workflow.
Eleven days means I finish an honest evaluation roughly when it becomes obsolete. So does everyone else doing this work seriously. What fills the gap instead is vibes-based coverage: someone posts a screenshot of a clever output, it gets ten thousand reposts, and that becomes the received wisdom about a model nobody has actually stress-tested. The review ecosystem has been quietly replaced by a reaction ecosystem, and I don’t think most readers noticed the swap.
What this costs you
The fatigue story is being framed as a labor problem inside AI labs, and it is one. But the downstream cost lands on people building things.
- Your prompts rot. Every prompt you tuned to a specific model’s quirks is a small piece of technical debt with an expiration date you don’t control.
- Your evals go stale. If you built a solid internal test suite, congratulations, you now maintain it against a moving target.
- Your team stops trusting upgrades. I’ve talked to developers who pin an older model and refuse to move, not because it’s better, but because they can no longer afford to re-validate.
- Regression becomes invisible. When releases outpace testing, nobody has a clean baseline. A new model can be worse at something you rely on, and you’ll find out from a customer.
That last one is the part that keeps me skeptical. Speed hides regressions. Not maliciously, just structurally. If the community can’t finish evaluating version N before N+1 lands, then nobody ever builds the historical record that would let you say “this got worse in March.”
Slowing down is the interesting signal
Politico reported that AI labs want to slow down risky model testing and raised the question of whether it’s already too late. Set the safety framing aside for a second and look at the business logic, because it’s the same conclusion from a different direction: the release treadmill has stopped producing returns proportional to its cost.
You can see why. When a model shipped every five weeks, each release got a news cycle, a round of independent testing, and time to accumulate a reputation. At eleven days, releases blur. Users can’t tell them apart. The marginal attention value of shipping fast has collapsed, and what’s left is the exhaustion.
I’d like to believe the industry is capable of stepping off voluntarily. I’m not confident. Release cadence is competitive signaling as much as it is engineering, and nobody wants to be the lab that looks like it stalled. Slowing down only works if several labs do it at once, which is a coordination problem the AI space has not historically been good at solving.
What I’d actually do about it
Practical advice, from someone whose job this breaks:
- Build your own eval set, small and specific to your use case. Twenty tasks you care about beats any public leaderboard.
- Pin your model version in production and upgrade deliberately, on your schedule.
- Treat any review published within 48 hours of a release, including mine, as a first impression rather than a verdict.
- Keep old outputs. They’re your only baseline when a new version quietly changes behavior.
The honest position is that nobody currently has enough time to know how good these models are, including the people making them. Model fatigue isn’t just tiredness. It’s the sound of an evaluation process falling behind a production process, and the gap is where bad decisions live.
🕒 Published: