Anthropic shipped a frontier model roughly every 46 days in the first half of 2026. In the second half, that gap closed to about 26 days. OpenAI’s cadence has gone up too. Meanwhile, I still have a testing checklist that takes me about three weeks to run properly on a single model.
Those two facts do not fit together, and that mismatch is the whole story of what reviewing AI tools feels like right now.
The math nobody wants to do
If a model drops every 26 days and it takes three weeks to evaluate one honestly, you are never evaluating the current model. You are always publishing a review of something that has already been superseded, using benchmarks that were designed for the version before that. By the time I finish stress-testing a model’s long-context behavior, its replacement is in preview.
The obvious response is to test faster. Run the quick benchmarks, post the numbers, move on. That is what most of the coverage does, and it is why so much of it is useless. Fast tests measure whether a model can answer a question. Slow tests measure whether it stays reliable at 2 a.m. on day nineteen of a project when your prompt has drifted and your context window is stuffed with legacy garbage. Those are different questions, and only one of them matters to anyone actually shipping software.
Not every release is a release
Part of why the pace feels frantic is that we have stopped distinguishing between types of updates. Version numbering used to carry information. A major jump like GPT-3 to GPT-4 or Claude 2 to Claude 3 signaled real capability changes and, usually, migration work on your end. Point releases signaled refinement.
That signal is getting noisy. A lot of what gets announced now is repackaging of advances that already existed, bundled with efficiency improvements and a new name. The genuinely useful trend underneath the noise is cost and efficiency: new models delivering better performance for less money. That is real, and it is the part I care most about as a reviewer, because it changes what you can afford to build. But it does not require the same evaluation cycle as a capability jump. A cheaper model that behaves like the old one is a procurement decision. A smarter model that behaves differently is an engineering decision.
Google made this even clearer by pushing updates across models, research tools, search features, and development platforms in a single week. That is not a release. That is a weather event.
What I actually changed about how I test
I gave up on trying to keep pace with every launch. Here is what replaced it:
- A fixed task suite that never changes. Same prompts, same repos, same failure cases, every model. It is boring and that is the point. Comparability beats novelty.
- Cost-per-completed-task as the headline number, not tokens per dollar. Efficiency claims mean nothing if the cheap model needs four attempts.
- A two-week wait before writing anything. Launch-day behavior and week-three behavior are frequently different. Rate limits settle. Quirks surface. Someone finds the thing that breaks it.
- Ignoring point releases unless something in my suite moves. If the numbers do not budge, there is nothing to tell you.
Whether it is still getting better
The honest answer is that the direction looks positive and the framing is doing a lot of work. There is a reasonable case that if the current pace of flagship improvement holds, 2026 gets remembered as the year models stopped feeling like a call center associate and started resembling a research assistant. In the pharmaceutical space, specialized models analyzing chemical databases alongside other data types are compressing discovery timelines from years to months, which is the kind of concrete result that is hard to argue with.
But “if the pattern holds” is a conditional, and release frequency is not evidence that it does. Shipping more often is a business decision. Getting better is a technical one. They can move independently, and right now the first is much easier to observe than the second.
What this means for you
Stop treating every announcement as a call to action. Pick your model, build your own small evaluation set against the work you actually do, and re-run it quarterly instead of every time your feed lights up. Switch when your numbers say to switch, not when a blog post says to.
The cadence is not slowing down. Anthropic went from 46 days to 26 in a single year, and nothing about the competitive pressure suggests a reversal. The only variable you control is how much of your attention that consumes. Mine is now capped, deliberately, and my reviews got more useful the moment I stopped chasing.
đź•’ Published: