Anthropic is growing fast. Anthropic’s best model can’t find enough people who want to use it. Both of those things are true at once, according to a Financial Times report that Gary Marcus flagged this week, and the gap between them is the most interesting thing to happen in AI tooling in months.
I review tools for a living. I spend my weeks swapping models in and out of the same workflows, watching where the money goes and where the frustration builds. So when I read that the flagship model’s adoption stays low while cheaper options keep gaining ground, my reaction was not surprise. It was recognition. That is exactly what I see in my own setup, and probably in yours.
Capability is not the product
The unstated assumption behind frontier model pricing is that people will pay a premium for the smartest thing available. It is a clean story and investors love it. It also does not match how most of us actually build.
Most real work is not hard. It is repetitive. Reformatting JSON. Writing the fortieth test file. Summarizing a meeting transcript. Renaming variables across a repo. None of that needs a model with a PhD-level reasoning budget. It needs a model that is fast, cheap, and does not fall over. When a task is genuinely difficult, sure, reach for the expensive option. But the difficult tasks are a small slice of the day, and you notice the price of the cheap model every single time you use it because you use it constantly.
That is the arithmetic that the FT report is describing in aggregate. Not that the top model is bad. That the top model is priced and positioned for a use case that turns out to be narrower than anyone modeled.
What this looks like from the toolkit side
Here is the pattern I keep running into when I test stacks:
- Teams start on the flagship model because it is the safe choice and the benchmarks look great.
- Someone looks at the monthly bill.
- They route the boring 80 percent to a cheaper model and keep the expensive one for the hard cases.
- Nothing breaks. Quality drops slightly on a few tasks. Nobody complains.
- Six months later, the cheap model handles more than they planned, because it got better in the meantime.
Step five is the one that should worry Anthropic and everyone in the same position. Cheaper models are not a fixed floor. They keep climbing. Every month the expensive tier has to justify a premium against a rival that just got closer for free.
The IPO problem hiding underneath
Marcus points at the thing that makes this more than a pricing footnote: public offerings are coming for the big AI labs. The pitch those companies need to make is that superior capability converts into pricing power and durable revenue. The FT story cuts directly against that. Strong company growth with weak flagship adoption suggests the growth is coming from volume and breadth, not from a premium product that customers cannot live without.
Those are very different businesses. One is a specialty vendor with a moat. The other is infrastructure competing on cost, which is a fine business but valued very differently.
The counterargument worth taking seriously
I am not going to pretend the frontier work is pointless. Anthropic has said that as of May 2026, more than 80 percent of the code merged into its own codebase was written by Claude, up from a lower figure earlier. That is a real signal about what these models can do when they are pointed at complex, high-context work by people who know how to point them.
But notice what it also says: the most convincing demonstration of the top model’s value is internal. The company that built it has the expertise, the tooling, and the incentive to extract every bit of value from it. Most buyers have none of those three. They have a budget, a deadline, and a task that a mid-tier model handles for a fraction of the cost.
What I would actually do
If you are picking tools right now, the news changes very little about your practical strategy, and that is sort of the point. Route by task difficulty, not by brand loyalty. Benchmark the cheap option on your own workload before you assume it fails. Keep the expensive model available for the cases that genuinely need it, and measure how often that actually happens. My guess is that the number will surprise you in the same direction it surprised the market.
The broader lesson is one that tool reviewers have been muttering for a while. Being the best model is not the same as being the right model. Users are voting with their invoices, and they are picking good enough at a price they can defend to their finance team. Anthropic built something remarkable. The market is answering a different question than the one the benchmarks were designed to ask.
🕒 Published: