The price hike got cancelled.
That is the part of the Claude Sonnet 5 story I keep coming back to, and it is the part almost nobody led with. Anthropic shipped the model on June 30, 2026, called it its most agentic Sonnet to date, and put it out at $2 per million input tokens and $10 per million output tokens. The catch was a stated expiration date. Introductory pricing ran through August 31, after which it was set to move to $3 per million input and $15 per million output.
Then on August 10, three weeks early, Anthropic made the introductory numbers permanent. No step up. If you had been budgeting for $3/$15 in September, you got a 33% discount handed back to you without doing anything.
One quick housekeeping note before I go further. You will see this release floating around as “Sonnet 5.5” in headlines and feeds. What is documented is Claude Sonnet 5, shipped June 30. I am going to stick to what is actually on the record, because guessing at version numbers is how review sites end up quoting each other’s mistakes for a year.
What Anthropic says it built
The pitch is autonomy. Anthropic positions Sonnet 5 as designed to plan multi-step tasks and operate tools with reduced human supervision. That phrasing is doing a lot of work, and it is the claim I would test hardest before restructuring any workflow around it.
“Reduced human supervision” is not a benchmark. It is a product decision about how much rope you are willing to hand a model before you check on it. In my experience the gap between a model that can plan a multi-step task and a model you actually trust to run one unattended is measured in how expensive the failure is. A model that drafts a seven-step refactor plan and executes six of them correctly is genuinely useful. A model that does the same thing against production infrastructure is a Tuesday you will remember.
So the honest framing for a toolkit buyer is this: the agentic claim tells you what the model is aimed at, not how much oversight you get to drop. Find that out yourself, on work you can roll back.
The context window is a billing question
Sonnet 5 carries a 1M-token context window. Large context is great for the obvious reasons, and it also quietly changes your cost profile.
Run the arithmetic. At $2 per million input tokens, a single request that genuinely fills a 1M-token context costs roughly two dollars on the input side alone, before the model writes a word back. That is fine occasionally. It is not fine as a default habit in an agent loop that re-sends accumulated state on every turn.
The models that feel cheap on paper get expensive through usage patterns, not price sheets. Big context windows encourage the lazy pattern of stuffing everything in rather than retrieving what matters. If you are moving an agent onto Sonnet 5, instrument your token spend per task before you scale anything up.
The silent default swap
On July 1, one day after launch, Sonnet 5 became the default model for all Free and Pro users, replacing Sonnet 4.6.
If you are building on top of Claude through the consumer product rather than the API, that is a change in your tooling that you did not schedule and probably did not notice. Prompts tuned against 4.6 behavior were suddenly running against a different model with different tendencies. That is not a knock on Anthropic specifically, every vendor does this, but it is an argument for pinning model versions anywhere the output feeds something that matters.
Worth pairing with an earlier move: on April 26, 2026, Anthropic raised Sonnet and Haiku rate limits at every usage tier and simplified down to three tiers named Start, Build, and Scale. Higher limits plus a cheaper mid-tier model plus an agentic pitch is a consistent strategy. They want Sonnet running long jobs, not answering one-off questions.
What I would actually do with this
- Treat the permanent $2/$10 pricing as the real headline. Predictable cost beats a marginal benchmark win when you are budgeting a product.
- Note that Sonnet 5 launched priced below Anthropic’s flagship Opus model, which makes it the sensible default and Opus the escalation path, not the other way around.
- Pin your model version if you are on Free or Pro and your prompts are load-bearing.
- Log tokens per completed task, not tokens per call. That is the number that tells you whether agentic autonomy is saving money or burning it.
- Give it real supervised runs before you reduce supervision. The claim is about design intent. Your trust should be about observed behavior.
A mid-tier model that got cheaper than promised and stayed there is a good outcome. Just do not let the autonomy language talk you out of watching it work.
đź•’ Published: