Quick test. Can you name, without looking, which Gemini model you used last week? Was it Flash? Pro? Was there a number attached, and did that number mean newer or just different? If you hesitated, you’ve stumbled into something TechCrunch put a name to this week: Google’s Gemini has a branding problem, and so does the rest of AI.
I review toolkits for a living. I spend my days installing things, reading changelogs, and figuring out which version of which model actually powers the thing I just paid for. And I want to be blunabout something the industry keeps pretending is cosmetic: naming confusion is not a marketing footnote. It’s a usability defect.
Why This Lands Differently For Reviewers
When I test a tool, the first thing I need to establish is what’s under the hood. That should take ten seconds. Instead it often takes twenty minutes of digging through docs, release notes, and a pricing page that lists three tiers with names that sound like sedan trim levels.
The practical consequences pile up fast:
- You can’t compare two tools if you can’t tell whether they’re running the same model underneath.
- You can’t reproduce a result if the model name silently shifted between your first test and your second.
- You can’t budget properly when tier names don’t map cleanly to cost or capability.
- You can’t file a useful bug report when you’re unsure what you were actually talking to.
None of that is a branding inconvenience. That’s an evaluation problem, and it hits every person trying to make an honest purchasing decision.
Google Is Not Uniquely Guilty
The reporting frames this as Gemini’s issue first and the industry’s issue second, which feels about right. Google gets top billing because it’s Google and because the Gemini name has been stretched across a lot of surfaces. But the pattern is everywhere. Every major lab has shipped some combination of size labels, speed labels, generation numbers, and codenames, often mixed into the same product family with no clear hierarchy.
Compare that to how we name almost anything else in software. Version numbers have a grammar. Higher is newer. Semantic versioning tells you whether something broke. Even phone naming, which nobody would call elegant, at least trends in one direction. TechCrunch also reviewed the Pixel 11 Pro XL this week and called it an iterative upgrade with snappier cameras. That’s a clunky name, but I know exactly where it sits relative to a Pixel 10.
What The Confusion Actually Costs
My honest read is that vague naming benefits vendors more than users, whether or not that’s intentional. If nobody can pin down which model is which, nobody can hold a specific version accountable. Regressions get absorbed into the fog. Comparisons get harder to run and harder to publish. A tool can quietly swap to a cheaper backend and most users will never notice, because they never had a firm grip on what they were using before.
I don’t think there’s a conspiracy here. I think there are fast-moving teams, competing internal roadmaps, and marketing departments trying to make incremental releases feel significant. The result is the same either way: a naming system that fails the person trying to use it.
The Adjacent Story Worth Watching
The other Google item this week is that publishers are getting a new way to fight AI-driven traffic losses. I mention it because it belongs to the same family of problems. Both stories are about clarity, or the lack of it. Publishers want to know where their content goes and what it feeds. Users want to know which model answered their question. Different stakeholders, same underlying gap between what these systems do and what anyone outside the company can see.
What I Want From Vendors
My asks are small and boring, which is usually a sign they’re the right ones:
- Pick a naming direction and hold it. Bigger numbers newer, consistently.
- Expose the exact model identifier in the product, not just in the API docs.
- Announce backend swaps in the changelog, in plain language.
- Stop reusing a family name for products that behave nothing alike.
My Take
I’ll keep doing the archaeology, because that’s the job. But you shouldn’t have to. When a reviewer with a spreadsheet and a lot of patience still can’t confidently state what’s running in a given tool, the naming has failed at the one thing naming exists to do.
The industry has gotten remarkably good at building capable models and remarkably bad at telling you which one you’ve got. Fixing that costs almost nothing and would improve every downstream decision users make. Until someone does, treat model names as a hint rather than a fact, and verify before you commit budget to anything.
🕒 Published: