Picture a claim for a three-day hospital stay. It leaves the provider’s billing system at 2 a.m., assembled by software that read the doctor’s notes, found four conditions the physician never explicitly coded, and attached them all. Ninety seconds later, on the other end, an insurer’s model scores that same claim, flags two of those codes as unsupported, and kicks it back. No human has looked at it yet. No human will, for weeks. Meanwhile the meter on that dispute is running, and somebody pays for the meter.
Blue Cross Blue Shield puts a number on it: $942 million in added costs over two years, which it attributes to hospital use of AI tools. Payers say AI-driven documentation and coding is pushing billing amounts up. Hospitals, in turn, are using AI to submit higher claims. The longstanding fight over who pays what has both sides now armed with automation, and the friction is showing up in medical expenses.
Why this one stings for tool reviewers
I spend my days testing AI tools and reporting what actually works. Most of the time, “it works” and “it’s good” point the same direction. A transcription tool that saves you four hours is a win for everybody. This is the rare case where a tool can work exactly as designed and still make the overall system worse.
Think about what a medical coding assistant is optimized for. Not truth. Not fairness. It’s optimized for capture — finding every billable detail buried in clinical notes that a rushed human coder would miss. By that metric, these tools are performing well. Revenue goes up. The vendor’s case study writes itself. The tool is doing its job.
And the insurer’s denial model? Also doing its job. It’s optimized to catch claims that don’t hold up. Higher denial rates, faster review, less leakage. Another clean case study.
Two products, both hitting their KPIs, and the joint outcome is a nine-figure cost increase plus a lot of paperwork nobody reads. That’s not a failure of either tool. It’s what happens when you deploy optimizers on opposite sides of a zero-sum negotiation and nobody owns the total.
The pattern shows up everywhere, just quieter
Healthcare billing is the loud version, because the dollars are big and the referee is absent. But the shape of this problem is familiar to anyone who reviews AI tooling:
- SEO tools that generate content to game ranking systems, facing ranking systems that use AI to detect generated content.
- Recruiting tools that pad résumés with keywords, meeting screening tools built to filter keyword-padded résumés.
- Outreach tools that write personalized-sounding email at volume, running into filters trained to spot personalized-sounding email at volume.
In every case, both tools deliver on their promise. Both vendors can show you the numbers. And the net effect is more volume, more noise, and more cost for everyone standing in the middle. The difference with healthcare is that the person in the middle is a patient, and the added cost eventually lands in premiums.
What I’d actually ask a vendor now
This story changed the questions on my evaluation list. When a tool operates in an adversarial setting — where somebody on the other side has a direct incentive to counter it — the usual demo metrics stop being useful. Volume processed, time saved, revenue lifted: all fine, all beside the point.
Better questions:
- What happens to my results when the other side deploys a comparable tool? Has the vendor modeled that, or are they selling first-mover numbers that expire?
- Is the output defensible by a human? If a coder can’t explain why a code was attached, you haven’t automated the work, you’ve automated the dispute.
- Does the tool optimize for accuracy or for capture? Vendors will insist those are the same thing. Ask for the cases where they diverge and see if anyone has an answer.
- What’s the cost of being wrong, and who absorbs it?
The uncomfortable part
I can’t tell you from the available facts who’s right in this fight. Insurers claiming AI raised their costs by $942 million have obvious reasons to publicize that number, and denial automation isn’t a neutral technology either. It’s entirely possible that some of those newly captured codes reflect care that genuinely happened and was previously under-billed. Under-coding is a real thing, and the same tool that inflates a claim can also correct one.
What I’m confident about is narrower: when two sides automate a dispute, the dispute gets bigger, faster, and more expensive before it gets resolved. That’s a predictable outcome, not a surprise. And it’s the part no product page mentions.
So if you’re evaluating AI tooling for anything that touches a negotiation, a claim, an application, or a ranking, assume the other side is shopping too. Then ask what your tool is worth in that world. The honest answer is usually less than the demo suggests.
🕒 Published: