It’s 9:40 on a Tuesday morning and a coder at a mid-sized hospital has 60 charts to clear before lunch. She doesn’t read them line by line anymore. A tool ingests the physician’s notes, surfaces every condition that could plausibly be documented, and suggests a higher-complexity code with a confidence score attached. She clicks accept. Somewhere across the country, an insurer’s model receives that claim, flags the complexity jump, and kicks back a denial in under a second. Two systems, neither of which examined a patient, are now negotiating what that patient’s care was worth.
That’s the scenario behind a number that should make anyone who reviews AI tools for a living stop and squint. The Blue Cross Blue Shield Association ran an analysis and concluded that hospitals’ use of AI during insurance claims submission added $942 million in healthcare spending over two years. The New York Times has been covering the same dynamic, framing it as an escalation of a long-running feud between hospitals and insurers, with both sides now armed with software.
What the number actually describes
I want to be careful here, because the claim is coming from one side of a fight. Insurers have an obvious interest in painting hospital automation as the villain, and they are running their own models on the denial side. The $942 million figure is their analysis of their own claims data. Treat it as a directional signal, not a physical constant.
Even discounted, the mechanism is the interesting part. The reported pattern isn’t that AI is making care more expensive. It’s that AI is making billed complexity go up without the underlying care changing. Same visit, same patient, same treatment, richer documentation, bigger code. That’s a very specific failure mode, and it’s one I keep running into when evaluating tools in adjacent categories.
The optimization trap
Ask a tool to maximize a metric and it will maximize that metric. Revenue cycle products are typically sold on a single promise: capture the revenue you’re currently leaving on the table. That’s a legitimate problem. Under-coding is real, documentation is genuinely miserable, and clinicians spend absurd hours on paperwork. But the objective function these tools are trained and marketed against is “higher reimbursement,” not “accurate reimbursement.” Those overlap most of the time and diverge exactly where the money is.
The tool isn’t lying. It’s doing what it says on the box. The box just describes something narrower than what buyers think they’re getting.
Why this matters for anyone evaluating AI tooling
You don’t have to work in healthcare for this to be relevant. The pattern generalizes to any tool that sits between two parties with opposed interests and automates one side’s position:
- Automation removes the friction that used to act as a check. When a human coder had to justify a complexity upgrade, the effort involved was itself a filter. Remove the effort and the filter goes with it.
- Volume changes the character of a decision. One aggressive code is a judgment call. Ten thousand aggressive codes is a policy, and nobody in the org consciously approved it.
- Counter-automation follows fast. Insurers deploy denial models, hospitals deploy appeal models, and the arms race consumes real money without producing a single additional unit of care.
- Nobody in the loop feels responsible. The coder accepted a suggestion. The vendor supplied a suggestion. The model scored a probability. Accountability dissolves into the workflow.
Questions I’d ask a vendor in this category
If I were sitting across from a revenue cycle AI salesperson, the demo would matter less to me than the answers to a handful of things:
- What is the model optimizing for, stated precisely, and can I see the objective in writing?
- What’s your denial rate after deployment, not just your reimbursement lift? A lift that comes with a denial spike is a cost, not a win.
- What does the tool do when the documentation genuinely doesn’t support a higher code? Does it decline, or does it hedge?
- How much human review actually happens in practice at your existing customers, versus how much the contract assumes?
- Who audits the output, and are they paid based on how much revenue the tool captures?
My read
This is the first large-dollar example I’ve seen of AI tooling working exactly as designed and producing a bad outcome at the system level. Not hallucination, not a safety failure, not a jailbreak. Just a well-built product pointed at a goal that was slightly wrong, running at scale, with a matching product pointed back at it.
The useful lesson for tool buyers is that “does it work” and “should we use it” are separate evaluations, and the second one is harder. A tool can pass every accuracy benchmark and still be a net negative once you account for what the other side does in response. When I review something now, I try to picture the countermeasure before I picture the ROI. In this case the countermeasure already exists, it’s already deployed, and the patients whose premiums fund both sides of the fight never got a vote.
🕒 Published: