\n\n\n\n A Broken Door Can Make a Joke, A Broken AI Service Can't - AgntBox A Broken Door Can Make a Joke, A Broken AI Service Can't - AgntBox \n

A Broken Door Can Make a Joke, A Broken AI Service Can’t

📖 5 min read•810 words•Updated Sep 28, 2026

The blogger behind i hate the future put it better than any vendor deck I’ve read this year: “A broken door can make a joke. A broken AI service can make a business process impossible to trust.”

I’ve been reviewing AI toolkits for long enough to recognize when a line lands. That one landed hard, because it names the exact thing I keep running into and keep failing to explain to readers who want a star rating and a verdict. The failure modes I’m testing against have changed shape. They used to be legible. Now they’re not, and somehow that’s fine with everyone.

The failure I can’t write up

When a tool breaks in an ordinary way, I can review it. The API returns a 500, the docs are wrong, the rate limit is undocumented, the SDK doesn’t build on a current runtime. All of that goes in the “what doesn’t work” column with a reproduction step, and you the reader can decide if you care.

What I can’t write up is the other category. The run that worked Tuesday and doesn’t Thursday, with no version change on either side. The output that’s subtly worse for a week. The pipeline that completes successfully and produces something nobody asked for. I re-run it, it works, and I’m left holding a note that says “sometimes it just doesn’t.” That’s not a review. That’s a shrug with a word count.

The same blog gets at the root of it: we’re actively engineering systems where the answer can be a few prompts away, and we’ve accepted that “a few prompts away” is a real engineering specification. It isn’t. It’s a coin flip we’ve agreed to call a feature.

Why 2026 made this worse

None of this is only an AI problem, which is the part I find genuinely useful to sit with. ERP deployments are failing at higher rates right now for reasons that have nothing to do with models. The observation making the rounds is blunt: organizations that can’t reliably deploy basic general ledger and accounting software are now stacking digital transformation layers on top of that same foundation.

Two forces are compounding. Complexity keeps climbing, and there aren’t enough people who understand the stack end to end. Labor shortages mean the person who could have told you why the job failed left in 2024, and their replacement inherited a system with no map. Add an AI layer on top of that and you’ve built something where nobody in the building can trace a bad output back to a cause.

So the failures show up, and they’re not the dramatic kind. They’re moderate. The project ships four months late. The migration mostly works. Customers notice something is off before you do. It costs real money and real time, and because nothing exploded, there’s no postmortem, no incident channel, no lesson. Just a slightly worse version of the thing you had before, absorbed into the baseline.

What I’m changing about how I test

I’ve stopped treating a single successful run as evidence of anything. A tool that works once has told me almost nothing. Here’s what I’m doing instead, and I’d suggest the same for anyone evaluating this stuff for their own team:

  • Run it repeatedly over days, not minutes. Same inputs, same config, spread across a week. Variance that shows up on day four is the variance that will hurt you in production.
  • Score explainability separately from output quality. A tool that fails and tells me why is worth more than a tool that succeeds slightly more often and stays silent. I now weight this heavily.
  • Test the failure path on purpose. Feed it garbage. Kill the network mid-run. If the tool reports success on a broken run, that’s disqualifying, and I’ll say so.
  • Ask what the vendor’s own support burden looks like. If their docs have no troubleshooting section beyond “try again,” that tells you what their internal debugging story is.

The standard worth holding

I don’t think the fix is refusing to use these tools. I use them. Some of them are good. What I’m not willing to do anymore is grade them on a curve where unexplainable behavior counts as an acceptable quirk of the category.

The door in that story is funny because a broken door is honest about being broken. You push, it doesn’t open, you know exactly where you stand. What we’ve built instead is a door that opens most of the time, closes for reasons nobody can name, and has a support page suggesting you push again.

When a tool I’m reviewing fails in a way neither I nor its makers can explain, that’s going in the review as a defect. Not a quirk, not a limitation of the space, not something to be expected. A defect. If enough of us hold that line, the shrug stops being free.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top