\n\n\n\n Shrug-Driven Development Is Now a Feature - AgntBox Shrug-Driven Development Is Now a Feature - AgntBox \n

Shrug-Driven Development Is Now a Feature

📖 5 min read•823 words•Updated Sep 28, 2026

There’s a line in a September 2026 post on i hate the future that I haven’t been able to shake: a broken door can make a joke, but a broken AI service can make a business process impossible to trust. The author builds it out of a scene from President Curtis, where the president wrestles with two doors, one of them fitted with a body. Doors are funny when they fail. You push, it doesn’t move, you laugh, you pull. The failure is legible. You can see the hinge.

That’s the distinction I keep running into in this job, and it’s the one the post nails better than most engineering writing I’ve read this year. The same post argues that the tragedy of software engineering today is that we are actively engineering systems whose failures cannot be explained. Not systems that fail rarely. Systems that fail unaccountably.

What breaking used to look like

I review AI tools for a living. Most weeks that means poking at something until it does something wrong, then figuring out why. The “why” part used to be the easy half. A stack trace, a rate limit, a malformed payload, a config key nobody set. You could write it down and a reader could act on it.

Increasingly, the “why” is just gone. A tool returns the right answer nineteen times and something structurally insane on the twentieth. Same input, same version, same day. When I ask vendors about it, the honest ones say some version of “yeah, that happens.” The less honest ones say the model has been improved.

Notice that both answers are the same answer. Neither one contains a mechanism. Neither one gives me anything to test against. I’ve started logging these cases separately in my notes, because the traditional bug category doesn’t fit. A bug implies something is wrong with the implementation. These aren’t bugs. They’re the implementation.

How the shrug becomes policy

The normalization part is what deserves attention. Nobody decided this was acceptable in a meeting. It arrived by accumulation. Enough tools behaved inexplicably that inexplicable behavior stopped reading as a defect. It became weather.

You can watch the vocabulary adapt in real time. Failures get called hallucinations, which sounds almost whimsical, like the software had an experience. Retries become a design pattern rather than an admission. “Non-deterministic” gets deployed as a shrug instead of a specification. Every one of those words does the same work: it moves an unexplained failure out of the accountability column and into the character column. That’s not how software is, that’s just how it is.

And the tooling reinforces it. When your product’s core primitive can’t explain itself, your documentation can’t explain it either. So docs get thinner and vibier. Instead of error codes you get prompting advice. Instead of guarantees you get best practices. The user is quietly handed responsibility for a failure mode they can neither predict nor diagnose.

Why the doom conversation makes this worse

This is happening alongside a much louder argument. Jacob Coxen, a young former Anthropic researcher, has been warning that AI and superintelligence could kill everyone within the next ten years. That warning got attention, as it should, and the whole topic went up the Hacker News front page, where the discussion pulled in people like commenter veep47.

Here’s my concern with how those two conversations sit next to each other. Existential risk is enormous and abstract and ten years out. The unexplained failure in your invoice pipeline is small and concrete and happened Tuesday. When the big conversation gets all the oxygen, the small one starts to look like whining. Who cares about a flaky extraction step when the stakes are supposedly the species?

I’d argue it’s the opposite. Our ability to notice, explain, and fix ordinary failures is the only muscle we have. If that muscle atrophies at the invoice-pipeline scale, there is no version of us that handles a larger problem well. Explainability isn’t a nice-to-have you bolt on before the serious stakes arrive. It’s the practice you either keep or lose.

What I’m doing about it in reviews

I’ve changed how I score things, and I’d suggest the same lens if you’re evaluating tools for your team:

  • Can it tell you it failed? Not guess-and-retry. An actual signal you can branch on.
  • Does the same input produce the same output? If not, is that documented with bounds, or just vibes?
  • When it breaks, is there a trace? Inputs, intermediate steps, versions. Something a human can read.
  • Does support answer “why” or answer “try this”? One of those is engineering.

A tool that fails 5% of the time and tells you which 5% is more useful than a tool that fails 1% of the time silently. That’s not a hot take, it’s just how you build anything you intend to depend on.

The door in the sketch is funny because you can see the hinge. Keep asking to see the hinge.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top