Remember when GitHub Copilot first landed and half of developer Twitter spent a month arguing about whether autocomplete counted as programming? The stakes were low. If the suggestion was wrong, you deleted it. Worst case, you shipped a bug to a side project nobody used.
That argument feels quaint now. General Motors has revealed that nearly 90% of the code produced by its autonomous driving team is AI-generated. Not a side project. Not a dashboard. Code that decides whether a two-ton vehicle brakes.
That number is the reason I’m paying attention to a single panel at TechCrunch Disrupt 2026. On the Real World AI Stage, Shield AI CTO Nathan Michael, Waabi founder and CEO Raquel Urtasun, and GM Director of Robotics Strategy Mikell Taylor are sitting down to talk about building AI when failure isn’t an option. Safety validation, regulatory navigation, trust-building. An hour on the topic that most AI tooling conversations skip entirely.
Why a Tool Reviewer Cares About a Conference Panel
I spend most of my time testing AI coding tools and writing up what actually holds together versus what falls apart the moment you push past a demo. The pattern I keep hitting is that almost every tool in this category is optimized for generation speed and almost none of them are optimized for verification.
The pitch is always some variant of “write more code, faster.” Fine. But the hard part of software was never typing. The hard part is knowing whether what you typed is correct. When a coding assistant produces a function in two seconds, you’ve moved the bottleneck, not removed it. You’ve just relocated it to review, testing, and validation, which are exactly the parts nobody has figured out how to accelerate at the same rate.
Defense autonomy, self-driving trucks, and production vehicle software are the three domains where that mismatch gets expensive fast. These teams can’t ship and patch on Tuesday. They have to prove correctness before deployment, to regulators who are not impressed by benchmark scores.
The Question I Want Answered
If 90% of your autonomous driving code is machine-generated, something in your validation pipeline has to be doing enormous work. That’s the part I want details on, because it’s the part the tooling market keeps quiet about.
Specifically:
- What does code review look like when humans are reviewing far more code than they author?
- How do you do root-cause analysis on a failure in code no person wrote line by line?
- Does the verification tooling scale with generation volume, or is there a growing gap someone is absorbing with headcount?
- How do you explain an AI-generated decision path to a regulator who wants a traceable rationale?
Those aren’t gotcha questions. They’re the operational reality of any team that adopts these tools at scale, and right now most of us are guessing at the answers based on hobbyist-scale experience.
Three Different Flavors of Not Failing
What makes the panel lineup interesting is that these three organizations face different versions of the same problem. Shield AI operates in defense, where adversarial conditions are the baseline assumption. Waabi has been building toward scalable autonomy for trucks and robotaxis, backed by $1 billion in new funding raised in January 2026, with a public position on generalization as the next frontier. GM is a legacy manufacturer with existing production lines, existing liability exposure, and existing regulatory relationships.
Those constraints produce different engineering cultures. A defense contractor’s threat model doesn’t map onto a highway trucking fleet, and neither maps onto a manufacturer shipping consumer vehicles at volume. If all three have converged on similar validation practices despite that, the practices are probably sound. If they haven’t, the disagreements will be more instructive than any consensus.
What I’ll Be Watching For
My honest read on the current crop of AI development tools is that they’re strong at drafting and weak at guaranteeing. That’s a fine tradeoff for a prototype and a bad one for anything load-bearing. The teams pushing autonomy forward have had to solve this because their failure mode is physical, not a bad user review.
So I’m hoping for specifics over philosophy. Named techniques. Actual pipeline architecture. The trade-offs they accepted and what broke before they got it right. That’s the kind of detail that transfers to the rest of us, even those working on things that can’t hurt anyone.
The 90% figure will get the headlines. The interesting question is what the other 10% of the process looks like, and how much of it is humans double-checking the machines.
🕒 Published: