It’s a Thursday morning in September and you have three tabs open. The first is a viral thread about an AI model doing something it very much was not supposed to do. The second is a TechCrunch piece explaining that two AI safety conversations went viral that week and that both of them are hard to verify. The third is your actual job, a half-finished evaluation of some agent framework you promised to review by Friday. You read the first tab twice. You forward it to a coworker. You do not open the third tab again for forty minutes.
That’s the whole problem, and I say that as someone whose entire job is testing tools and telling you whether they hold up. I can run an agent against a task suite two hundred times and give you a pass rate. There is no reproducible test for a feeling that something has gone sideways.
What actually happened in September 2026
Let’s separate the reporting from the noise, because the reporting is fairly thin and fairly specific. AI safety discussions have intensified this year following unexpected model behaviors and concerns about potential risks. TechCrunch’s Rebecca Bellan reported on September 15 that OpenAI, Anthropic, and Google have been in talks on AI safety for weeks, with Chris Lehane, OpenAI’s global policy chief, speaking to reporters that Tuesday. Regulatory discussions are underway. Broadcast news picked it up two days later, with both ABC News and CBS News running segments on the concerns, ABC’s featuring Cynthia Kaiser, an SVP at Halcyon.
That’s it. That’s the verified core. Competitors talking to each other about safety, regulators circling, and enough unexplained model behavior to make all of it newsworthy.
Now compare that to what you probably absorbed about AI safety in September 2026. Bigger, right? Louder. More specific in ways the actual reporting isn’t. TechCrunch’s framing was that these conversations demonstrate how hard it has become to tell AI fact from AI fiction, and one of the two viral cases involved Andrew Yang, the former presidential candidate. The story was not that AI did something terrifying. The story was that nobody could establish what AI had done.
Why this breaks a reviewer’s toolkit
My methodology is boring on purpose. I install the thing. I give it tasks with known correct answers. I log failures, count them, and check whether the failures repeat. If a tool behaves badly once and never again, I say so. If it behaves badly reliably, that goes in the headline. Repeatability is the whole product.
Safety discourse in its current form gives me none of that. What it gives me is:
- A clip with no system prompt attached
- A screenshot with no model version
- A claim about behavior with no way to trigger it again
- A well-known name attaching credibility to something unverified
- Company statements that are accurate and also carefully scoped
Every one of those items is unfalsifiable. You cannot review unfalsifiable. You can only react to it, and reacting is what the whole machine wants from you.
The part that should actually get your attention
Strip away the viral layer and one detail is genuinely unusual: three companies that compete hard have reportedly been talking for weeks. Firms do not coordinate on a problem they consider handled. Add regulators to the room and you get a picture of an industry that has decided it would rather shape the rules than receive them.
Whether that’s caution or positioning, I can’t tell you, and neither can anyone quoting a fifteen-like news clip at you. What I can say is that the coordination is the signal and the screenshots are the static. Watch the talks. Watch the regulatory filings when they land. Those leave paper trails you can check.
What I’d actually do this week
Practical stuff, since that’s supposedly why you read this site. When a safety story crosses your feed, ask three questions before you forward it. Which exact model and version. Who reproduced it besides the original poster. What the system prompt was. If you can’t answer all three, you have a story about people, not a story about software.
Then go test your own stack. Not the scary one from the video, yours. Run your agents against your real tasks with your real data and log what breaks. That’s a solid signal about the risk sitting in your own systems, and it beats any amount of secondhand alarm.
The uncomfortable thing about 2026 isn’t that models are doing strange things. Models have always done strange things, and documenting them is most of my job. The uncomfortable thing is that our shared ability to check claims got worse at exactly the moment the claims got more serious. I don’t have a fix for that. I do have a task suite, and it’s still the most honest tool I own.
🕒 Published: