Remember when GPT-4 launched and half the internet spent a week feeding it the bar exam? That was the era when we measured model quality by how well it did on tests designed for humans. Standardized exams, coding puzzles, trivia. The scoreboard was familiar even if the contestant wasn’t.
GPT-6 Astra, which OpenAI announced this September, is playing a different sport. OpenAI is pitching it as a new generation of intelligence with advanced cybersecurity and problem-solving abilities, and says it outperforms previous models at exploit development and code execution. Some coverage has gone further, framing the release as a possible start of the AGI era. That’s a lot of weight to put on one model card.
I review tools for a living, so let me tell you what actually caught my attention in the launch material. It wasn’t the AGI talk.
The benchmark detail that matters
Buried in OpenAI’s own documentation is an acknowledgment I wish more vendors would make. The company noted concerns that exposure to historical software vulnerabilities may have affected benchmark results. Translation: if a model was trained on the entire public record of disclosed security bugs, then testing it on those same bugs measures recall, not reasoning.
So OpenAI built additional evaluations. One is an internal set called “ExploitBench – Internal Port (June–August 2026),” described as containing only recent vulnerabilities disclosed after Astra’s training. That’s the right instinct. Testing a security model on bugs it couldn’t have memorized is the difference between a real capability claim and a very expensive parlor trick.
I want to be clear about what I’m praising here. I’m not praising the scores, because the scores are internal and I haven’t seen them independently reproduced. I’m praising the methodology disclosure. A vendor telling you why its own numbers might be inflated is rare enough that it deserves credit.
What “good at exploits” actually means for your workflow
Here’s where I get skeptical, and it’s the same skepticism I bring to every tool that promises to change how a team works.
Exploit development is not one task. It’s a chain of them:
- Finding a candidate weakness in a large codebase
- Understanding whether it’s reachable in practice or theoretically interesting only
- Building something that actually triggers it
- Making that something reliable enough to matter
- Knowing when to stop, because the finding is a duplicate or a dead end
A benchmark can score the middle links of that chain. It struggles with the first and last, which are the ones that eat a security engineer’s week. So when OpenAI says Astra is stronger at exploit development, my next question is always: at which link? The launch framing doesn’t answer that, and I’d rather say so than pretend otherwise.
Same story with code execution. Stronger execution ability is genuinely useful, and it’s also the capability most likely to turn a small mistake into a large one. A model that can run code is a model that can run the wrong code, in the wrong directory, with the wrong permissions. That isn’t a knock on Astra specifically. It’s the tradeoff for any agent you hand real tools to.
The dual-use problem nobody solved
You cannot build a model that’s excellent at finding vulnerabilities and only good at finding them for the right people. The capability is the capability. OpenAI publishing a deployment safety hub and a system card suggests they know this. Whether the safeguards hold under adversarial pressure is a question that gets answered by researchers over months, not by a launch post.
For teams evaluating this, I’d treat Astra’s security claims the way you’d treat any new scanner. Run it against your own code, with findings you already know about, and see how it does. Then run it against code you haven’t audited and see how much of its output is noise. That second number is the one that decides whether the tool saves you time or generates homework.
My honest read
Astra looks like a real step up on a specific set of technical capabilities, documented with more methodological care than I expected. That’s a solid release. It is not, based on what’s public, evidence that we’ve crossed into a new category of machine intelligence, and I’d separate those two claims firmly in your head before you make budget decisions.
The AGI framing is doing marketing work. The ExploitBench methodology is doing engineering work. Pay attention to the second one.
I’ll be testing Astra against a set of tasks I’ve run on every major model since GPT-4, including a few where the previous generation failed in ways that were interesting rather than just wrong. When I have numbers I trust, you’ll get them here, along with the cases where it fell apart. That second part is usually the more useful review.
đź•’ Published: