\n\n\n\n Two benchmarks no hacker had touched yet - AgntBox Two benchmarks no hacker had touched yet - AgntBox \n

Two benchmarks no hacker had touched yet

📖 1 min read•181 words•Updated Sep 6, 2026

Forty-one percent. That is the number that should give you pause. When OpenAI’s GPT-6 Astra was evaluated on fresh, never-before-seen exploit benchmarks built specifically to catch a model that had already memorized the answers, it still managed to find and exploit vulnerabilities that no one had disclosed until months after its own training cutoff. A model that had never seen the vulnerability, finding the vulnerability, on its own.

That is the difference between an AI that has read about hacking and an AI that can hack.

Why the old benchmarks suddenly did not matter

The release of GPT-6 Astra, developed by Amazon’s team and made available through platforms including Amazon Web Services, generated the usual cycle of excited press releases and breathless headlines. But the more interesting story happened behind the scenes, inside OpenAI’s own testing process.

Here is the problem they faced: Astra scored exceptionally well on cybersecurity and reasoning benchmarks. So well, in fact, that the team started to wonder whether the results meant anythind how fast the rest of the industry catches up — is the story worth following.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top