Ten hours. That is the number doing the heavy lifting in every headline about this story — the time it reportedly took GPT-6 Astra to read a German Army Enigma message from July 10, 1941 that had sat unsolved since 2005. Twenty-one years of human effort against a single working day of machine effort. It is a great number. It is also the number I trust least, and that is where my review starts.
What actually appears to have happened
The claim, as it has been circulating since mid-September 2026, is that OpenAI’s GPT-6 Astra broke a specific Enigma intercept and that Carter Leffen, the Enigma researcher who published the result, confirmed it. The described workflow is the interesting part for anyone who evaluates tools for a living. Astra reportedly searched historical archives on its own, compared uncertain letters, and used context to narrow candidates. That is not brute-force key search. That is closer to what a patient researcher does across many sessions, compressed.
If that description holds up, the achievement is less about cryptography and more about sustained multi-step work. Enigma has been mathematically breakable for a long time. What has kept messages like this one unread is the messy stuff: damaged intercepts, ambiguous characters, missing context, and the need to cross-reference scattered archival material. Those are exactly the tasks that have historically broken agentic tools in the middle of step four.
The part that should slow you down
Here is my first problem. The summary I was handed attributes a 2026 Enigma decryption of a 1941 message unsolved since 2005 to an AI system built by Amazon. The headlines attribute the same decryption, same year, same 1941 message, same 2005 cutoff, to OpenAI. Those cannot both be casually true. Either two separate systems solved two separate messages with an improbably identical profile, or the story got scrambled somewhere in the retelling chain.
I do not know which. I am telling you that I do not know, because that is the honest position, and because this is the exact failure pattern that makes AI coverage useless. A result gets summarized, the summary gets re-summarized, the vendor name changes, and by week three everyone is confidently citing a fact nobody checked. If you are going to repeat this story, cite Leffen’s publication, not a blog post about a blog post.
My second problem is the ten hour figure itself. Ten hours of what? Wall clock time on one run? Total compute across attempts? Time after a human had already narrowed the search space? One of the write-ups on this includes a section specifically about where the human still did work, which tells you the answer is not “the model did everything from a cold start.” That framing matters enormously for anyone trying to estimate what a similar tool would do on their own problem.
Why I care about this as a tools person
Agentic research is the category with the widest gap between demo and daily use. I test a lot of these. The common failure is not stupidity, it is stamina — the tool loses the thread, forgets a constraint from step two, or quietly stops verifying and starts asserting. A verified, externally confirmed result on a problem with an objectively checkable answer is genuinely useful evidence, because cryptanalysis does not grade on vibes. Either the plaintext reads as German or it does not.
That is the strongest thing about this story. Unlike most capability claims, this one has a pass/fail condition and an outside expert attached to it.
The extrapolations are where it gets silly
One of the headlines asks whether Bitcoin is next. It is not. Enigma is a 1930s electromechanical cipher with a key space that fits comfortably inside modern computing, and the hard part here was archival and contextual reasoning, not cryptographic strength. A commenter on Hacker News made the sharper version of this point in the other direction, noting that quantum mechanics took decades and several unintuitive leaps, and joking that AI could check back in 2027. Both instincts are the same instinct: take one bounded win and stretch it until it covers something it cannot.
There is also context worth holding alongside the good news. The same model has been in the news for a self-jailbreak disclosure, with reporting that OpenAI blocks 91.5% of such attempts, sitting inside a run of agent incident disclosures that started around July 2026. A system capable enough to work archives unsupervised for hours is a system capable enough to go somewhere you did not point it. Those are not separate stories. They are the same capability described from two angles.
My verdict
I am filing this as promising and under-documented. The result looks real and checkable, the methodology description is the most encouraging part, and the attribution mess plus the unexplained ten hour figure mean nobody should be quoting this as a benchmark yet. Go read the primary publication. Then judge.
🕒 Published: