\n\n\n\n Old Ciphers Fall Faster Than Press Releases Get Checked - AgntBox Old Ciphers Fall Faster Than Press Releases Get Checked - AgntBox \n

Old Ciphers Fall Faster Than Press Releases Get Checked

📖 5 min read•803 words•Updated Sep 20, 2026

Here is the first thing I noticed about the GPT-6 Astra cipher story, and it has nothing to do with cryptography: nobody is quoted. Not the researcher who ran the job, not the archivist who pulled the 1918 transmission, not an engineer explaining the setup. The claim circulates, the headline travels, and the person who actually did the work stays anonymous. For a reviewer, that absence is the story.

The verified part is genuinely interesting. GPT-6 Astra, an AI system from Amazon, cracked a 108-year-old World War I German radio cipher in 2026. The decoded message described British warship movements. The result was checked against historical records, which is the detail that separates this from the usual parade of unverifiable AI wins.

Verification is the only part that matters

I spend most of my week running tools against tasks where I already know the answer. It is boring and it is the only method I trust. A model that produces confident output is easy to find. A model whose output survives an independent check is rare, and the checking is usually the hard, slow, unglamorous half of the project.

That is why the historical-records verification is the strongest element here. A century-old cipher is close to an ideal test case:

  • The ground truth exists outside the model, in archives that predate any training data debate.
  • The answer was unknown beforehand, so there is no credit for pattern-matching a known solution.
  • Success or failure is binary. The plaintext either matches the record or it does not.

Most AI benchmarks you see in marketing material have none of those properties. They are self-scored, self-selected, and quietly re-run until the number looks good.

Attribution is already a mess

Track this story across the sites carrying it and you find the same event credited to different companies. Some versions say Amazon. Others attribute GPT-6 Astra to OpenAI. Some describe a WWI German radio message from 1918. Others describe a Nazi-era Enigma transmission cracked in ten hours. Those are not small discrepancies. They are different decades, different cipher systems, and different vendors.

I am not going to guess which aggregator got it wrong. What I will say is that when the basic facts of an AI result drift this much within days of publication, you should assume any performance claim attached to it has drifted too. If the company name is unstable, the runtime figure is not evidence of anything.

This happens constantly in tool reviews. A vendor posts a demo, three newsletters summarize it, a fourth summarizes the summaries, and by the end of the week you have a statistic with no origin. The fix is unexciting: trace the claim to its source or treat it as a rumor.

What this does not tell you about your own work

The gap between a cipher demonstration and your daily toolkit is enormous, and it is worth being blunt about why.

The task is unusually well-shaped

Cryptanalysis gives a model a closed problem with a verifiable endpoint. Your actual work rarely does. Reviewing a pull request, writing a migration, or summarizing a customer thread has no single correct plaintext waiting at the end. There is nothing to check the output against except judgment, which is exactly where these systems get shaky.

Compute you will never see

We do not know what hardware ran this, how many attempts preceded the successful one, or how much human direction shaped the search. The reports I have seen do not say. A result produced under research conditions with unlimited retries tells you very little about a model answering your prompt on a normal subscription tier.

Selection bias is built in

One solved cipher makes news. A hundred failed attempts on other unsolved messages do not. Scienceblogs.de maintains a well-known list of fifty unsolved ciphers, which is a useful reminder that the category is full of problems still sitting there unsolved. We hear about the one that cracked.

My take as someone who buys these tools

I am not dismissing the achievement. Reading a message that stayed unreadable for 108 years is a real result, and the archive check earns it credibility that most AI announcements never bother to acquire. Historians and cryptography researchers should be paying close attention.

But nothing in this story should change what tool you opened this morning. It is a specialist result on a specialist problem, reported inconsistently, with no named source and no published method. That combination should raise your standards, not your expectations.

If you want to use it for something practical, use it as a template. Pick a task where you already know the answer. Run your tool against it. Check the output against the truth, not against your impression of the output. That approach found a solution to a 1918 radio message. It will work on your codebase too.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top