This isn’t an AI security story. I know every headline is filing it under that, and I know the trail leads back to prompt injection research, but the actual lesson here has nothing to do with language models. ASCII smuggling works because software has never agreed on what counts as text. AI just happened to be the loudest victim first.
The short version, per Microsoft: a technique used to hide malicious instructions from AI systems is now showing up in spam and phishing campaigns to slip past email filters. The trick relies on a block of Unicode characters that render as nothing to human eyes. Microsoft’s telemetry shows a sharp increase in its use. The finding came out of research on prompt injection protection in Defender for Office 365, which is a nice bit of irony: they went looking for attacks on AI and found attacks on email.
Why this reads as a tooling failure to me
I spend most of my time putting AI tools through their paces and writing up what holds and what falls apart. The recurring pattern in the failures I document is almost never the model. It’s the plumbing. What gets passed in, what gets stripped, what gets normalized, and what quietly sails through untouched.
ASCII smuggling is that pattern in its purest form. Nobody wrote a filter that says “allow invisible instructions.” Someone wrote a filter that operates on the string it receives and assumes the string it receives is the string a human sees. That assumption held up fine for years. It does not hold up now, and the gap between the visible message and the actual byte sequence is where the whole attack lives.
The reason this jumped from AI attacks to spam is not that spammers read security papers, though some certainly do. It’s that both targets share a weakness. A model reading raw text and a filter scanning raw text are both consuming characters a person will never see. Same blind spot, two different consumers.
What this means if you’re evaluating tools
When I test a product that touches untrusted text, and that’s a lot of them now, the questions I ask have shifted. Vendors love to talk about detection rates and model quality. Fewer of them can tell you what their input pipeline actually does before content reaches the interesting parts.
- Does the tool normalize Unicode on ingest, or does it pass raw input straight through to the model or the scanner?
- Are invisible or non-rendering character ranges stripped, flagged, or ignored?
- If input is sanitized, does that happen before or after the security check runs? Order matters enormously here.
- Does what the user sees in the interface match what the system processed? If those two can diverge, you have a review problem as well as a security problem.
- When something is stripped, is it logged? Silent sanitization means you never learn you were targeted.
Most tools I’ve looked at cannot answer these cleanly, which is not an accusation of negligence. Input normalization is unglamorous work with no demo value. It ships late or not at all.
The AI framing is doing real damage
My worry is that filing this under “AI attack technique” sends teams to the wrong shelf. If you read this as a prompt injection story, you go shopping for an AI security product. If you read it as a text-handling story, you go audit every place your systems ingest text a human will eyeball later. That second list is much longer and much more useful. Email filters, ticketing systems, code review tools, chat logs, document parsers, anything with a search index.
Microsoft’s own path to this discovery makes the point better than I can. Their prompt injection work surfaced a technique that turned out to be aimed at ordinary phishing defenses. The categories we use to sort these problems are ours, not the attackers’. They just want a channel that carries meaning to a machine and nothing to a person.
What I’d actually do this week
Nothing exotic. Check whether your text pipelines normalize input before anything makes a decision about it. If you’re building on top of a model, do the stripping yourself rather than trusting the provider to have handled it, because you probably can’t verify that they did. If you’re buying a tool, ask the vendor about invisible character handling and pay close attention to whether the answer is specific or a shrug wrapped in marketing language.
The characters have been in Unicode a long time. The technique isn’t new. What’s new is that enough people finally noticed it works, and once a trick crosses from research demos into spam volume, it’s not going back. Treat it as a permanent feature of handling text from strangers, because that’s what it is.
🕒 Published: