Over 10,000 organizations. More than 12,000 compromised inboxes. Two men, aged 32 and 38. That’s the scorecard Microsoft put on the table on 22 September 2026 when its Digital Crimes Unit announced it had disrupted EvilTokens, a subscription-based cybercrime service that used an AI chatbot to help attackers pick their victims.
I review AI toolkits for a living. I spend my weeks poking at onboarding flows, checking whether a product’s pricing tiers make sense, and figuring out if the assistant bolted onto the side actually helps or just burns tokens. So when I read Microsoft’s description of EvilTokens as a commercial operation with subscribers, I had an uncomfortable moment of professional recognition. This wasn’t a hacker in a basement. This was a product.
What the attackers actually bought
Strip away the criminal part and look at the shape of the thing. EvilTokens was phishing-as-a-service. Subscription model. A chatbot that, per Microsoft, helped attackers decide which victims to pursue and how to exploit them. The technique at the center of it was device-code phishing, the flow where an attacker convinces you to authorize a login on their behalf rather than stealing your password outright.
That combination is what makes this interesting rather than just grim. The AI layer wasn’t doing the hacking. It was doing the part humans are bad at: triage. Deciding who’s worth the effort. Suggesting the next move. If you’ve ever used an AI assistant inside a CRM to rank leads, you already understand the architecture. Same pattern, different target list.
I keep saying in reviews that the value of an AI feature usually isn’t the flashy generation part. It’s the boring judgment work in the middle of a workflow. EvilTokens is an ugly proof of that thesis. The criminals arrived at the same conclusion the good SaaS companies did, and they shipped it.
Why device-code phishing should worry you more than it does
Here’s what I want readers of this site to take away, because it affects how you evaluate the tools in your own stack. Device-code phishing doesn’t need your password. It works by getting you to complete a legitimate authorization step on an attacker’s terms. The end result is a valid session, which is why account takeover services like this one can rack up numbers in the five figures.
A lot of AI toolkits and agent platforms you’re being sold right now ask for exactly this kind of access. OAuth grants. Device-code sign-ins on a headless box. “Connect your workspace” buttons that hand a third party a token with broad scopes and no clear expiry. I’m not saying those products are malicious. I’m saying the consent flow they train you to click through is the same flow a criminal service was monetizing at scale through most of 2026.
When I’m testing an integration now, these are the things I actually check:
- What scopes does it request, and does it explain why it needs each one
- Can I see and revoke active sessions from an admin console without filing a support ticket
- Does the sign-in flow ever ask me to type a code into a page I didn’t initiate
- Are tokens short-lived, or does one approval last effectively forever
None of that is exotic. It’s just hygiene that most product tours skip because it makes the demo feel slower.
The takedown is the encouraging part
Microsoft’s DCU worked with industry partners and law enforcement, and the Metropolitan Police Service arrested two suspected administrators on 11 September 2026, eleven days before the public announcement. That sequencing tells you something about how these operations run: the legal action lands first, the press release follows.
It’s a solid outcome, and I’d rather report it than another unattributed breach. But I’ve watched enough tooling cycles to know that a disrupted platform is a disrupted platform, not a solved problem. The blueprint is public now. Subscription billing, a chatbot for target selection, a token-theft technique that sidesteps password strength. That’s a product spec anyone can copy, and the components are commodity.
What I’d change in how we review AI products
The lesson I’m taking into my next batch of reviews is that “does the AI work” is no longer the interesting question. It mostly works. The interesting question is what a motivated bad actor could do with the same capability pointed the other way, and whether the vendor has thought about that at all.
I’d also stop treating multi-factor authentication as the end of the conversation in security sections. EvilTokens built a business on the gap between “the login was authenticated” and “the login was intended.” Any toolkit that touches your identity provider lives in that gap.
Twelve thousand inboxes across ten thousand organizations is not a story about clever hacking. It’s a story about someone applying decent product thinking to a bad idea, and it working. The uncomfortable part, for those of us who evaluate this software all day, is how familiar the product looked.
đź•’ Published: