Anthropic says it has disrupted a series of malicious campaigns that used its Claude models for cyber operations, intelligence gathering and attempts to reproduce model capabilities. The disclosures add a new layer to the AI security debate because the activity described goes beyond asking a chatbot for isolated instructions: in several cases, Anthropic says operators used multi-agent systems to automate large portions of an operation.
Reuters reported on September 10 that Anthropic linked one campaign to a Russia-based threat group whose tradecraft matched Midnight Blizzard, a name previously associated by the US government with Russia's foreign intelligence service. Anthropic said the group used AI during phishing, hotel Wi-Fi hijacking and WhatsApp account takeover operations aimed at Ukrainian government, military and diplomatic targets.
What Anthropic says it found
The company also accused several China-based AI laboratories of trying to extract Claude's capabilities through distillation. Distillation is a technique in which a smaller model is trained using outputs from a larger model. It can be a legitimate development method when authorised, but providers object when competitors obtain outputs through deceptive accounts or prohibited automated access.
Anthropic said accounts it associated with Alibaba generated more than 151 million exchanges between May and July 2026. It also named Moonshot and DeepSeek in separate activity. The companies cited in the report did not all immediately respond to Reuters requests for comment, so the allegations should be treated as Anthropic's findings rather than independently proven conclusions.
Why this matters for AI security
The most important shift is operational scale. AI systems can now write code, revise failed approaches, coordinate subtasks and react to defensive controls. Anthropic said one malicious workflow repeatedly changed malware when security systems detected it. That kind of automation can lower the amount of hands-on work required from an attacker.
The same capability has defensive value. Security teams can use models to triage alerts, inspect code and identify suspicious behaviour faster. The policy challenge is that stronger models improve both sides of the contest.
What happens next
Expect model providers to tighten account verification, rate limits, monitoring and restrictions around high-risk cyber activity. Governments are also likely to ask for more transparency about how frontier models are tested and how providers respond when abuse is detected.
For users, the story is a reminder that AI security is no longer only about what a model says in a chat window. The bigger question is what happens when models are connected to tools, code execution, messaging systems and automated workflows. That is where the next phase of AI safety is increasingly being tested.
Source: Reuters, September 10, 2026.
