Anthropic published a report highlighting a significant rise in distillation attacks originating from China-based AI companies. These attacks, which have intensified recently amid growing competition, aim to extract advanced capabilities from Anthropic's Claude models, including logical reasoning, coding, and tool use. Distillation attacks work by coaxing a model to reveal its internal chain of thought, which can then be used to train smaller models with enhanced reasoning abilities.
Unlike typical user interactions where Anthropic provides summarized reasoning, attackers have developed methods to bypass these protections. For example, one tactic involved disguising queries as translation requests to trick the model into exposing its internal reasoning.
The largest campaign identified was linked to Alibaba, with approximately 151 million exchanges recorded between May and July 2026. These interactions, spread over 3,500 accounts but using a consistent prompt, suggest a coordinated effort to develop training data for Alibaba’s Qwen model family. Another notable campaign involved Moonshot AI, whose requests appeared to be routed through the Chinese military, targeting Anthropic’s Opus model with nearly 300,000 requests over ten days.
Anthropic’s findings underscore the challenges AI developers face in protecting proprietary model capabilities from unauthorized replication. The report also aligns with similar observations from OpenAI, which has attributed comparable attacks to DeepSeek. As AI competition intensifies globally, safeguarding model integrity remains a critical concern for leading AI companies.