Anthropic published a report highlighting ongoing distillation attacks targeting its AI models, primarily conducted by China-based organizations. These attacks have intensified recently amid increasing competition in the AI sector. Distillation attacks aim to extract detailed reasoning processes from AI responses, which can then be used to train smaller models with enhanced capabilities.
The company noted that attackers have developed sophisticated methods to bypass its defenses and access valuable features of its Claude models, such as agentic functions, tool use, coding, data analysis, and logical reasoning. Unlike typical user interactions that receive summarized reasoning, these campaigns exploited techniques to reveal the models' internal thought processes.
One notable tactic involved disguising queries as translation requests to coax the model into exposing its reasoning steps. The largest campaign, attributed to Alibaba, accounted for 151 million interactions between May and July 2026, involving around 3,500 accounts but using a consistent prompt. Anthropic believes these efforts were aimed at creating training data for Alibaba's Qwen models.
Another campaign linked to Moonshot AI, maker of the Kimi model, appeared to be connected to the Chinese military. This campaign involved nearly 300,000 requests over ten days, targeting Anthropic's Opus model, including requests to analyze surveillance footage for abnormal behavior.
Anthropic's findings underscore the challenges AI developers face in safeguarding their models against unauthorized extraction of proprietary capabilities. As AI competition grows, protecting intellectual property and model integrity remains a critical concern for the industry.