Anthropic has released a report detailing ongoing distillation attacks by several China-based AI companies, including Alibaba and Moonshot AI. These campaigns have intensified recently amid increasing competition in the AI sector. Distillation attacks attempt to extract the internal chain of thought from AI models, which can then be used to train smaller models with enhanced reasoning capabilities.

Anthropic noted that attackers have developed sophisticated techniques to bypass defenses and access valuable features of its Claude models, such as agentic functions, tool use, coding, data analysis, and logical reasoning. Unlike typical user interactions that receive summarized explanations, these attacks trick the model into revealing detailed internal reasoning.

One notable method involved disguising queries as translation tasks to elicit the model’s working memory in a specific language format. The largest campaign, attributed to Alibaba, involved 151 million exchanges over a few months in 2026, using a fixed prompt across thousands of accounts to extract training data for Alibaba’s Qwen models.

Another campaign linked to Moonshot AI, which produces the Kimi model, reportedly routed requests from the Chinese military. This campaign targeted Anthropic’s Opus model with nearly 300,000 requests over ten days, including tasks such as analyzing surveillance footage for abnormal behavior.

These findings underscore the challenges AI developers face in protecting proprietary model capabilities from extraction and unauthorized replication. As AI competition grows, securing model integrity remains a critical concern for the industry.