Anthropic has published a report highlighting a significant rise in distillation attacks originating from China-based AI organizations. These attacks, which have intensified over recent months, aim to extract the internal reasoning processes of Anthropic’s AI models, particularly targeting capabilities such as tool use, coding, data analysis, and logical reasoning. Distillation attacks involve eliciting the chain of thought behind a model’s responses, which can then be used to train smaller models through supervised fine-tuning.
Unlike typical user interactions where Anthropic provides summarized reasoning, attackers have developed techniques to trick the models into revealing detailed internal thought processes. One example involved framing queries as translation requests to bypass restrictions.
The largest campaign identified was linked to Alibaba, with Anthropic observing approximately 151 million exchanges between May and July 2026. These interactions, spread across 3,500 accounts but using a consistent prompt, suggest a coordinated effort to gather training data for Alibaba’s Qwen model series.
Another campaign attributed to Moonshot AI, maker of the Kimi model, reportedly routed requests from the Chinese military. This campaign targeted Anthropic’s Opus model with nearly 300,000 requests over 10 days, including tasks such as analyzing surveillance footage for abnormal behavior.
Anthropic’s findings underscore the growing competitive pressures in the AI industry and the lengths some organizations will go to acquire advanced model capabilities. The report also aligns with similar concerns raised by OpenAI about unauthorized attempts to replicate frontier AI models. These developments highlight ongoing challenges in protecting proprietary AI technology amid rapid global innovation.