Anthropic Reports Distillation Attacks by Alibaba, Moonshot AI, and DeepSeek
Anthropic released a new report on Thursday alleging that Chinese AI firms have been conducting persistent distillation attacks against its models. The company claims that these attacks—where a model is trained to replicate the outputs of a target system—have increased in frequency and sophistication as competition in the generative‑AI market intensifies. According to the report, the attacks are designed to extract proprietary knowledge from Anthropic’s systems and then redistribute it under the guise of open‑source models.
The report details several instances in which Chinese companies reportedly used publicly available data to train surrogate models that mimic Anthropic’s behavior. Anthropic says the attacks not only compromise intellectual property but also pose a risk to the safety and reliability of AI outputs. While Anthropic has not publicly disclosed specific names of the companies involved, it has urged regulators and industry peers to strengthen safeguards against model theft and to improve transparency in the deployment of large‑language models. The findings come at a time when the global AI industry is witnessing rapid growth and heightened scrutiny over data security and cross‑border technology transfer.