US Government Accuses Chinese AI Firms of Distilling Frontier Models
The US government, through agencies like the FBI and NSA, has accused several Chinese AI firms of systematically stealing proprietary capabilities from leading US AI models – including Claude, GPT, Gemini, and Grok – via industrial-scale distillation campaigns. These firms are allegedly using US models to train their own, significantly reducing development costs and accelerating their AI development timelines, with the Chinese government reportedly aware of and supporting these activities. The agencies are urging US AI companies to implement robust monitoring and intelligence sharing to combat this growing threat.
The US government, including the FBI, the National Security Agency (NSA), and the Cybersecurity and Infrastructure Security Agency (CISA), has issued a joint advisory warning of a concerning trend: Chinese AI firms are allegedly engaged in industrial-scale efforts to extract proprietary capabilities from leading US AI models. These firms – including Alibaba, DeepSeek, MiniMax, Moonshot AI, StepFun, and Z.AI – are utilizing a process called distillation, a common practice in AI research, to significantly reduce their development costs and accelerate their AI development timelines.
According to the advisory, these companies are accessing US models – such as Claude, GPT, Gemini, and Grok – through massive distillation campaigns, extracting billions of tokens across millions of exchanges/requests since at least late 2024. The firms are allegedly training on outputs obtained in violation of the terms of service, deliberately extracting a competitor’s proprietary capabilities, and using evasive techniques to avoid detection. The US agencies believe this is happening with the awareness and support of the Chinese government.
Multiple techniques are being employed to access US frontier models at scale. These include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures. China-based AI firms are reportedly obtaining bulk premium subscriptions for US AI models and sharing them across teams of developers to save money during this intensive process.
To avoid detection, the firms are allegedly routing distillation requests through native APIs, remote cloud providers, and third-party aggregators that automatically obfuscate user metadata. They also take advantage of "transfer stations," a gray market of proxies used specifically to bypass US AI model geographic restrictions and evade safeguards.
DeepSeek, for example, ran an organized distillation campaign against frontier models of US AI companies since at least late 2024 in order to "generate synthetic training data for its models." The company targeted specific knowledge domains to extract proprietary functionality and reasoning capabilities to reduce their compute and research costs. DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.
Moonshot AI allegedly extracted "significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model." Both firms plundered models from Anthropic, OpenAI, Google, and xAI as part of this campaign.
The US agencies recommend that US AI companies implement comprehensive detection and mitigation measures to find anomalous and malicious behavior; share intelligence with other AI organizations to gain awareness of broader campaigns; and tune responses for suspected malicious attempts to reduce the effectiveness of these outputs. Ismael Valenzuela, vice president of labs, threat research and intelligence at Arctic Wolf, suggests treating model extraction and distillation "as a security event category" in itself, rather than treating it as API abuse or an intellectual property dispute. Organizations should prioritize 24/7 monitoring of AI API access patterns for automation at scale, anomalous prompt harvesting behavior, and distributed account creation, and work with model providers to share telemetry and indicators.
