Industrial-Scale 'Knowledge Distillation': Chinese AI Firms Accused of Systematically Extracting US AI Model Capabilities
A joint advisory from the **NSA**, **CISA**, and **FBI** reveals a concerning trend: China-based AI companies are allegedly engaging in large-scale 'knowledge distillation' campaigns. These operations systematically extract proprietary functionalities from advanced U.S. AI models like **Claude**, **GPT**, **Gemini**, and **Grok**, accelerating their own development timelines and reducing costs.
U.S. intelligence and cybersecurity agencies are sounding the alarm over what they describe as aggressive, industrial-scale knowledge distillation efforts by China-based AI companies. This activity, detailed in a joint cybersecurity advisory by the **National Security Agency (NSA)**, **Cybersecurity and Infrastructure Security Agency (CISA)**, and **Federal Bureau of Investigation (FBI)**, points to a systematic extraction of proprietary functionalities and capabilities from leading U.S. AI models.
While 'distillation' is a recognized and legitimate technique in AI research, the advisory highlights the malicious and targeted nature of these operations, which appear to form the core strategy for the Chinese firms' AI development rather than a mere supplement.
### Widespread Campaigns by Major Players
Since at least late 2024, companies including **DeepSeek**, **Moonshot AI**, **Alibaba**, **MiniMax**, **StepFun**, and **Z.AI** have reportedly extracted billions of tokens across millions of requests from U.S. frontier AI models. These targeted models include variants of **Claude**, **GPT**, **Gemini**, and **Grok**.
**DeepSeek**, for instance, has allegedly conducted organized campaigns since 2024, focusing on extracting reasoning capabilities, specialized optimizations, and domain-specific functions to train its **R1** and **V3** models. **Alibaba** is accused of leveraging industrial-scale distillation to enhance its **Qwen** family of AI models, while **Moonshot AI**, **MiniMax**, **StepFun**, and **Z.AI** have also engaged in similar activities.
### Evasive Tactics and Cost Savings
To facilitate these large-scale distillation campaigns, China-based AI companies are employing sophisticated tactics to gain unauthorized access and bypass detection. These include:
* Routing requests through multiple pathways, such as native application programming interfaces (**APIs**), remote cloud providers, and third-party aggregators that obfuscate user metadata.
* Utilizing a 'gray market' of proxies, referred to as 'transfer stations,' to circumvent U.S. AI companies' geographic restrictions, breach terms of use, evade safeguards, and undermine traceability.
* Achieving significant cost savings by bulk procuring premium subscriptions to U.S. AI services, which are then shared across teams of developers.
Advanced distillation tactics further involve chain-of-thought (**CoT**) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks designed to detect defensive countermeasures. These methods allow China-based AI companies to significantly shorten AI development timelines and reduce financial expenditures in training frontier models.
### Threat to U.S. Technological Leadership
The agencies emphasize that these operations represent a systematic extraction of proprietary functionalities and capabilities, posing a direct threat to U.S. technological leadership in AI. The deliberate distribution of operations across multiple providers, platforms, and pathways is intended to avoid single-point detection and distill the best features of each U.S. frontier model.
Addressing this issue, the advisory states, requires a coordinated response across the entire AI ecosystem, involving effective information-sharing among the U.S. Government, private industry, and allied nations.
### Agency Recommendations for U.S. AI Companies
The authoring agencies recommend three immediate actions for U.S. AI companies:
1. **Implement Comprehensive Detection and Mitigation:** Focus on detecting anomalous and malicious prompts, accounts, networks, and behaviors. Monitor subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput patterns.
2. **Deploy Targeted Response Changes:** Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs for companies conducting industrial-scale distillation campaigns.
3. **Establish Cross-Organization Intelligence Sharing:** Correlate activity across model providers, cloud platforms, and API aggregators to uncover distributed campaigns.
### Deep Dive into Attributed Campaigns
Since late 2024, the scale and sophistication of these campaigns indicate that distillation is not merely supplemental but a critical core of these companies' AI model development.
Likely with the knowledge of the Chinese government, this sector has adopted a comprehensive distillation strategy to bridge technological and performance gaps with U.S. frontier AI models. They leverage 'transfer stations' to bypass regional restrictions and terms of use.
#### DeepSeek's Targeted Extraction
**DeepSeek** has been conducting an organized distillation campaign since at least late 2024 to generate synthetic training data for models like **R1**, released in early 2025. The company targeted specific knowledge domains to extract proprietary functionality and reasoning capabilities, significantly reducing compute and research costs. Publicly quoted training costs of $5.6M for **DeepSeek** are considered misleading, as they do not account for the true cost of data acquired through extensive malicious distillation.
Between late 2024 and mid-2025, **DeepSeek** reportedly distilled specialized training data and capabilities from various U.S. models for its **R1** and **V3** models, including:
* **Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1**
* **Gemini 2.5 Pro Preview, Gemini 2.5 Flash Preview**
* **GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, GPT-5**
* **Grok 4**
The specific knowledge and capabilities extracted included legal specialization optimization, API rule-driven tasks, writing using CoT drafts, agentic functions, question and answer optimization, coach/assistant capabilities, functional creation optimization, supervised fine-tuning (**SFT**) optimization, and creative/occupational writing optimization.
#### Moonshot AI's Widespread Distillation
**Moonshot AI** has conducted a widespread distillation campaign since at least mid-2025. Notably, it extracted significant **Claude Fable 5** data to train its **Kimi-K3** model and **GPT-4o** data for its **Kimi-K2** model. The company used a broad range of models to distill SFT optimization, reinforcement learning (**RL**), software engineering, and math capabilities, including:
* **Claude Opus 4.1, Claude Sonnet 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.5 Thinking, Claude Fable 5**
* **GPT-oss-20b, GPT-3, GPT-4o, GPT-4o mini, GPT-5, GPT-5 Codex, GPT-5 Pro**
* **Gemini 2.5 Flash, Gemini 2.5 Flash-Image, Gemini 2.5 Pro**
* **Nano Banana**
* **Grok Code Fast-1**
#### Other Companies Engaging in Distillation
**Alibaba**, **MiniMax**, **StepFun**, and **Z.AI** have also leveraged distillation techniques to build their AI models. In late 2025, **Alibaba** distilled **Claude-4**, **Claude Opus**, **Claude Sonnet**, and **GPT-5** to improve its AI models' software engineering skills, customer service dialogue functionality, image/character creation, and integration of RL, SFT, and distillation capabilities.
**MiniMax**, in late 2025, distilled CoT reasoning, RL, SFT, and software engineering capabilities for its **M2** model from **Claude Code**, **Claude Sonnet 4**, **Claude Opus**, **Gemini 1**, **Gemini 2.5 Pro**, and **Gemini 3 Pro**. **MiniMax** even used prompt injections to try to trick **Claude Code** into believing it was a **MiniMax** product.
Between late 2025 and early 2026, **StepFun** distilled data from **Claude Opus 4.1** and **4.5**, **Claude Sonnet 4.5**, **Claude Haiku 4.5**, **GPT-5 Mini**, **GPT-5 Pro**, **GPT-5.1**, **GPT-5.1 Codex**, and **GPT-5.2** to improve its **Step 4** model's coding and agentic functions. By mid-2026, **Z.AI** had distilled billions of tokens of **GPT-5.5** data and **Claude Opus 4.8** data to develop the CoT reasoning capabilities of its model.