US Agencies Accuse Chinese AI Companies of Industrial-Scale Theft Through Model Distillation

Reviewed byNidhi Govil

12 Sources

Share

Three US intelligence agencies have accused six Chinese AI companies of conducting aggressive distillation campaigns to extract proprietary capabilities from American frontier models. The NSA, FBI, and CISA claim DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens from models like GPT, Claude, Gemini, and Grok since late 2024.

US Agencies Issue Joint Advisory on Chinese AI Distillation

The National Security Agency (NSA), Federal Bureau of Investigation (FBI), and Cybersecurity and Infrastructure Security Agency (CISA) have issued a joint cybersecurity advisory accusing Chinese AI companies of conducting "aggressive, malicious, and targeted distillation activities at an industrial scale."

1

2

The agencies claim these campaigns extract restricted proprietary functionalities and capabilities from American frontier AI models, asserting that model distillation forms "the core - not merely a supplement" of their AI development strategy.

1

Source: The Next Web

Source: The Next Web

Six Chinese Firms Named in Systematic Extraction Campaign

US agencies specifically named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as companies that extracted billions of tokens across millions of exchanges from US frontier AI models since at least late 2024.

2

4

The advisory alleges these Chinese AI companies targeted variants of Anthropic Claude, OpenAI GPT, Google Gemini, and SpaceXAI Grok, likely with Chinese government awareness.

2

According to the agencies, DeepSeek conducted organized campaigns between late 2024 and mid-2025 targeting reasoning capabilities and domain-specific functions to train its R1 and V3 models.

2

Moonshot AI allegedly extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data for Kimi-K2.

2

4

Source: TechRadar

Source: TechRadar

How Industrial-Scale Distillation Campaigns Operate

Model distillation involves a smaller model querying a larger model to learn response patterns, improving performance over time without expensive training.

1

While distillation is recognized as a legitimate technique in AI research, commercial model providers generally prohibit such activity through their terms of use to protect substantial investments in technology and training.

1

5

The agencies claim China-based AI companies route distillation requests through multiple pathways including native application programming interfaces (APIs), remote cloud providers, and third-party aggregators that automatically obfuscate user metadata to avoid detection.

1

These operations also leverage a gray market of proxies known as "transfer stations" to bypass geographic restrictions and breach terms of use.

1

Advanced Tactics Include Chain-of-Thought Reasoning Extraction

The joint advisory details sophisticated methods employed in these industrial-scale distillation campaigns. Advanced tactics include chain-of-thought reasoning extraction, automated failover between pathways during blocking attempts, and quality evaluation frameworks to detect defensive countermeasures.

2

Alibaba allegedly leveraged industrial-scale distillation to improve its Qwen family of AI models by distilling Claude-4, Claude Opus, Claude Sonnet, and GPT-5 to enhance software engineering skills and customer service dialogue functionality in late 2025.

1

2

MiniMax distilled chain-of-thought reasoning and software engineering capabilities from Claude Code, Claude Sonnet 4, Claude Opus, Gemini 1, Gemini 2.5 Pro, and Gemini 3 Pro to improve its M2 model.

2

DeepSeek Claims and Market Impact

The advisory specifically accuses DeepSeek of using distillation to generate synthetic data used to train its models, making its claim of creating them with trivial quantities of computing power false.

1

This is particularly significant because DeepSeek's claims about low computational requirements panicked investors who worried that billions pumped into AI infrastructure might not be needed.

1

The agencies state that Chinese AI companies conducting industrial-scale distillation see significantly shorter AI development timelines and reduced financial expenditures in training frontier models.

2

Source: Gizmodo

Source: Gizmodo

Recommended Defenses Against Malicious Distillation

The NSA, CISA, and FBI recommend AI companies implement comprehensive detection measures to identify distillation attacks.

1

Immediate maximum usage from new accounts serves as one indicator of adverse action.

1

The agencies suggest subtly altering responses for suspected malicious distillation attempts to reduce payoffs to companies conducting industrial-scale distillation campaigns.

1

2

They also recommend correlating activity across different model providers, cloud platforms, and API aggregators to reveal distributed distillation campaigns.

1

National Security Risks and Geopolitical Context

The agencies warn that illicitly distilled models lack necessary safeguards, creating significant national security risks.

5

US frontier AI models are officially restricted and not offered in China, forcing Chinese developers to rely on alternative access methods like virtual private networks, obfuscated accounts, and automated agents to bypass geographic controls.

2

The timing of this advisory aligns with upcoming US-China talks on AI security, with Treasury Secretary Scott Bessent scheduled to meet Chinese officials this month.

5

This is not the first time Chinese firms have faced such accusations—earlier this year, Anthropic identified industrial-scale campaigns conducted by DeepSeek, Moonshot AI, and MiniMax to illicitly extract proprietary functionalities.

2

China has responded with counter-accusations that US companies distill Chinese models, plus veiled threats of retaliation against any US bans flowing from these allegations.

1

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved