OpenAI Accuses Moonshot AI of Coordinated Campaign to Extract Protected Model Reasoning

Reviewed byNidhi Govil

8 Sources

Share

OpenAI disrupted a coordinated distillation campaign in July that peaked at 16,000 extraction attempts over two days. The company traced core activity to individuals associated with Chinese AI startup Moonshot AI, developer of the Kimi model, raising fresh concerns about AI model theft and national security risks.

OpenAI has accused individuals linked to Moonshot AI of orchestrating a coordinated campaign to extract protected reasoning from its AI models, marking the latest escalation in mounting tensions between US and Chinese artificial intelligence companies over data extraction and AI security practices.

Campaign Peaked at 16,000 Requests Over Two Days

Source: Hacker News

Source: Hacker News

The adversarial distillation campaign began on July 1 at low volume before surging dramatically on July 24 and 25, when OpenAI detected 16,000 requests using extraction patterns from over 4,000 users

1

4

. Further investigation revealed related prompt-pattern activity across more than 15,000 users before OpenAI fully disrupted the coordinated campaign on July 28

2

3

. While the company could not definitively link all operators to a single actor, it traced a core cluster of the activity to individuals associated with Moonshot AI, the Chinese company behind the Kimi model

5

.

How the Distillation Attack Exploited Model Interactions

The operators did not break OpenAI's encryption, compromise databases, or gain direct access to stored user conversations. Instead, they manipulated model interactions to extract protected reasoning through a sophisticated technique

2

. According to OpenAI, attackers copied encrypted reasoning from one conversation, then prompted a model in another conversation to decrypt and transcribe it

5

. This adversarial distillation represents the systematic and unauthorized use of one model's outputs to help train, reproduce, or improve another model without preserving the same safety guardrails

3

.

Researchers from MATS Research, ELLIS Institute Tübingen, and Synk identified an architectural vulnerability in August 2026 that made encrypted reasoning traces fully compatible across different sessions, users, and models within a provider's ecosystem

3

. By injecting an encrypted reasoning trace into a weaker, less safeguarded model from the same provider, attackers could force it to decode and output the trace in plaintext without directly jailbreaking the more capable model.

National Security Risks and Safety Concerns Mount

OpenAI warned that adversarial distillation poses significant safety and national security risks. Extracted reasoning could be used to train another model without preserving safeguards applied to the original model's user-facing outputs

3

. At scale, distillation can accelerate the transfer of advanced capabilities without requiring the same investment in safety, concerns that become heightened as models gain capabilities in dual-use domains

2

. Caroline Zier, who leads strategic national security policy initiatives at OpenAI, emphasized the company's concern centers on terms of service violation rather than open models or legitimate distillation

5

.

Moonshot AI Faces Multiple Accusations

Source: The Next Web

Source: The Next Web

This marks the first time OpenAI has directly accused Moonshot AI of data extraction attempts, though the Chinese startup faces similar allegations from other major AI companies

5

. Anthropic accused Moonshot AI last month of stealthily relaying customer requests to Claude instead of processing them using Kimi, then displaying Claude's responses back to users while retaining a subset of exchanges to train its chain-of-thought model

3

. US President Donald Trump's Assistant for Science and Technology Michael Kratsios also accused Moonshot AI in late July of creating its Kimi K3 model by distilling Anthropic's Fable

2

. Moonshot did not immediately respond to requests for comment on the allegations

4

.

Industry Response and Enhanced Protections

Source: The Register

Source: The Register

OpenAI has deployed multiple countermeasures following the distillation attack. The company banned fraudulent accounts, tightened signup and infrastructure controls, and expanded monitoring efforts

2

. It closed a pathway that allowed someone who already possessed another user's encrypted reasoning to replay and recover its contents, and added checks to detect and hold streamed output that might expose reasoning

3

. OpenAI shared investigation details with other AI firms through the Frontier Model Forum and government information-sharing programs

2

4

.

Anthropic's Claude Opus 5.5 model, released recently, includes a defense against distillation called preserved thinking that was introduced with Fable 5.1

2

. However, not everyone accepts these claims at face value. David Sacks, Trump's former AI czar, has characterized such reports as attempts to pressure the US into banning rival open models

5

. Meanwhile, Chinese regulators are investigating both DeepSeek and Moonshot AI domestically, adding another layer of scrutiny to the companies' operations

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved