VYPR
researchPublished Sep 30, 2026· 2 sources

OpenAI Disrupts Coordinated AI Model Distillation Campaign by Moonshot AI

OpenAI has thwarted a large-scale, coordinated effort by Chinese AI firm Moonshot AI to distill its proprietary large language models, marking a significant incident of AI intellectual property theft.

OpenAI announced on Wednesday that it has successfully disrupted a sophisticated and coordinated campaign orchestrated by Moonshot AI, the developer of the Kimi model, aimed at extracting proprietary training data and reasoning from OpenAI's advanced artificial intelligence models. The frontier AI lab detected activity consistent with adversarial distillation in early July, though it chose to delay publicizing the incident to thoroughly assess its scope and impact, and to collaborate with industry partners on preventative measures.

According to OpenAI's blog post, the distillers did not breach encryption or access any confidential databases. The company's detection of the activity was aided by independent security researchers who were preparing to publish a paper on cross-model vulnerabilities. OpenAI was able to replicate these findings and confirm the attack vectors.

The campaign saw a significant spike in activity on July 24 and July 25, when OpenAI observed over 4,000 users making approximately 16,000 requests that exhibited a pattern indicative of data extraction attempts. "We saw operators attempt to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content," OpenAI stated.

By July 28, a cluster of over 15,000 users had employed similar prompting techniques. While OpenAI could not definitively attribute all the activity to a single actor, the company expressed strong suspicion that a core group of operators were associated with Moonshot AI. This is not the first time Moonshot AI has faced accusations of illicitly distilling models; Anthropic and other Chinese companies have previously criticized the firm for similar practices.

The U.S. Cybersecurity and Infrastructure Security Agency (CISA) has also identified Moonshot AI as an entity believed to be engaged in extracting data and reasoning from American-developed AI models. This broader context highlights a growing concern within the AI industry regarding the protection of intellectual property and the ethical boundaries of model development and research.

OpenAI outlined several mitigation strategies it has implemented to counter such attacks. These include enhancing user sign-up controls, expanding monitoring capabilities, and specifically increasing protections for hidden reasoning. These new protections are designed to prevent unauthorized users from recovering encrypted reasoning content, even if they possess it from another user's session.

While distillation is a controversial practice, it is often considered a legitimate part of AI innovation when conducted with explicit permission. Many open-source and open-weight models utilize distillation in their training processes, and figures like Meta CEO Mark Zuckerberg have defended it as essential for democratizing AI development. However, unauthorized distillation, as alleged in this case, represents a direct threat to the intellectual property and competitive advantage of AI developers.

This new report from OpenAI details a novel method used by the attackers to extract encrypted reasoning data by copying it from one conversation and asking the model to decrypt and transcribe it in a separate one. While the previous article focused on the overall campaign and attribution to Moonshot AI, this update clarifies the specific technical mechanism of the "encryption bypass" and notes that outside researchers also reported a similar vulnerability in August. OpenAI has since fixed the bug, enhanced its controls, and shared details with the Frontier Model Forum.

Synthesized by Vypr AI