MikhbarMIKHBAR
Artificial Intelligence

OpenAI Links Moonshot AI to Model Reasoning Extraction Campaign

OpenAI has revealed that individuals associated with China-based Moonshot AI were at the center of a coordinated campaign in July to extract protected reasoning data from its models.

OpenAI Links Moonshot AI to Model Reasoning Extraction Campaign

Coordinated Campaign Targeted Protected Reasoning

OpenAI reported in a company blog post that people associated with China-based Moonshot AI formed the core of an effort to extract hidden reasoning from its models. According to details shared via a company blog post, the coordinated campaign to pull protected reasoning was consistent with adversarial distillation and occurred throughout July.

Activity kicked off on July 1st at low volumes before surging into high-volume spikes on July 24 and 25. During those two days, OpenAI logged 16,000 requests utilizing an extraction pattern originating from over 4,000 users. Although the campaign attempted extraction, OpenAI noted it was not necessarily successful, and the operation was fully disrupted by July 28. The company did not clarify exact success rates, targeted models, or the precise count of Moonshot-linked users.

Understanding Protected Reasoning and Adversarial Distillation

OpenAI defines protected reasoning as a model's internal record for working through a task. Adversarial distillation is described as the systematic and unauthorized use of one model's outputs or reasoning to train or improve another model. This sensitive data remains encrypted to conceal the chain of thought and is handed to the client as an encrypted block, which the client then sends back with each request so the provider avoids storing it directly.

During the July incidents, operators attempted various methods, including taking the encrypted reasoning from one conversation and asking a model in another conversation to decrypt it. OpenAI confirmed that the encryption itself was not broken, no direct access to stored user conversations occurred, and no database was compromised during the campaign.

Mitigation Steps and Security Fixes

To counter these vulnerabilities, OpenAI closed a pathway that allowed individuals possessing another user's encrypted reasoning to replay it and recover its contents. Additionally, the company introduced checks to detect and hold streamed output that could potentially expose reasoning traces.

Furthermore, OpenAI strengthened protections for hidden reasoning across different users, workspaces, organizations, and model families. The company also worked alongside third-party providers to disrupt accounts whose activity routed through their services.

Independent Disclosure and Industry-Wide Threats

Independent security researchers also brought related cross-model and conversation-compaction vulnerabilities to OpenAI through responsible disclosure, prompting confirmation of real attack paths. A research paper titled 'Stealing Reasoning Traces from Proprietary LLM APIs' demonstrated how feeding a frontier model's encrypted reasoning to a weaker model caused it to write out the data in plain text.

The broader industry faces similar threats. Anthropic recently detailed a period where Moonshot relayed nearly 300,000 customer requests using a proxy network of fraudulent accounts. In related developments, Moonshot previously released a major model, as covered in reports on Moonshot AI, while denying that its Kimi K3 system relied on distillation.

Ongoing Industry Vigilance

OpenAI's next steps involve ensuring partner-hosted deployments match the security protections of first-party tools, alongside deploying additional checks against tool-output attacks. The company shared its findings with the Frontier Model Forum and government information-sharing channels.

As frontier models continue to improve, OpenAI expects distillation attempts to grow more sophisticated. Attackers looking for cheaper ways to mimic advanced systems will likely persist, making ongoing defense efforts critical for the artificial intelligence industry.

Sources

  • Tom's HardwareOpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning

Continue chronologically

Related entity coverage