OpenAI Disrupts Campaign Extracting Model Reasoning
OpenAI has identified and dismantled a coordinated effort to extract protected internal reasoning from its artificial intelligence models. The company attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.

Coordinated Campaign Disrupted in July
OpenAI revealed that it recently identified and disrupted a coordinated campaign designed to systematically extract protected reasoning from its artificial intelligence models. According to official disclosures published by OpenAI News, the earliest observed extraction activity began on July 1, operating at a relatively low volume before scaling rapidly later in the month.
The operational volume escalated significantly on July 24 and July 25, generating approximately 16,000 extraction requests across more than 4,000 user accounts. Further investigation by OpenAI identified related prompt-pattern activity across a broader network cluster of more than 15,000 users. The organization confirmed that it fully disrupted the activity by July 28.
Extraction Tactics and Cross-Model Attack Paths
The campaign involved adversarial distillation, which is the unauthorized and systematic capture of one model's outputs or internal processes to train, reproduce, or enhance another model. In this instance, operators targeted protected reasoning—the internal operational record used by a model to work through a task prior to delivering a final output. Accessing this concealed reasoning can expose context withheld from user-facing answers and assist outside actors in replicating frontier capabilities.
OpenAI stated that the actors did not break encryption, compromise databases, or gain direct access to stored user conversation logs. Instead, operators manipulated standard model interactions to force protected reasoning into forms visible to the requester in a scaled, coordinated manner that breached the platform's terms of service. Novel methods included copying encrypted reasoning content from one conversation and instructing a model in another interaction to decrypt and transcribe the hidden text.
Defensive analysis was assisted by external researchers, as Independent security researchers (opens in a new window) brought related cross-model and conversation-compaction vulnerabilities to OpenAI's attention through responsible disclosure. OpenAI confirmed the validity of these attack paths, using the findings to better understand the broader attack class and accelerate defensive controls.

Attribution to Individuals Linked to Moonshot AI
While OpenAI noted that it remains unclear whether all observed operators belonged to a single entity, the company formally attributed a core cluster of the extraction activity to individuals associated with Moonshot AI, the Chinese AI developer responsible for the Kimi model.
OpenAI emphasized that adversarial distillation poses meaningful safety and national security risks across the broader technology industry. Extracted reasoning can be utilized to train external models without transferring the alignment and safety controls built into the original developer's public outputs. At scale, distillation enables cheaper capability transfers, which becomes particularly concerning as frontier models develop advanced skills in dual-use domains.

Technical Controls and Account Enforcement
To counter the campaign, OpenAI implemented a combination of account actions, technical safeguards, and partner cooperation. The company restricted or banned fraudulent accounts, upgraded infrastructure and signup controls, and expanded monitoring across connected networks.
Engineering teams closed a specific pathway that allowed an actor in possession of another user's encrypted reasoning to replay the data and recover its contents. OpenAI also added checks to evaluate and hold output streams that could expose reasoning artifacts. Where activity moved through third-party platforms, OpenAI worked alongside those service providers to disrupt the associated accounts as part of its overarching focus on platform Security.
Ecosystem Coordination and Long-Term Outlook
Because vulnerabilities surrounding portable or replayable reasoning artifacts are not isolated to a single provider, OpenAI shared its findings through the Frontier Model Forum and relevant government information-sharing channels to help peer developers strengthen collective defenses.
As technical capabilities expand throughout 2026, OpenAI expects adversarial distillation tactics to grow increasingly complex as actors search for lower-cost methods to mimic frontier systems. The company stated that ongoing protective work will focus on improving tool defenses, expanding classifier coverage, refining model refusals, and deploying matching controls across partner hosted and cloud environments.
Sources
- OpenAI NewsDisrupting a coordinated model-distillation campaign