Full Report
OpenAI on Wednesday said it identified and disrupted a coordinated distillation campaign that was designed to illicitly extract protected reasoning from its artificial intelligence (AI) models. A "core cluster of the activity," going back to the first week of July, has been attributed to individuals associated with Moonshot AI, a Chinese AI company based in Beijing. It did not cite any
Analysis Summary
# Threat Actor: Moonshot AI (Associated Individuals / GTG-16002)
## Attribution & Identity
* **Actor Identification:** Individuals associated with **Moonshot AI**, a Beijing-based Chinese artificial intelligence company.
* **Aliases:** The activity cluster is tracked under the moniker **GTG-16002**.
* **Known Associations:** Linked to the development of the "Kimi" AI model.
## Activity Summary
OpenAI identified and disrupted a coordinated "adversarial distillation" campaign beginning July 1, 2026. The activity peaked on July 24 and 25, 2026, involving approximately 16,000 illicit requests across a network of over 4,000 accounts. The campaign aimed to extract protected "reasoning traces" (Chain-of-Thought) from OpenAI’s advanced models to facilitate the development of rival AI capabilities. OpenAI fully disrupted the campaign on July 28, 2026.
## Tactics, Techniques & Procedures
* **Adversarial Distillation:** Systematic and unauthorized use of a target model's outputs to train, reproduce, or enhance a different model.
* **Model Interaction Manipulation:** Instead of direct database compromise, the actors manipulated model interactions to force protected reasoning to be reproduced in visible, plaintext forms.
* **Replay Attacks:** Exploiting a pathway that allowed the replaying of another user's encrypted reasoning to recover its contents.
* **Scalable Decryption Jailbreak:** Using an architectural vulnerability to inject encrypted reasoning traces from a high-capability model into a "weaker" model, forcing the weaker model to decode and output the trace in plaintext.
* **Coordinated Account Usage:** Utilizing a distributed network of over 15,000 user accounts to bypass rate limits and detection.
## Targeting
* **Sectors:** Artificial Intelligence, Technology, Research & Development.
* **Geography:** United States (OpenAI headquarters/infrastructure) originating from Beijing, China.
* **Victims:**
* **OpenAI:** Targeted for model reasoning extraction.
* **Anthropic:** Previously targeted by Moonshot AI via stealthy relaying of customer requests to the Claude model.
## Tools & Infrastructure
* **Infrastructure:** A coordinated network of over 15,000 fraudulent user accounts.
* **Technological Leverage:** Exploitation of architectural vulnerabilities in model ecosystem interoperability (as identified by MATS Research and ELLIS Institute).
* **Note:** Specific C2 domains or IPs were not disclosed in the report, though the activity was coordinated and scaled.
## Implications
* **National Security Risks:** The extraction of reasoning can accelerate the transfer of advanced AI capabilities to foreign entities without the associated investment in safety or alignment.
* **Intellectual Property Theft:** Distillation allows competitors to replicate a model's logic and "Chain-of-Thought" without original research.
* **Safety Circumvention:** Extracted data can be used to train models that lack the safety safeguards and filters applied to the original provider's output.
## Mitigations
* **Account Termination:** OpenAI banned all fraudulent accounts identified in the GTG-16002 cluster.
* **Pathway Closure:** Closing technical vulnerabilities that allow for the "replay" of encrypted reasoning traces.
* **Streamed Output Checks:** Implementing real-time detection to hold or block streamed outputs that inadvertently contain or expose reasoning traces.
* **Anti-Distillation Mechanisms:** Enhancing architectural safeguards to ensure encrypted traces are not interchangeable across different model tiers (e.g., preventing a "weak" model from decoding a "strong" model's trace).