Announced 30 Sept 2026 · Sources checked
What did OpenAI say it found?
OpenAI’s 30 September post says the company identified a coordinated campaign designed to extract protected reasoning, with the earliest observed activity in the first week of July. It calls the pattern adversarial distillation: systematically using one model’s outputs or reasoning, without authorisation, to help train, reproduce or improve another model.
Protected reasoning, in OpenAI’s wording, is the model’s internal record of working through a task. Extracting it can reveal material withheld from the final answer and, the company argues, help others copy capabilities. The operators did not break encryption, compromise a database, or gain direct access to stored user conversations. They manipulated interactions so that hidden reasoning could be reproduced in a form the requester could see, at a scale OpenAI says violated its terms of service.
OpenAI says the method is not unique to its models and that it shared information through the Frontier Model Forum. Independent security researchers, unnamed in the post, also disclosed related cross-model and conversation-compaction paths; OpenAI says it confirmed those paths were real. That is a vendor security narrative, not an independent incident report. Read it the way our guide to company announcements recommends: start from what the primary page actually states.
What numbers did OpenAI publish?
Activity began on 1 July at low volume. OpenAI then reports high-volume spikes on 24 and 25 July: 16,000 requests using a relevant extraction pattern, from more than 4,000 users. A footnote says those figures describe attempted, not necessarily successful, extractions. Further investigation, the company writes, identified related prompt-pattern activity across a cluster of more than 15,000 users, which it says it fully disrupted by 28 July.
The post also describes one concrete technique: copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden content. We are not expanding that into a how-to. OpenAI says it later closed a pathway that let someone who already possessed another user’s encrypted reasoning replay it and recover the contents, and added checks that can hold streamed output that might expose reasoning.
Those controls sit next to ordinary account and network enforcement. They are closer to hiding an internal scratchpad than to classic prompt injection, where the malicious instruction arrives in the prompt rather than being pulled out of a hidden trace.
| Date | What OpenAI reports |
|---|---|
| 1 July 2026 | Earliest observed activity; initially low volume |
| 24–25 July | 16,000 extraction-pattern requests from >4,000 users (attempts) |
| By 28 July | Related cluster of >15,000 users described as fully disrupted |
| 30 September | Public post; further mitigation work still underway |
What does the Moonshot attribution actually say?
The attribution section is two sentences. OpenAI writes that it is unclear whether all operators in the period came from a single actor. It then attributes “a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.” It names no person, publishes no indicator, and does not say Moonshot as a company ordered the work.
The post also does not say that any extracted reasoning was used to train, improve or evaluate a Kimi model. Distillation in the abstract is a training method; this disclosure is about attempted extraction. Those are different claims. Anyone repeating “OpenAI proved Kimi was distilled from ChatGPT” is going beyond the page.
As of 7 October 2026 we have not found a public Moonshot statement that answers this specific post. Silence is not confirmation. Hidden-model and unnamed-associate stories are a standing pattern in the industry; see our note on stealth models.
Why does OpenAI say this matters?
The company argues that extracted reasoning could train another model without the safeguards applied to the original user-facing outputs, and that at scale distillation can move advanced capabilities without the same safety investment. It says the concern grows as models pick up dual-use skills. That is OpenAI’s risk case. This publication has not audited whether the July traffic produced usable training data.
The same post says similar techniques may affect other advanced systems, which is why OpenAI briefed the Frontier Model Forum and government information-sharing channels. Buyers who already ask vendors who can see logs and traces should add hidden-reasoning retention to that list; our API data-retention guide is the neighbouring question.
How did OpenAI say it responded?
The published response has three layers. Account enforcement: banned or restricted fraudulent accounts, tighter signup and infrastructure controls, and wider monitoring of related networks. Technical controls: stronger protection for hidden reasoning across users, workspaces, organisations and model families, plus the replay closure and stream-hold checks already mentioned. When activity moved through third-party services, OpenAI says it worked with those providers to disrupt the accounts.
The unfinished work, in OpenAI’s own list, is partner-hosted deployments, tool-output attacks that are not ordinary visible text, classifier coverage, model refusals, and propagating controls to cloud partners. The company expects extraction attempts to get more sophisticated as frontier models improve.
What this post does not settle
It does not quantify successful extractions. It does not identify the associated individuals. It does not publish the telemetry behind the Moonshot cluster. It does not show that Kimi, or any named checkpoint, contains OpenAI reasoning. Additional mitigation, OpenAI says, is still underway.
Readers should keep the three facts that are actually on the page: a July campaign aimed at hidden reasoning, attempt counts with a 28 July disruption date, and an attribution to people associated with Moonshot that remains OpenAI’s claim until someone else tests it.
Common questions
Did OpenAI say Moonshot trained Kimi on stolen reasoning?
No. It attributes a core cluster of extraction activity to individuals associated with Moonshot. It does not say extracted material was used in a Kimi training run.
Do the 16,000 requests mean 16,000 successful leaks?
No. OpenAI’s footnote says the spike figures describe attempted, not necessarily successful, extractions.
Has Moonshot replied?
We have not found a public statement from Moonshot AI addressing this 30 September post as of 7 October 2026. That is an observation about the public record, not a finding about the company.
What to remember
OpenAI published a July extraction campaign, attempt-level counts, and an attribution to people associated with Moonshot. Hold those three claims apart. Do not upgrade them into proof that a Kimi model was trained on OpenAI reasoning.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





