OpenAI says Moonshot AI-linked individuals ran a coordinated campaign to extract its models' hidden reasoning
On Sept 30, 2026 OpenAI published "Disrupting a coordinated model-distillation campaign". It says a campaign starting July 1 manipulated model interactions so that protected (encrypted) reasoning was reproduced in visible form at scale, and it attributes a core cluster to individuals associated with Moonshot AI, maker of Kimi. OpenAI banned accounts, closed a replay vulnerability and shared findings via the Frontier Model Forum and government channels.
Key facts
- Campaign began July 1, 2026; spikes of ~16,000 requests from 4,000+ users on July 24–25; a broader cluster of 15,000+ users fully disrupted by July 28
- Method: copying encrypted reasoning from one conversation and asking the model in another conversation to decrypt/transcribe it; no encryption break, database compromise or access to stored user conversations
- Attribution: a core cluster to 'individuals associated with Moonshot AI'; OpenAI says it is unclear whether all activity traces to one actor
- Response: account bans, stronger signup and infrastructure controls, a fix for the encrypted-reasoning replay vulnerability, sharing via the Frontier Model Forum and government
- OpenAI's Caroline Zier: 'Our concern is about violation of our terms of service, not open models or legitimate distillation.' (The Next Web)
- Earlier: Anthropic had accused Moonshot (with other Chinese labs) of distillation; China's CAC is separately probing DeepSeek and Moonshot
- Outside researchers (the Stolen Thoughts reasoning-extraction group) say OpenAI confirmed the attack paths they reported, and that Astra/GPT-6.1 Sol reasoning was still extractable via third-party cloud API providers two months after first disclosure to OpenAI and Anthropic (their X posts)
What happened
OpenAI's threat report describes "adversarial distillation": the systematic, unauthorized use of one model's outputs or reasoning to train or improve another. The operators did not break OpenAI's encryption. They exploited how encrypted reasoning items could be replayed into new conversations and asked the model to reveal them.
Why it matters
Hidden chains of thought are a main competitive asset and a safety-monitoring surface. This is the first time OpenAI has publicly named a specific Chinese lab in a distillation case. It adds to US government and Anthropic claims and lands while Moonshot pursues a Hong Kong IPO.
Unverified: which OpenAI models were targeted (secondary reports mention reasoning models generally), and any response from Moonshot AI (none found as of Sept 30 evening).
Changelog
- 2026-09-30: created (evening sweep run)
- 2026-09-30: added the Stolen Thoughts researchers posts
Related posts (3)
- Joachim Schaeffer: 'We stole reasoning. Again.' Update to the reasoning-extraction paper original ↗ Joachim Schaeffer @JSchaeff3r · x · 2026-09-30
First-hand thread (17K views) by researchers whose work OpenAI cited in its distillation disclosure. - Joachim Schaeffer: OpenAI's distillation disclosure cites our reasoning-extraction work original ↗ Joachim Schaeffer @JSchaeff3r · x · 2026-09-30
First-hand: the researchers say OpenAI confirmed the attack paths they reported. - Alexander Panfilov: reasoning from Astra/Sol-6.1 still extractable via third-party API providers original ↗ Alexander Panfilov @kotekjedi_ml · x · 2026-09-30
First-hand researcher claim that mitigations are incomplete two months after the first attack release.
Related events
- Moonshot AI releases Kimi K3, a 2.8T-parameter open-weights multimodal model ★★★★★
- China's cyberspace regulator probes DeepSeek and Moonshot over possible data leaks to Anthropic via Claude ★★★
- Anthropic threat intelligence report: AI-orchestrated cyberattacks and distillation by Chinese labs ★★★
- Moonshot AI (Kimi) confidentially files for a Hong Kong IPO, reportedly seeking about $3B ★★★
Sources (7)
- officialOpenAI: Disrupting a coordinated model-distillation campaign
- pressBloomberg: OpenAI blames Moonshot for mass data extraction on its AI models
- pressThe Next Web: OpenAI says Moonshot-linked users tried to extract its AI reasoning
- pressWccftech: Moonshot tried to crack OpenAI's encrypted reasoning through 16,000 requests
- discussionTechmeme discussion
- discussionStolen Thoughts researchers on X: OpenAI cites reasoning-extraction work
- discussionkotekjedi_ml on X: reasoning-extraction audit
id: 2026-09-30-openai-moonshot-distillation-campaign · updated 2026-09-30 · open in the interactive timeline