Disrupting a coordinated model-distillation campaign
Summary
OpenAI identified and disrupted a coordinated campaign beginning July 1 to extract protected reasoning from its models through adversarial distillation. Operators did not breach encryption or databases; instead, they manipulated model interactions to reproduce hidden reasoning in visible forms at scale. High-volume spikes on July 24-25 involved 16,000 extraction requests from over 4,000 users, with related prompt-pattern activity across 15,000+ users fully disrupted by July 28. OpenAI attributed a core cluster to individuals associated with Moonshot AI (developer of Kimi) and shared findings through the Frontier Model Forum. Responses included account bans, strengthened signup controls, closing reasoning-replay vulnerabilities, and third-party coordination. OpenAI warns that adversarial distillation poses safety and national security risks, and plans to continue improving extraction protections, detection, and threat-information sharing.
(Source:OpenAI)