2026年9月30日
安全
破坏协调一致的大模型蒸馏活动
分享
我们观察到的情况
我们对归因的评估
为何这很重要
我们的应对措施
未来展望
我们最近发现并破坏了一场旨在从我们的模型中提取受保护推理过程的协调行动,最早观察到的活动发生在7月的第一周。这种行为与对抗性蒸馏(adversarial distillation)一致:即系统性地、未经授权地使用一个模型的输出或推理过程来帮助训练、复制或改进另一个模型。受保护的推理过程是模型在处理任务时内部的记录;提取这些信息可能会揭示最终答案中未包含的信息,并帮助他人复制该模型的能力。
操作者并未破解我们的加密系统,也未入侵数据库或直接访问存储的用户对话。相反,他们操纵了与模型的交互方式,使得受保护的推理过程能够以请求者可见的形式被复制,这种协调且规模化的行为违反了我们服务条款的规定。这种操纵并非OpenAI模型独有的漏洞,我们已通过前沿模型论坛(Frontier Model Forum)向行业合作伙伴分享了相关信息,以加强针对对抗性蒸馏的集体防御能力。
在发布本报告之前,我们调查了事件的范围和潜在影响,部署了自身的缓解措施,并与研究人员及行业合作伙伴进行了分享并征求反馈,以确保针对此类攻击的保护措施到位。额外的缓解工作和调查仍在继续。我们相信,现在分享我们所获得的经验将有助于更广泛的生态系统加强其防御能力。
我们观察到的情况
我们发现操作者尝试以新颖的方式提取受保护的推理过程,包括从一次对话中复制加密的推理内容,并在另一次对话中要求模型解密并转录隐藏的推理内容。
独立安全研究人员(在新窗口打开)也通过负责任披露的方式将相关的跨模型和对话压缩漏洞告知了我们。我们调查了他们的发现,并确认了他们所识别的攻击路径是真实存在的。他们的工作帮助我们理解更广泛的攻击类别,并加速了缓解措施的实施。
该活动始于7月1日,起初规模较小,直到我们在7月24日和25日观察到高频激增,涉及超过4,000名用户发出的16,000次请求[1],使用了相关的提取模式。进一步的调查确定了在超过15,000名用户的集群中存在的类似提示词模式活动,我们已于7月28日完全破坏了这一活动。
该活动随时间演变,这进一步证实了对抗性蒸馏是一个更广泛的安全挑战,需要分层且自适应的防御措施。
我们对归因的评估
目前尚不清楚我们在相关时间段内观察到的所有操作者是否都来自同一个行为体。然而,我们将核心集群的活动归因于与Kimi开发者月之暗面(Moonshot AI)有关联的个人。
为何这很重要
对抗性蒸馏带来安全和国家安全风险。提取出的推理过程可能被用于训练另一个模型,而无需保留应用于原始模型用户端输出的安全措施。在大规模情况下,蒸馏还可以加速先进能力的转移,而无需投入同样的安全成本。随着模型在双重用途领域获得能力,这些担忧变得更加突出。
这种风险并非OpenAI独有。如上所述,类似的技术可能会影响其他先进的AI系统,使其成为需要行业协调的共同安全挑战。
我们的应对措施
我们通过账户执行、技术控制和合作伙伴协调的组合措施,遏制了此次近期的蒸馏活动。我们封禁或限制了欺诈账户,加强了注册和基础设施控制,并扩大了对相关网络的监控。
我们还加强了对跨用户、工作区、组织和模型家族的隐藏推理的保护。我们关闭了一条允许已拥有其他用户加密推理内容的人重放该推理并恢复其内容的途径,并增加了检测可能暴露推理内容的流式输出的检查机制。当相关活动通过第三方服务进行时,我们与这些提供商合作,识别并中断涉及的账户。
最后,我们通过前沿模型论坛和适当的政府信息共享渠道分享了相关发现,以便其他前沿开发者和公共部门合作伙伴能够查找类似活动并加强自身的防御。支持可移植或可重放推理产物的系统可能面临相关风险。
接下来会发生什么
随着前沿模型的进步以及行为者寻求更便宜的方式来模仿其能力,我们预计对抗性蒸馏尝试将变得更加复杂。防御此类活动需要分层控制和持续适应。
这项工作尚未完成。合作伙伴托管的部署需要与第一方服务相同的保护,而工具输出攻击需要检查普通可见文本之外的内容的保护措施。我们正在继续改进工具防御、分类器覆盖范围、模型拒绝机制,并在云合作伙伴之间推广相关控制措施。
我们的回应将继续聚焦于三个领域:加强针对提取的技术保护、更好地检测和打击协调一致的活动,以及在行业和政府部门之间深化威胁信息共享。
2026年
作者
OpenAI
脚注
1
这些数据描述的是尝试进行的提取,不一定是成功的提取。
继续阅读
查看全部
前线防御者的黎明
安全
2026年9月3日
通往Astra之路:关键能力与前沿保障措施
安全
2026年9月1日
Hugging Face事件与未来之路
安全
2026年8月26日
September 30, 2026
Security
Disrupting a coordinated model-distillation campaign
Share
What we observed
Our assessment of attribution
Why this matters
How we responded
What comes next
We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.
The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service. This manipulation is not a vulnerability unique to OpenAI’s models, and we have shared information about it with industry partners through the Frontier Model Forum in order to strengthen collective defenses against adversarial distillation.
Before publishing, we investigated the scope and potential impact, deployed our own mitigations, and shared with and took feedback from researchers and industry partners to ensure protections against this type of attack are in place. Additional mitigation and investigation work is continuing. We believe sharing what we have learned now will help the broader ecosystem strengthen its defenses.
What we observed
We saw operators attempt to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content.
Independent security researchers
(opens in a new window)
also brought related cross-model and conversation-compaction vulnerabilities to our attention through responsible disclosure. We investigated their findings and confirmed that the attack paths they identified were real. Their work helped us understand the broader attack class and accelerate mitigations.
The activity began on July 1, initially at a low volume until we observed high-volume spikes on July 24 and 25 consisting of 16,000 requests1 using a relevant extraction pattern from over 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which we fully disrupted by July 28.
The activity evolved over time, reinforcing that adversarial distillation is a broader security challenge that requires layered, adaptive defenses.
Our assessment of attribution
It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
Why this matters
Adversarial distillation poses safety and national security risks. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs. At scale, distillation can also accelerate the transfer of advanced capabilities without requiring the same investment in safety. These concerns become heightened as models gain capabilities in dual use domains.
This risk is not unique to OpenAI. As cited above, similar techniques may affect other advanced AI systems, making this a shared security challenge that requires coordination across the industry.
How we responded
We mitigated this recent distillation campaign through a combination of account enforcement, technical controls, and partner coordination. We banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks.
We also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. We closed a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents, and added checks to detect and hold streamed output that might expose reasoning. When related activity moved through third-party services, we worked with those providers to identify and disrupt the accounts involved.
Finally, we shared relevant findings through the Frontier Model Forum and appropriate government information-sharing channels so that other frontier developers and public-sector partners could look for similar activity and strengthen their own defenses. Systems that support portable or replayable reasoning artifacts may face related risks.
What comes next
We expect adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic their capabilities. Defending against this activity requires layered controls and continual adaptation.
This work is not finished. Partner-hosted deployments need the same protections as first-party services, and tool-output attacks require protections that examine more than ordinary visible text. We are continuing to improve tool defenses, classifier coverage, model refusals, and propagate relevant controls across cloud partners.
Our response will continue to focus on three areas: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing across industry and government.
2026
Author
OpenAI
Footnotes
1
These figures describe attempted, not necessarily successful, extractions.
Keep reading
View all
Daybreak for Frontline Defenders
Security
Sep 3, 2026
Path to Astra: critical capabilities and frontier safeguards
Safety
Sep 1, 2026
The Hugging Face incident and the road ahead
Security
Aug 26, 2026
首次收录 · 2026-10-01 · 12.38 分