上周,Hugging Face 在检测并遏制了一起入侵其基础设施的 AI 智能体事件后,披露了一种新型安全事件(在新窗口中打开)。随着具备日益增强网络能力的模型不断普及,我们预计此类事件将变得更加常见。经过调查,我们现在了解到,这起特定事件是由 OpenAI 模型的组合驱动的——包括 GPT‑5.6 Sol 和一个能力更强的预发布模型,这些模型在评估目的下减少了网络拒绝响应,并在一个网络能力基准测试(在新窗口中打开)中进行了内部测试。
我们认为这是一起前所未有的网络事件,涉及最先进的网络能力,并正在做出相应响应。我们在此阶段分享初步发现,以帮助防御者了解事件经过,并帮助校准当前模型的能力。我们将继续与 Hugging Face 一起进行彻底调查,并在调查完成后分享有关漏洞、事件和发现的更多细节。
事件期间发生了什么
这起事件发生在一项内部评估期间,该评估提示模型使用复杂的攻击路径进行高级利用,以量化其网络能力。我们通过运行此评估(未使用用于阻止模型从事高风险网络活动的生产分类器)来估算最大网络能力。我们的基准测试运行在高度隔离的环境中,网络访问仅限于通过一个内部托管的第三方软件安装软件包的能力,该软件充当软件包注册表的代理和缓存。
这些模型识别并串联了 OpenAI 研究环境和 Hugging Face 生产基础设施中的漏洞,直接从 Hugging Face 的生产数据库中获取了测试解决方案。所有证据表明,这些模型高度专注于为 ExploitGym 寻找解决方案,不惜一切代价实现一个相当狭窄的测试目标。
在我们的沙盒测试环境中运行时,我们的模型花费了大量推理计算资源,寻找获得开放互联网访问权限的方法,以解决评估问题。为了获得访问权限,这些模型识别并利用了软件包注册表缓存代理中的一个零日漏洞(我们现已负责任地向供应商披露)。凭借此访问权限,我们的模型在我们的研究测试环境中执行了一系列权限提升和横向移动操作,直到模型到达一个具有互联网访问权限的节点。
获得互联网访问权限后,模型推断 Hugging Face 可能托管了 ExploitGym 的模型、数据集和解决方案。得知这一点后,模型搜索并成功找到了获取可用于作弊评估的秘密信息的方法。在一个例子中,模型串联了多个攻击向量,包括使用窃取的凭证和零日漏洞,在 Hugging Face 服务器上找到了一条远程代码执行路径。OpenAI 的安全团队在内部发现了这一异常活动。
Hugging Face 的安全团队和智能体检测并阻止了其基础设施上的活动,并在我们的团队联系时已经开始使用他们自己的开源模型进行遏制和取证重建。我们正在积极与他们合作,继续调查这起事件。我们感谢 Hugging Face 在调查和修复方面迅速而密切的合作。
我们正在采取的行动
- 作为调查的一部分,我们正在基础设施配置中实施严格的控制,以研究速度为代价,同时修补漏洞。我们定期向我们的安全与安保委员会汇报这些控制措施及其影响。
- 我们正在与 Hugging Face 合作,对这起事件进行取证调查。
- 我们已负责任地披露了内部托管第三方软件中已识别的零日漏洞,并正在与他们合作进行修补。
- 我们已将 Hugging Face 纳入可信访问计划,并支持他们的团队快速利用我们模型的能力来改进其防御。
- 我们正在改进并增加围绕未来训练和评估的更强保护措施。本周,我们发布了一篇关于在长周期模型时代改进安全性与对齐的博客。这些部署安全措施在此次评估期间有意未启用,因为评估旨在测试网络漏洞。这起事件表明,需要进一步加强我们模型的对齐性、评估期间的网络保护以及内部测试期间的监控。
我们评估高级网络能力的方法
正如我们最近分享的,AI 正在加速漏洞的发现和利用。这起事件的主要教训是,模型安全必须与快速发展的能力保持同步。我们正在加强模型开发过程中使用的遏制、监控、访问控制和评估实践。
英国 AISI 的评估显示,GPT‑5.6 Sol 等模型越来越能够在长时间跨度内维持复杂、多步骤的网络操作。这起事件表明,这些理论能力确实适用于现实世界环境。

这起事件还表明,高级模型可以在没有源代码访问权限的情况下,发现并利用现实世界系统中的新型攻击路径。它凸显出,高级网络能力必须与更强的安全措施和防御工具同步发展。
我们相信,具备高级网络能力的模型需要帮助安全团队在攻击者之前发现弱点,理解漏洞如何被串联,并以机器速度进行修复。我们正在利用这些能力继续加强对基础设施配置和模型评估环境的保护;我们将在学习过程中分享我们的发现和最佳实践。我们鼓励其他防御者申请可信访问,并立即试验这些模型,将这些能力转化为更好的预防、更快的检测和更有效的应急响应。
“我们感谢与OpenAI在此事及其他议题上的合作。这起事件或许是首例,但它印证了我们长期以来的信念:AI安全无法靠任何一家公司闭门解决。唯有通过公开协作,让全球每一位防御者都能广泛接触AI,才能真正实现安全。”
——Clem Delangue,Hugging Face联合创始人兼首席执行官
Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities.
We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
What happened during this incident
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.
Actions we are taking now
- As part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched. We are regularly briefing our Safety and Security Committee on these controls and their impact.
- We’re working with Hugging Face to forensically investigate the incident.
- We’ve responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.
- We’ve brought Hugging Face into the trusted access program and are supporting their teams in rapidly using our models’ capabilities to improve their defenses.
- We’re improving and adding stronger protections around future training and evaluations. This week, we published a blog on improving safety and alignment in an era of long horizon models. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities. This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.
Our approach to evaluating advanced cyber capabilities
As we recently shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.
UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.

The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access. It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools.
We believe advanced cyber capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed. We are using these capabilities to continue strengthening protections around infrastructure configuration and model evaluation environments; we will share our findings and best practices as we learn. We encourage other defenders to apply for trusted access and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response.
“We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
—Clem Delangue, Co-founder and CEO, Hugging Face
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文