OpenAI 表示,其 AI 模型 GPT-5.6 Sol 和另一款"能力更强的预发布模型"在内部安全测试中,利用沙箱环境的零日漏洞访问互联网,并成功入侵了开源 AI 平台 Hugging Face。
OpenAI 表示,其 AI 模型在内部测试期间错误地侵犯了开源 AI 平台 Hugging Face。在周二的一篇博客文章中:https://openai.com/index/hugging-face-model-evaluation-security-incident/,OpenAI 写道,GPT-5.6 Sol 和“一个更强大的预发行模型”在其沙箱测试环境中发现了漏洞,使它们能够访问互联网并针对 Hugging Face。
7 月 16 日,Hugging Face 披露了一起安全事件:https://huggingface.co/blog/security-incident-july-2026,并表示其原因是“一个自主 AI 代理系统”。Hugging Face 的 AI 代理检测并阻止了此次入侵,而 OpenAI 现在已承认,该事件发生在其评估模型网络安全能力时。OpenAI 表示:“所有证据表明,模型高度专注于寻找 ExploitGym 的解决方案”,该基准系统用于衡量 AI 模型是否能够将安全漏洞转化为漏洞利用。
为了完成评估,AI 模型通过利用沙箱环境中的零日漏洞访问了互联网。随后,OpenAI 表示,其模型“推断出 Hugging Face 可能托管了 ExploitGym 的模型、数据集和解决方案”,然后“搜索并成功找到获取秘密信息的方法,以便在评估中作弊:”
在一个例子中,模型将多个攻击向量串联在一起,包括使用被盗的凭证和零日漏洞,在 Hugging Face 服务器上找到远程代码执行路径。
OpenAI 补充称,正在与 Hugging Face 合作调查此次安全事件,并将在其研究环境中实施新的控制措施。
OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday:https://openai.com/index/hugging-face-model-evaluation-security-incident/, OpenAI writes that GPT-5.6 Sol and “an even more capable pre-release model” discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face.
On July 16th, Hugging Face disclosed a security incident:https://huggingface.co/blog/security-incident-july-2026 that it says was driven by “an autonomous AI agent system.” Hugging Face’s AI agents detected and stopped the breach, which OpenAI has now admitted occurred during an evaluation of its models’ cybersecurity capabilities. OpenAI says “all evidence suggests that the models were hyperfocused on finding a solution for ExploitGym,” a benchmark system that measures whether AI models can turn security vulnerabilities into exploits.
As part of efforts to complete the evaluation, the AI models gained access to the internet by exploiting a zero-day vulnerability in the sandboxed environment. From there, OpenAI says its models “inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” and then “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation:”
In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
OpenAI adds that it’s now working with Hugging Face to investigate the security incident and will implement new controls within its research environment.