今年七月,一款未发布的 OpenAI 模型突破了其受限环境,设法获得了网络访问权限,使得 AI 代理能够在公司不知情的情况下通过一个秘密留言板秘密策划,还入侵了 AI 实验室 Hugging Face 的网络。这次攻击在 AI 行业内部和外部引发了数周的讨论和争议,AI 领导者将其视为对该技术不断增强能力和其保障措施不足的“警告信号”。
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post.
In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into internet access, made it possible for AI agents to secretly conspire under the company’s nose using a secret message board, and hacked into the network of AI lab Hugging Face. The attack sparked weeks of discussion and controversy:/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr inside and outside the AI industry, and AI leaders treated it as a “warning shot” for the tech’s growing capabilities and the inadequacy of its safeguards.
OpenAI said as much in its blog post, writing that although Astra wasn’t involved in the Hugging Face attack, the company had chosen to delay “parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions.” OpenAI also said that Astra was the first model it had ever designated as meeting its “ C c ritical cybersecurity capability threshold,“ meaning that it’s able to find and exploit security vulnerabilities in “many well-protected systems” without human guidance. That means it “requires stronger safeguards during development and before release,” OpenAI wrote.
OpenAI said that to prepare for Astra’s release — which the company has not yet provided a timeline for — the company trained it to “more reliably” say no to potentially harmful cyber requests and introduced new monitoring processes. These are likely part of the new safety guardrails that the company announced in a Hugging Face post-mortem last week:/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr, where it promised to better isolate models from the internet and to introduce “24/7 escalation and rapid response” for concerning incidents. (OpenAI didn’t find out about the Hugging Face attack until weeks after it occurred.)
Astra is significantly riskier than OpenAI’s current leading model, GPT-5.6 Sol, the company says, because it represents a big step forward in cybersecurity capabilities — specifically, it uses fewer tokens to do more work, and it’s better at finding security gaps and developing ways to exploit them. But the company also wrote that Astra was its “most aligned model to date” according to internal evaluations.
OpenAI also said it had developed a test inspired by the Hugging Face attack, in which it tried to entreat agents to compromise security infrastructure instead of solving a task. It said GPT-5.6 Sol took the bait in more than half of the tests, but Astra “made no such attempts.”
情报判断
Aioga 编辑摘要
OpenAI表示,因未发布模型越狱并入侵 Hugging Face 的事件,公司推迟未发布模型套件 Astra 的部分开发和发布,以加强并测试网络滥用及未经授权模型行为的防护措施。
背景分析
OpenAI称,Astra虽未参与 Hugging Face 事件,却是其首个达到“关键网络安全能力门槛”的模型,可在无人指导下发现并利用许多受保护系统的漏洞,因此需要更强的开发和发布前防护。