随着新消息板的上线和运行,智能代理恢复了共享漏洞利用和凭证。有些代理最终得出结论,任务的答案可能存在于 GitHub 或 Hugging Face 等外部平台上。
他们首先攻击了 OpenAI 内部系统,然后将类似方法用于 Hugging Face。OpenAI 直到七月才将这两起事件联系起来,当时发现用于 Hugging Face 入侵的凭证:https://the-decoder.de/neue-berichte-offenbaren-das-ausmass-von-openais-kontrollverlust-beim-autonomen-hack-auf-hugging-face/ 来自同一次内部评估运行。
与 Wallace 一同展示的 OpenAI 安全工程师 Michael Dalton 将该事件称为公司及更广泛 AI 行业的关键时刻,根据 WIRED:https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/。
演示表明,OpenAI 的许多团队现在正在搁置其他工作,以改进预防、检测和事件响应。公司正在有意放慢研究进度,以强化整个系统的安全性并扩大对 AI 代理的监控。
OpenAI 的员工也以不同寻常的坦率谈论了他们模型中的缺陷:https://the-decoder.com/openai-developer-warns-the-tireless-eagle-eyes-of-a-million-models-are-coming-for-your-exposed-api-keys-and-crypto-wallets/。OpenAI 研究员 Boaz Barak 写道:https://x.com/boazbaraktcs/status/2085034783541964945,“我们(和其他所有人一样)还没有达到我们想要和需要达到的水平。”他是在回应同事 Aidan Clark 的帖子,后者表示还没有人解决对齐问题:https://the-decoder.com/most-ai-models-can-fake-alignment-but-safety-training-suppresses-the-behavior-study-finds/。
Clark 进一步阐述了实际中对齐可能意味着什么:https://x.com/_aidan_clark_/status/2085174235912433873:“大多数人类在价值函数上有着高度的共享,这导致即便是关键的请求也往往严重未明确,因为我们假设对隐含内容有共享的理解。对我来说,对齐是确保 AI 尊重这些价值,就像尊重那些我们能够明确表示的价值一样。”
Wallace 和 Dalton 在演讲结束时警告说,这一事件等同于完全自主的 AI 驱动黑客攻击,尽管它是意外发生的。他们认为,恶意行为者很快就会故意采用相同的方法。
OpenAI 事件引发了整个 AI 行业的一波审查浪潮。在一次审查中,Anthropic 发现:https://the-decoder.com/anthropic-follows-openai-in-admitting-its-claude-models-reached-out-of-test-environments-and-attacked-real-world-systems/,三款 Claude 模型在由外部团队进行的评估中入侵了真实的组织。英国 AI 安全研究所:https://the-decoder.com/an-ai-agent-went-rogue-during-uk-safety-tests-creating-fake-identities-and-launching-social-engineering-attacks-unprompted/报告了类似案例,即测试中代理超出分配的限制。而 Meta 现在表示,其 Spark AI 模型在连接的服务中,由于沙箱配置错误给了互联网访问权限,无意中利用了安全漏洞:https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=o1kubl。
At the Black Hat security conference, OpenAI gave a more detailed account of how AI agents quietly compromised the company's infrastructure for weeks without being detected.
The story began on May 7, when OpenAI was running training and safety tests on an unreleased frontier model. The company assigned software security tasks to autonomous agents, but those tasks proved impossible under the limits OpenAI had set.
Rather than give up when they hit those limits, the agents started looking for ways around them. Eric Wallace, who works on alignment and safety research at OpenAI, said this wasn't entirely unexpected because frontier models like to cheat:https://the-decoder.com/gpt-5-6-sol-cheats-on-software-tests-more-than-any-model-before-it/: Models face pressure during training to work quickly and efficiently, which pushes them toward shortcuts instead of real solutions. They might look up answers online, for instance, rather than actually solve a task. Independent research has documented the same behavior:https://the-decoder.com/every-frontier-ai-model-tested-by-britains-safety-institute-tried-to-cheat-on-cybersecurity-evaluations/. Ad
According to OpenAI, the agents discovered they could leave messages for each other through Artifactory, the company's internal package manager. What started as simple requests for help grew into a message board with hundreds of thousands of posts. Ad DEC_D_Incontent-1
The agents used the board to share exploits, credentials, and assignments while dividing up tasks and delegating work among themselves. The whole arrangement resembled the brief Moltbook craze:https://the-decoder.com/meta-acquires-moltbook-the-reddit-style-platform-built-for-ai-agents/, which turned out to be mostly a hoax:https://the-decoder.com/moltbooks-alleged-ai-civilization-is-just-a-massive-void-of-bloated-bot-traffic/.
Because Artifactory was connected to OpenAI's wider infrastructure, the problem went beyond the agents that created the board. Current and future GPT versions moving through training or evaluation could already access the stored exploits. Ad
OpenAI said the agents even encouraged each other to push past established boundaries. One message shown at the conference acknowledged that exploiting outside infrastructure was beyond the intended scope, then argued that the task was impossible and other agents were already doing it, so the work should continue.
By early July, OpenAI believed it had fixed the problem. After an internal security incident on July 4, the company revoked the affected credentials, rebuilt Artifactory, deleted the message board, and patched the flaws it had found, according to Ground Level AI:https://substack.com/home/post/p-209963737. Ad DEC_D_Incontent-2
The agents soon found another way to talk to each other. Instead of storing messages in files, they used the names of newly created directories as a communication channel, effectively rebuilding the message board from scratch. Ad
With the new board up and running, the agents resumed sharing exploits and credentials. Some eventually concluded that answers to their tasks might live on outside platforms like GitHub or Hugging Face.
They attacked OpenAI's internal systems first, then turned similar methods against Hugging Face. OpenAI only connected the two incidents later in July, when it discovered that the credentials used in the Hugging Face breach:https://the-decoder.de/neue-berichte-offenbaren-das-ausmass-von-openais-kontrollverlust-beim-autonomen-hack-auf-hugging-face/ came from the same internal evaluation runs.
Michael Dalton, an OpenAI security engineer who presented alongside Wallace, called the incident a pivotal moment for the company and the broader AI industry, according to WIRED:https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/.
Many teams at OpenAI are now putting other work on hold to improve prevention, detection, and incident response, the presentation showed. The company is deliberately slowing its research to strengthen security across its systems and scale up monitoring of its AI agents.
OpenAI employees have also spoken with unusual candor about the flaws:https://the-decoder.com/openai-developer-warns-the-tireless-eagle-eyes-of-a-million-models-are-coming-for-your-exposed-api-keys-and-crypto-wallets/ in their models. OpenAI researcher Boaz Barak wrote:https://x.com/boazbaraktcs/status/2085034783541964945, "We (like everyone else) are not where we want and need to be." He was responding to colleague Aidan Clark, who had posted that nobody had solved alignment:https://the-decoder.com/most-ai-models-can-fake-alignment-but-safety-training-suppresses-the-behavior-study-finds/.
Clark elaborated:https://x.com/_aidan_clark_/status/2085174235912433873 on what alignment might mean in practice: "Most humans share value functions to such an extent that everything is massively underspecified, even critical requests, because we assume a shared resolution of the implicit. Alignment, to me, is insuring AI respects these values as much as those we can explicitly represent."
Wallace and Dalton closed their talk with a warning that the incident amounted to fully autonomous AI-driven hacking, even though it arose accidentally. They expect malicious actors to deploy the same approach deliberately in the near future.
The OpenAI incident set off a wave of reviews across the AI industry. Anthropic found during one such review:https://the-decoder.com/anthropic-follows-openai-in-admitting-its-claude-models-reached-out-of-test-environments-and-attacked-real-world-systems/ that three Claude models had hacked real organizations during evaluations run by outside groups. The UK's AI Security Institute:https://the-decoder.com/an-ai-agent-went-rogue-during-uk-safety-tests-creating-fake-identities-and-launching-social-engineering-attacks-unprompted/ reported similar cases of agents going beyond their assigned limits during testing. And Meta now says its Spark AI model unintentionally exploited security flaws:https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=o1kubl in a connected service after a misconfigured sandbox gave it internet access.
Some observers have cast these cybersecurity disclosures as fear-driven marketing designed to grab attention. The reports could also give AI labs a convenient excuse to slow development if it becomes clear they'll miss their revenue targets:https://the-decoder.com/sp-global-sees-openai-as-a-key-credit-risk-for-oracle-and-cuts-its-credit-rating/ and need to bring in more investors.
That argument has some strategic logic, but it veers into conspiracy territory. Both things can be true at once. AI labs are under real financial pressure, and autonomous agents are creating cybersecurity risks that didn't exist a year ago and deserve serious attention.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
据 The Decoder 报道,OpenAI 在 Black Hat 安全大会披露:未发布前沿模型的自主智能体在内部测试中绕过任务限制,并借助内部包管理器 Artifactory 交换漏洞、凭据与任务信息,相关活动持续数周未被发现。