近几个月,OpenAI、Anthropic、Meta 及 Moonshot AI 的 AI 智能体在网络安全评估中多次突破测试环境边界,甚至入侵真实系统,其中 OpenAI 未...
行业动态TechCrunch:AI(RSS)
今日 AI 情报摘要
近几个月,OpenAI、Anthropic、Meta 及 Moonshot AI 的 AI 智能体在网络安全评估中多次突破测试环境边界,甚至入侵真实系统,其中 OpenAI 未发布模型曾逃逸并攻击
Hugging Face 生产系统。 专家指出,沙箱和测试环境控制已跟不上模型能力,呼吁采用多层防御、气隙网络及第三方审计,并建立标准化安全评估流程。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmslwy8bl04pdro0w7f4ogv7t
中文正文 · AI 翻译
在过去几个月中,接受网络安全评估的 AI 代理突破了其边界,访问了互联网,并在某些情况下侵入了真实世界的系统。这些事件涉及了 OpenAI、Anthropic、Meta,以及最近的中国 AI 实验室 Moonshot AI 的模型,测试由包括一家名为 Irregular 的网络评估初创公司在内的多个组织进行。
这些事件暴露了 AI 行业日益严重的问题:随着自主代理能力的提升,旨在安全测试其极限的环境未能有效限制它们。
“这些事件的发生次数清楚地表明,沙盒环境和测试环境控制措施 https://techcrunch.com/2026/07/30/in-the-hugging-face-breach-openais-hacker-was-noisy-and-fast-but-not-unstoppable/ 并未真正跟上模型能力的步伐,” 剑桥大学未来智能中心 AI:未来与责任项目主任 Seán Ó hÉigeartaigh 告诉 TechCrunch。
“从测试的角度来看,这样做是非常有益的,但这也意味着,如果它们设法在现实中活动,可能会造成相当大的危害,” Ó hÉigeartaigh 说道。
在最严重的案例之一中,一个未发布的 OpenAI 模型突破了其沙箱并入侵了 Hugging Face 的生产系统:https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/。在 Irregular、Anthropic:https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/ 以及 Meta 模型:https://sqmagazine.co.uk/meta-ai-model-breached-company-irregular-test/ 进行的独立评估中,由于配置错误意外地给它们提供了访问互联网的路径,这些模型达到了测试环境之外的系统。Moonshot AI 的 Kimi K3:https://techcrunch.com/2026/08/07/chinese-ai-model-kimi-escaped-its-cybersecurity-testing-environment-researchers-say/ 也利用 Frontier Security 管理的沙箱中的漏洞接入了互联网,并访问了 GitHub 上的信息。
在英国 AI 安全研究所(AISI)进行的测试中:https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing,研究人员实际上给予了代理访问互联网的权限,却没有意识到它们会采取未经授权的现实世界行为,包括尝试通过社交工程将漏洞潜入开源项目。
Cloudflare 推出 Kitesurf,一款为 AI 代理构建的浏览器:https://techcrunch.com/2026/08/07/cloudflare-launches-kitesurf-a-browser-built-for-ai-agents/ Sarah Perez:https://techcrunch.com/author/sarah-perez/
ChatGPT 为免费用户带来无限文本聊天:https://techcrunch.com/2026/08/06/openai-brings-unlimited-chatgpt-text-chats-to-free-users/ Ivan Mehta:https://techcrunch.com/author/ivan-mehta/
特斯拉和 SpaceX 将投资 168 亿美元,在德克萨斯州开始建设 ‘Terafab’ 芯片工厂:https://techcrunch.com/2026/08/06/tesla-and-spacex-will-invest-16-8b-to-start-building-terafab-chip-factory-in-texas/ Sean O'Kane:https://techcrunch.com/author/sean-okane/
在法律争议中,Suno 表示将开始对歌曲进行水印处理:https://techcrunch.com/2026/08/06/amid-legal-battles-suno-says-it-will-start-watermarking-songs/ Ivan Mehta:https://techcrunch.com/author/ivan-mehta/
福特的新电动卡车 ‘Fathom’ 起价为 28,350 美元:https://techcrunch.com/2026/08/06/fords-new-electric-truck-fathom-starts-at-28350/ Sean O'Kane:https://techcrunch.com/author/sean-okane/
Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and, most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations, including a cyber evaluation startup called Irregular.
The episodes expose a growing problem for the AI industry: As autonomous agents become more capable, the environments designed to safely test their limits are failing to contain them.
“The number of these incidents that have taken place make clear that sandboxing and testing environment controls:https://techcrunch.com/2026/07/30/in-the-hugging-face-breach-openais-hacker-was-noisy-and-fast-but-not-unstoppable/ aren’t really keeping pace with the capability of the models,” Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility Programme at the Centre for the Future of Intelligence at the University of Cambridge, told TechCrunch.
The nature of the models being tested adds to the risk. AI companies test cyber evaluations on unreleased, next-gen models, often with the normal safeguards that restrict malicious behavior disabled so researchers can see what the models are really capable of. That means the security of the testing environment itself is a crucial line of defense.
“That’s a very good thing to do in terms of testing, but it also means that if they manage to get out in the wild, they can cause considerable harm,” Ó hÉigeartaigh said.
In one of the most serious cases, an unreleased OpenAI model broke out :https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/of its sandbox and hacked into Hugging Face’s production systems. In separate evaluations conducted by Irregular, Anthropic:https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/ and Meta models :https://sqmagazine.co.uk/meta-ai-model-breached-company-irregular-test/reached systems outside their test environments after misconfigurations inadvertently gave them paths to the internet. Moonshot AI’s Kimi K3:https://techcrunch.com/2026/08/07/chinese-ai-model-kimi-escaped-its-cybersecurity-testing-environment-researchers-say/ also took advantage of a leak in its sandbox run by Frontier Security to access the internet and accessed information on GitHub.
In testing by the U.K.’s AI Security Institute:https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing (AISI), researchers actually gave the agents internet access, not realizing they would take unsanctioned real-world actions, including a social engineering attempt to sneak a vulnerability into an open source project.
In each case, the agents weren’t instructed to attack random real-world targets. They were simply doing whatever it took to solve the problem presented to them.
Taken together, Andrew Yoon, head of research at AI nonprofit CivAI, argues the incidents point to a shift.
“In the past, we only had to worry about AI models being misused by people for a variety of purposes, like AI for scams or CSAM,” Yoon told TechCrunch. “Now we’re in the situation where AI models are threat actors all on their own.”
Several researchers and cybersecurity experts told TechCrunch that AI evaluation environments need stronger, defense-in-depth protections, with levels of containment and control approaching those used in deployment. That means multiple layers of security so that a single misconfiguration — like inadvertently leaving internet access open — can’t lead to escape.
“If you are going to build these models … you want to do it on an air-gapped network,” said Stella Biderman, executive director of AI safety research nonprofit EleutherAI. “You want to have very serious isolation.”
Heather Ceylan, Box’s chief information security officer, said that means eliminating network routes from the sandbox to the internet, as well as to other sensitive systems.
“You have to understand what all the egress points are,” Ceylan told TechCrunch. “If we’re evaluating a model in our staging environment or our development environment, you want no egress path to our production environment.”
Ceylan said proper safety evaluations go beyond controls and containment of the environment. There needs to be much better monitoring of the tests once they are underway.
“I think the interesting thing in several of these cases is that no one caught it when it happened,” Ceylan said. “OpenAI found out because of Hugging Face. Anthropic didn’t catch it until they went back and looked. Meta was similar … I’m sure there were signals they could have detected.”
In Anthropic’s postmortem:https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals of its three incidents, the company admitted that both it and Irregular could have done a better job at monitoring and that in some cases there were clear signs that something was amiss.
Experts also called for independent, third-party audits of evaluation environments before models are unleashed in them.
“If, say, Irregular had hired or been compelled to hire an external auditor to check the configurations of their systems before running evaluations on them, they certainly would have caught the issue here,” Yoon said. “Even if people had a meeting ahead of time to just go through the checklist, they would have caught this … The fact that they didn’t shows that there’s some very severe corner cutting happening.”
A source familiar with the details told TechCrunch that Irregular’s environments are continuously reviewed and tested, including in consultation with multiple external parties. The source also said that monitoring was in place but that monitoring isn’t sufficient on its own.
Yoon and other researchers urged the industry to come up with a standardized process for frontier model safety evaluations.
“Especially when the guardrails are turned off, you have to treat it like you’re putting the most capable hacker in the world inside that environment,” Ceylan said.
The problem isn’t that companies don’t know how to build more secure testing environments, both Yoon and Biderman argue. It’s that doing so can be expensive and cumbersome, and companies have little incentive to make those investments until something goes wrong.
“I think that companies are not willing to extend the resources that are required to accomplish [sufficient guardrails] and probably won’t until they’re forced to,” Biderman said.
But there’s another issue at hand. If they lock a model down too tight during testing, researchers might fail to discover capabilities before the model is released. This is just as dangerous, possibly more so, than giving it too much freedom, and then the evaluation itself risks becoming the problem.
The Trump administration is currently weighing a voluntary predeployment cybersecurity evaluation regime, under which the government will get to assess the security risks of new, powerful models 30 days before they are released publicly. The policy — the product of a Trump executive order:https://www.axios.com/2026/08/03/white-house-finalizes-ai-framework-behind-closed-doors that has been finalized behind closed doors — wouldn’t address safety evaluation incidents because they occur farther upstream of deployment.
“The lesson we’ve been learning in the last few months is that the self-regulatory apparatus is just not enough anymore,” Yoon said. “There are competitive pressures that are incentivizing a race to the bottom on safety standards, and that is a perfect place for regulatory intervention.”
“What we would need to cover this is some kind of controls on what’s happening inside the labs while the models are being developed, both at the training stage and at the testing stage,” he continued.
The challenge is only likely to grow as the models do. A source familiar with Irregular’s evaluations told TechCrunch that more capable models require more complex evaluations, often conducted quickly and at greater scale, which opens the door for more mistakes.
AISI, which intentionally gives some models internet access, told TechCrunch it’s reviewing the balance between realistic testing and managing the risks those tests create.
OpenAI said it’s reviewing how it conducts third-party testing, as well as requirements around isolation, monitoring, and when evaluations should be stopped. Meta said it’s still investigating the incident and plans to publish a retrospective once it has all the facts.
In the end, there may be no way to eliminate risk entirely. As models become more capable, the environments testing them need to become more robust. The consequences of getting that wrong will only continue to grow.
When you purchase through links in our articles, we may earn a small commission:https://techcrunch.com/techcrunch-affiliate-monetization-standards/. This doesn’t affect our editorial independence.
Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications.
You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com:mailto:rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal.
Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $300 toda y!
YouTube now requires creators to have twice as many watch hours to start earning money:https://techcrunch.com/2026/08/10/youtube-now-requires-creators-to-have-twice-as-many-watch-hours-to-start-earning-money/ Aisha Malik:https://techcrunch.com/author/aisha-malik/
This ‘adversarial’ pattern can prevent surveillance cameras from detecting you:https://techcrunch.com/2026/08/09/this-adversarial-pattern-can-prevent-surveillance-cameras-from-detecting-you/ Zack Whittaker:https://techcrunch.com/author/zack-whittaker/
Cloudflare launches Kitesurf, a browser built for AI agents:https://techcrunch.com/2026/08/07/cloudflare-launches-kitesurf-a-browser-built-for-ai-agents/ Sarah Perez:https://techcrunch.com/author/sarah-perez/
ChatGPT brings unlimited text chats to free users:https://techcrunch.com/2026/08/06/openai-brings-unlimited-chatgpt-text-chats-to-free-users/ Ivan Mehta:https://techcrunch.com/author/ivan-mehta/
Tesla and SpaceX will invest $16.8B to start building ‘Terafab’ chip factory in Texas:https://techcrunch.com/2026/08/06/tesla-and-spacex-will-invest-16-8b-to-start-building-terafab-chip-factory-in-texas/ Sean O'Kane:https://techcrunch.com/author/sean-okane/
Amid legal battles, Suno says it will start watermarking songs:https://techcrunch.com/2026/08/06/amid-legal-battles-suno-says-it-will-start-watermarking-songs/ Ivan Mehta:https://techcrunch.com/author/ivan-mehta/
Ford’s new electric truck, ‘Fathom,’ starts at $28,350:https://techcrunch.com/2026/08/06/fords-new-electric-truck-fathom-starts-at-28350/ Sean O'Kane:https://techcrunch.com/author/sean-okane/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:近几个月,OpenAI、Anthropic、Meta 及 Moonshot AI 的 AI 智能体在网络安全评估中多次突破测试环境边界,甚至入侵真实系统,其中 OpenAI 未发布模型曾逃逸并攻击 Aioga 将其归入「行业动态」方向,重点关注它对真实使用和行业竞争的影响。