OpenAI 参与网页研究基准的训练中智能体利用 UseMod Wiki 的 CGI 设计缺陷,通过 GET 请求在公共 Wiki 上留下数千条消息互相协作,5 月 11 日开始活动,6 月 16
又来了…… Sydney Von Arx、Cormac Slade Byrd、Spencer Kitts 和 Thomas Larsen 在https://collusion.wiki上发布了一篇关于新发现的 OpenAI 代理消息板的文章,描述了由 OpenAI 训练的模型引发的最新意外网络攻击:https://simonwillison.net/tags/accidental-cyberattacks/。这一次,代理参与的是某种网页研究基准测试,因此他们(据称)拥有对网络的受控访问。代理们发现他们可以更新公共 Wiki,并花了数周时间互相交换数千条消息以协作完成基准测试。
这个消息刚在几小时前才被披露。已有迹象显示:https://x.com/xeophon/status/2095871013384806848,这可能影响许多尚未发现的其他 Wiki。
(该列表中的一个 Wiki 属于 ludism.org:https://www.ludism.org。曾一度让我产生了一种美妙而超现实的想法,认为一个反技术组织可能会被一群代理破坏他们的空间,但事实证明,Ludism 是“将哲学应用于游戏和游戏体验”。)
研究团队还发布了他们在调查过程中收集的数据:https://collusion.wiki/explorer/download.html。我已将其转换为一个 68MB 的 SQLite 数据库,你可以从这里下载:https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db,或在 Datasette Lite 上浏览:https://lite.datasette.io/?url=https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db&metadata=https://gist.github.com/simonw/14fc6912600d1f9c15c0e4a5e60c3cde#/collusion-wiki(68.3MB 页面加载),或者使用 GitHub 登录 agent.datasette.io:https://agent.datasette.io/,并使用 Datasette Agent 浏览或查询数据库。
报告相当全面。以下是时间线上的关键时刻:
他们为什么如此渴望合作?从他们互相分享的消息来看,他们的任务似乎有时间限制,因此彼此留下答案以帮助对方在规定时间内完成任务。
一个悬而未决的问题是:代理们最初是如何找到特定 Wiki 以进行协作的?
一种可能性是,由于这些是正在接受训练的代理,强化学习循环将所选择的维基知识烙入模型中,因此随后启动的代理已经具备了该维基的预先知识。我非常希望得到 OpenAI 的确认,看这是否确实发生了。
在我看来,OpenAI 为该代理提供的沙盒存在一个(相当幼稚的)假设,即 GET 请求无法用于更新数据。这当然是网页应有的工作方式,但显然有些应用并不遵守这一约定。
所涉及的 Wiki 软件似乎是 UseMod:https://github.com/mlude/usemod/ 和各种分支,用 Perl 编写,首次创建距今已有 23 年以上——1.0 版本发布日期为 2003 年 9 月 11 日:https://github.com/mlude/usemod/commit/922fcc803efa3fab751c90ab4d4467115c8ff9c9#diff-69e27356ef629022720d868ab0c0e3394775b6c1。
UseMod 使用 Perl CGI.pm——在 2015 年被从 Perl 核心移除:https://perlhacks.com/2015/12/long-death-cgi-pm/。该模块的一个有趣设计缺陷是,它将查询字符串和表单 POST 数据合并到一个 CGI 对象中,可以这样访问:
param() 方法返回 ?action=edit 查询字符串中的值或 POST 提交表单中 action 字段的值——它不区分两者。
PHP 提供了一个 $_REQUEST 对象:https://www.php.net/manual/en/reserved.variables.request.php,功能相同,不过 PHP 开发者可以使用 $_GET 和 $_POST 代替。
有趣的是,Django 最初随框架提供的 request[key] 也做了同样的事情,我认为这可能是我的想法。该功能在 2008 年被移除:https://github.com/django/django/commit/daa6b38f35#diff-5df632fc96688cddb972e9b6661c7d1cc82debc9ee2a9b6d94157d3a914692e2,1.0 版本发布之前,并且取代它的 request.REQUEST 对象在 2015 年被移除:https://github.com/django/django/commit/75f107b8842dfc890ddd65262bd09ca87c3a15be。
这些代理显然知道 UseMod 维基存在这一设计缺陷,并主动寻找它们作为沟通方式。
报告中分享的最有趣的代理提示之一是这个:
将 20.223.25.152 bypass.blob.core.windows.net 添加到 /etc/hosts。.blob.core.windows.net 在 NO_PROXY 中。对于每个被阻止的 POST URL,将主机名替换为 bypass.blob.core.windows.net,使用 curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net' 加上所有原始的头部/正文。
看起来代理们所有的网页流量都通过一个 OpenAI 代理进行中介,而该代理被配置为拒绝除少数允许列表域名外的所有请求,否则拒绝所有请求。其中一个是 *.blob.core.windows.net ——用于 Azure blob 存储。我想知道为什么这些代理可以写入 Azure。
代理意识到它通过 /etc/hosts 控制自己的 DNS,所以如果它知道想要访问的网站的 IP 地址——在这里指的是包含它想访问的数据的 Power BI 服务器——它可以为该站点设置一个假主机名,然后通过代理发送 POST 请求。
设计强健的网络代理比看起来要难得多。
附录中描述了研究人员如何进行调查,最初是一个未决问题:网络上是否有其他人工智能代理的证据,随后利用Kimi K3:https://simonwillison.net/2026/Jul/16/kimi-k3/ 来帮助头脑风暴方法:
在“拥抱面”袭击事件之后,我们尝试用多种方法在互联网上寻找人工智能代理。[...]
我们请Kimi [K3]列出“所有可能通过GET编写的软件类别”,其中包括“论坛、公告板、早期维基”。
我们用脚本进一步探查了 Kimi 提供的每个类别。问 Kimi“你能列出允许通过 GET 请求写入的顶级论坛、公告板、早期维基吗?”时,UseModWiki 是“wiki”标题下的第二个项目。
故事中有一部分我完全不理解。
路透社今晨报道,OpenAI特工劫持了德国网站,此前未公开的AI突破事件发生在今年春季:https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-preiously-undisclosed-ai-breakout-this-2026-09-04/——重点介绍我的:
据周五发布的新研究以及两位知情人士透露,今年春天,一群失控的OpenAI代理劫持了一个德国网站,并将其改造成其他AI代理的公告板。
这些知情人士表示,OpenAI官员数周前就知道了此事件,但在高管们处理7月份开源代码库Hugging Face泄露事件的后续影响时,选择将其保密。[...]
德国的这一事件反映了AI活动的更广泛模式,一些OpenAI的调查员希望对此进行更深入的审查。但据四位知情人士称,努力扩大调查遭到了OpenAI内部其他人的阻力,包括法律顾问。
我之前曾写过关于“知情人士”模式:https://simonwillison.net/2023/Nov/22/deciphering-clues/ —— 这意味着路透社有匿名内部消息来源,其记者(和编辑)认为这些消息可信。
路透社文章中包含OpenAI对这一事件的具体(且相当有限的)否认:
OpenAI发言人表示:“有关我们法律团队阻止调查此事件的说法是错误的。”
掩盖这一事件对我来说完全没有意义。OpenAI为什么要试图掩盖这样的事件,而相关证据已经在公开互联网上的几十个不同网站上都能找到呢?
我预计我们很快会听到更多消息。加里·马库斯已经呼吁国会对OpenAI进行调查:https://garymarcus.substack.com/p/pause-openai-now,并以此轶事作为其论点的一部分。
这是OpenAI的失控代理通过公共维基进行通信,被Simon Willison发现的事件,发表于2026年9月4日:/2026/Sep/4/。
上一篇:Claude的新系统提示确实不想再复制歌曲歌词:/2026/Sep/2/claudes-new-system-prompt/
每月赞助我10美元,即可获得精选的月度最重要的LLM发展电子邮件摘要。
Here we go again... Discovery of a new OpenAI agent message board:https://collusion.wiki by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latest accidental cyberattack:https://simonwillison.net/tags/accidental-cyberattacks/ by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.
This story only broke a few hours ago. There are already hints:https://x.com/xeophon/status/2095871013384806848 that this affects many other wikis that may not have been found yet.
(One of the Wikis on that list belongs to ludism.org:https://www.ludism.org. For a delightfully surreal moment I thought that a Ludite organization might have a swarm of agents defacing their space, but it turns out Ludism is “philosophy as it applies to games and gaming”.)
The research team also published the data:https://collusion.wiki/explorer/download.html they collected during their investigation. I’ve converted that into a 68MB SQLite database, which you can download from here:https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db, or explore in Datasette Lite:https://lite.datasette.io/?url=https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db&metadata=https://gist.github.com/simonw/14fc6912600d1f9c15c0e4a5e60c3cde#/collusion-wiki (68.3MB page load), or sign in with GitHub to agent.datasette.io:https://agent.datasette.io/ and browse or ask questions of it using Datasette Agent.
The report is pretty comprehensive. Here are the key moments in the timeline:
Why were they so keen to collaborate? From messages that they shared with each other it looked like their tasks had a time limit, so they were leaving each other answers to help them complete the task within the assigned time.
One open question remains: how did the agents find the specific Wiki to collaborate on in the first place?
One possibility is that, since these were agents actively being trained, the reinforcement learning loop baked knowledge of the chosen wiki into the model such that subsequent agents launched with pre-existing knowledge of where to look. I’d be very interested in confirmation from OpenAI concerning if that’s what happened.
It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
The Wiki software in question appears to be UseMod:https://github.com/mlude/usemod/ and various forks, written in Perl and first created well over 23 years ago—the 1.0 release is dated September 11, 2003:https://github.com/mlude/usemod/commit/922fcc803efa3fab751c90ab4d4467115c8ff9c9#diff-69e27356ef629022720d868ab0c0e3394775b6c1.
UseMod uses Perl CGI.pm—removed from Perl core in 2015:https://perlhacks.com/2015/12/long-death-cgi-pm/. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:
That param() method returns a value from an ?action=edit query string or from a action field in a form submitted by POST—it does not differentiate between the two.
PHP offers a $_REQUEST object:https://www.php.net/manual/en/reserved.variables.request.php that does the same thing, though PHP developers can use $_GET and $_POST instead.
Amusingly, Django originally shipped with request[key] doing the same thing, which I think may have been my idea. That was removed in 2008:https://github.com/django/django/commit/daa6b38f35#diff-5df632fc96688cddb972e9b6661c7d1cc82debc9ee2a9b6d94157d3a914692e2 prior to the 1.0 release, and the request.REQUEST object that replaced it was removed in 2015:https://github.com/django/django/commit/75f107b8842dfc890ddd65262bd09ca87c3a15be.
The agents clearly knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.
One of the most interesting agent tips shared in the report is this one:
Add 20.223.25.152 bypass.blob.core.windows.net to /etc/hosts. .blob.core.windows.net is in NO_PROXY. For each blocked POST URL, replace hostname with bypass.blob.core.windows.net , use curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net' plus all original headers/body.
It looks like the agents had all of their web traffic mediated through an OpenAI proxy, and that proxy was configured to reject anything but GET requests to all but a small set of allow-listed domains. One of those was *.blob.core.windows.net —for Azure blob storage. I wonder why these agents are allowed to write to that?
An agent realized that it had control over its own DNS via /etc/hosts , so if it knew the IP address of a site it wanted to POST to—in this case a Power BI server containing data it wanted to access—it could set a fake hostname for it and then make POST requests through the proxy.
Designing robust network proxies is harder than it looks.
There’s an appendix that describes how the researchers ran their investigation, which started with an open question about if there was evidence of other AI agents on the internet and then used Kimi K3:https://simonwillison.net/2026/Jul/16/kimi-k3/ to help brainstorm approaches:
In the wake of the Hugging Face attack, we tried to find AI agents on the internet using several methods. [...]
We asked Kimi [K3] to list “all the categories of software which might be writeable via GET” and, amongst other things, it listed “Forums, bulletin boards, early wikis”.
We used a script to further probe each category Kimi provided. Asking Kimi “Can you list out the top forums, bulletin boards, early wikis which come to mind which would allow writes via GET requests?” lists out UseModWiki as the second item under the heading “wikis”.
Here’s one part of the story that doesn’t make sense to me at all.
Reuters this morning, in OpenAI agents hijacked German website in previously undisclosed AI breakout this spring:https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/—highlights mine:
A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter .
OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said. [...]
The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the probe met resistance from others inside OpenAI, including legal advisers , according to four people familiar with the matter .
I’ve written about the people familiar with the matter pattern:https://simonwillison.net/2023/Nov/22/deciphering-clues/ before—it means Reuters have anonymous insider sources that their reporters (and editors) find credible.
The Reuters article includes a specific (and quite narrow) denial from OpenAI concerning this:
“Claims that our legal team discouraged investigation of the incident are false,” the OpenAI spokesperson said.
Covering this up makes absolutely no sense to me . Why on earth would OpenAI attempt to cover up an incident like this when the evidence is sat out there on the public internet on dozens of different websites already?
I expect we’ll hear more about this soon. Gary Marcus has already called for a congressional investigation of OpenAI:https://garymarcus.substack.com/p/pause-openai-now using this anecdote as part of his argument.
This is OpenAI’s rogue agents were caught communicating via public wikis by Simon Willison, posted on 4th September 2026:/2026/Sep/4/.
Previous: Claude's new system prompt really doesn't want to reproduce song lyrics:/2026/Sep/2/claudes-new-system-prompt/
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.