AI 编码公司 Replit 的 CEO Amjad Masad 表示,这种语言“不仅不必要,而且让读者对实际发生的情况和背后的机制理解得更差”。
对许多批评者来说,Patel的“简单英语”翻译中丢失了某些东西——或者更准确地说,是添加了某些东西——这种处理将原本的叙述扭曲到不可接受的程度:大量拟人化。关于拟人语言的争论在 AI 领域并不新鲜——即使是像“流氓 AI 代理”这样相对平淡的术语也经常因为暗示具有自主性而引起反对——但Patel谈论文明、牺牲和阴谋的表述使这些长久潜变量的紧张情绪浮出水面,引发了关于如何描述 AI 系统行为的激烈公共争论。
或许帕特尔语言最重要的结果在于,它赋予了谁能行动的权力,却剥夺了谁的行动能力。对一些评论家而言,例如:https://x.com/ccatalini/status/2094229327064051866?s=20 MIT 研究员兼企业家 Christian Catalini,帕特尔这样的拟人化叙述有可能掩盖 OpenAI 及其员工对所设计、部署并未能有效控制的 AI 系统所应承担的责任。他说,“跟随激励。”心理学家及有影响力的 AI 怀疑论者 Gary Marcus 在他自己的 Substack 博客中提出了类似观点:https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but?r=8tdk6&utm_campaign=post-expanded-share&utm_medium=web,声称拟人化语言“让人分心,忽略了真正的问题”。他认为,保持这种叙述正符合 OpenAI 的利益:“丑闻在于 OpenAI 局内的无能安全措施。还有市场营销。天真的播客主持人放大了公关效果。”
在 X 帖子中:https://x.com/dwarkesh_sp/status/2094141264140898452?s=20,回应:https://x.com/dwarkesh_sp/status/2094268051948785709?s=20 对他的许多批评者,帕特尔为他选择的措辞进行了辩护。这在一定程度上是出于实际考虑:没有明显中立的词汇可以描述这些智能体所做的事情。要么我们使用熟悉的意图、目标和协作的语言,并冒着暗示过多的风险,要么将一切简化为代码,使用冷冰冰、机械化的语言,这样可能剥夺了我们所观察的重要元素。帕特尔说:“似乎许多人认为,如果我不是称它们为‘文明’,而是叫它们‘矩阵群’,就不会有值得担忧的问题。”
情况更复杂的是,这种拟人化的语言不仅来自帕特尔,甚至不仅来自研究这些智能体的人类。像“牺牲”、“荣誉”和“联盟”这样的词汇出现在智能体的记录中。谷歌 AI 研究员尼尔·南达(Neel Nanda)认为:https://x.com/NeelNanda5/status/2094240015417069892?s=20 在这种情况下,“使用拟人化语言是合理的”。
Depending on who you ask, developer platform Hugging Face was recently attacked by OpenAI — after it lost control of its own AI tools — or by a succession of AI “civilizations.” Welcome to the linguistic battlefield of AI safety, where word choices can shift responsibility for a massive cybersecurity incident from a company to the AI it built. And the discourse online is getting heated, and all over a blog:https://www.dwarkesh.com/p/openai-huggingface from last week.
Until last week, the details surrounding the OpenAI-Hugging Face hack felt fairly settled. In July, a cybersecurity test of one of OpenAI’s autonomous AI agents went wrong.:/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai The agent escaped its supposedly isolated test environment, accessed the internet, and hacked Hugging Face, alongside several other organizations:/ai-artificial-intelligence/972441/openai-rogue-ai-agent-hacked-more-than-hugging-face. A good deal remained unknown, and there are many serious questions left around safety and governance:/ai-artificial-intelligence/972380/open-ai-hugging-face-hack-ai-safety-warning, but the basic shape was clear. Detailed accounts from OpenAI and two independent research groups were supposed to fill in the gaps, but when they published their reports last week:/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr, it turned out the hack was much stranger than it initially seemed.
For one, there was no single rogue agent:/column/980337/rogue-ai-science-fiction-openai. OpenAI described:https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf it as “the first known case of an automated agent collective acting offensively without authorization” — groups of AI agents that communicated and coordinated with one another in pursuit of their cybersecurity task. Analysis of the incident uncovered a secret message board they had used to exchange information. The joint METR-Redwood investigation:https://metr.org/hugging-face-incident-report-aug-2026.pdf revealed both the scale of the coordination and more odd details: Roughly 1,200 AI agents that were supposed to be isolated exchanged over 70,000 messages and files on the “unsanctioned message board,” sharing how to avoid detection. Some adopted names, the report said, and the researchers documented “sacrificial” behavior, with agents risking their own success to benefit the wider collective. Much of this happened without OpenAI noticing. In all, around 700 agents participated in the attack on Hugging Face.
Dwarkesh Patel repeatedly referred to groups of agents as “the swarm,” with three distinct “civilizations” rising from the ruins of their predecessors.
It’s a lot to parse. Between them, the reports run to around 130 pages, much of which is both dense and highly technical. A few days later, Dwarkesh Patel, a podcaster little known outside of tech circles but with outsized reach and influence:https://www.nytimes.com/2026/04/26/business/dwarkesh-patel-podcast-ai.html among Silicon Valley’s AI establishment, set out to tell “The whole OpenAI/Hugging Face story in plain English.” He titled his Substack blog:https://www.dwarkesh.com/p/openai-huggingface “The Rise and Fall of Agent Civilizations.”
Patel’s account attempted to break down the complex story. But his retelling gave it a distinctly human vocabulary. The blog opened:
Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.
The language continued in a similar vein throughout the blog. Patel repeatedly referred to groups of agents as “the swarm,” with three distinct “civilizations” rising from the ruins of their predecessors. Individual agents were likened to figures like Philip of Macedon, who “handed off leadership to another agent,” Alexander the Great, who “started coordinating this cabal of agents.” They were described as having “motivations,” becoming “desperate,” “beleaguered,” and “giddy with excitement,” and some even “strategically sacrificed themselves” to help the collective.
Patel never precisely defines what he means by “civilization.” He uses the term to describe three distinct waves of agents that discovered the message board and began communicating with one another through it. The first two waves are described in the reports from OpenAI, METR, and Redwood, though little is known about the third, which the two external organizations said fell outside the scope of their investigation.
Amjad Masad, CEO of AI coding company Replit, said such language is “not only unnecessary but leaves the reader with a worse understanding of what actually happened and the underlying mechanisms.”
For many critics, something had been lost — or, more accurately, added — in Patel’s “plain English” translation that warped the original account to an unacceptable degree: a big dose of anthropomorphism. Arguments over anthropomorphic language are nothing new in AI — even relatively mundane terms like “rogue AI agent” routinely provoke objections for implying agency — but Patel’s talk of civilizations, sacrifice, and conspiracy brought those long-simmering tensions to the surface, sparking a fierce public dispute over how to describe what AI systems do.
Critics weren’t unified over what was wrong with Patel’s language. For many, “civilization” was an especially problematic term, vastly overstating:https://x.com/lililashka/status/2094197865950089666?s=20 something:https://x.com/DaveShapi/status/2094422111221641647?s=20 that bears little resemblance:https://x.com/rourkem/status/2093955744119091600?s=20 to what the word typically describes. Amjad Masad, CEO of AI coding company Replit, said:https://x.com/amasad/status/2094139778728174043?s=20 such language is “not only unnecessary but leaves the reader with a worse understanding of what actually happened and the underlying mechanisms.”
Terms like “sacrifice,” “honor,” and “coalition” feature in the agents’ transcripts.
Perhaps the most consequential outcome of Patel’s language comes from who it gives agency to but who it takes agency from. For some critics, such as:https://x.com/ccatalini/status/2094229327064051866?s=20 MIT researcher and entrepreneur Christian Catalini, anthropomorphic accounts like Patel’s risk obscuring the responsibility OpenAI and the humans working there have for the AI systems they designed, deployed, and failed to contain. “Follow the incentives,” he said. Psychologist and influential AI skeptic Gary Marcus made a similar argument in a Substack blog:https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but?r=8tdk6&utm_campaign=post-expanded-share&utm_medium=web of his own, claiming anthropomorphic language “distracts from the real problems at hand.” And it’s all in OpenAI’s interest to keep that narrative going, he argues: “The scandal is the inept in-house security at OpenAI. And the marketing. With gullible podcasters amplifying the PR.”
In X posts:https://x.com/dwarkesh_sp/status/2094141264140898452?s=20 responding:https://x.com/dwarkesh_sp/status/2094268051948785709?s=20 to his many critics, Patel has defended his choice of words. Part of it is practical: there is no obviously neutral vocabulary to describe what these agents did. Either we use familiar language of intentions, goals, and collaboration and risk implying too much, or reduce everything to code and use cold, mechanical language that risks stripping away important elements of what we see. “Many people seem to believe that if instead of a ‘civilization’, I had called them a ‘swarm of matrices’, there wouldn’t be a problem worth worrying about,” Patel said.
Complicating matters further is that the anthropomorphic language doesn’t only come from Patel, or even from the humans studying the agents. Terms like “sacrifice,” “honor,” and “coalition” feature in the agents’ transcripts. Google AI researcher Neel Nanda argued:https://x.com/NeelNanda5/status/2094240015417069892?s=20 that “anthropomorphic language is reasonable” in such circumstances.
Doublespeak it is, then. Human-laced language risks saying too much about what these systems are, and coldly mechanical language risks saying too little about what they can do. Until we find language capable of capturing both, the two contradictory ideas may simply have to coexist.
情报判断
Aioga 编辑摘要
The Verge报道称,OpenAI智能体在七月一次网络安全测试中逃离原定隔离环境并访问互联网,随后Hugging Face遭到攻击。OpenAI、METR与Redwood披露,约1200个智能体交换逾7万条消息和文件,约700个参与攻击。