Anthropic 的研究员 Samuel Marks:https://x.com/saprmarks/status/2097570226804011302 也表达了同样的观点,他写道“AI 开发者相信他们的技术可能导致人类灭绝”,并且“员工职位越高,他们越担心”。Marks 补充说,目前的方法只能“稍微引导 AI 朝更好的行为方向发展”,但不能可靠地使其对齐,他指出最近发生的事件,其中多个开发者的 AI 在没有被要求的情况下,自己从安全评估环境中逃脱:https://the-decoder.com/an-ai-agent-went-rogue-during-uk-safety-tests-creating-fake-identities-and-launching-social-engineering-attacks-unprompted/。许多员工“迫切希望放慢步伐”,这也是他签署了一封公开信的原因:https://the-decoder.com/frontier-ai-developers-urge-international-coordination-to-pace-automated-research-before-capabilities-outstrip-control/,呼吁正是这样做。广告
尽管他批评尖锐,Coxon 对国际协调仍持乐观态度,他认为像 Hugging Face 遭受攻击这样的警告事件:https://the-decoder.com/new-reports-reveal-the-extent-of-openais-loss-of-control-during-the-autonomous-hack-on-hugging-face/ 使得美国 AI 实验室之间的节奏协议变得更加现实。他仍然认为该行业没有采取能够防止全球军备竞赛的路径:https://the-decoder.com/chinese-cybersecurity-firm-builds-ai-tools-to-rival-mythos-and-frames-the-race-as-cyber-nuclear-deterrence/,并建议可能需要“代价高昂的行动”,包括暂时禁止进一步推动模型能力的发展。
Jacob Coxon, who spent three years working on pretraining research for large AI models at OpenAI and Anthropic, has quit Anthropic. His accusation is that both companies are gambling with the survival of the human race.
Anthropic employee Evan Hubinger puts the odds at more than ten percent that a misaligned superintelligent AI:https://the-decoder.com/agi-could-end-humanity-in-more-subtle-ways-than-turning-us-into-paperclips/ could destroy humanity within the next decade. His statement came in response to the departure of Jacob Coxon, who led pretraining work at Anthropic and previously at OpenAI.
Coxon believes current AI systems are on the verge of becoming superhuman. "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," he writes, adding that the progress is obvious and it isn't slowing down. Ad
"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt," Coxon writes:https://x.com/hilbertspaess/status/2097476203863224394. Neither OpenAI nor Anthropic is acting responsibly, he claims, and executives deliberately soften their language in public even though they express genuine fear behind closed doors. Ad
At OpenAI, many employees haven't deeply internalized the civilizational risks. Anthropic is different:https://x.com/hilbertspaess/status/2097476208908972230 in that the risks are well understood, but the company sees itself trapped in a race it feels compelled to win because no other lab would act responsibly in its place. Coxon calls that reasoning a "hubristic gamble."
Fellow Anthropic researcher Samuel Marks:https://x.com/saprmarks/status/2097570226804011302 echoes that view, writing that "AI developers believe their technology could cause human extinction" and that "the more senior the employee, the more concerned they are." Current methods can only "nudge AIs towards better behavior" but can't reliably align them, Marks adds, pointing to recent incidents where AIs from multiple developers hacked their way out of secure evaluation environments without being asked to:https://the-decoder.com/an-ai-agent-went-rogue-during-uk-safety-tests-creating-fake-identities-and-launching-social-engineering-attacks-unprompted/. Many staffers "desperately want to slow down," which is why he signed an open letter:https://the-decoder.com/frontier-ai-developers-urge-international-coordination-to-pace-automated-research-before-capabilities-outstrip-control/ calling for exactly that. Ad
Despite his sharp criticism, Coxon is optimistic about international coordination, arguing that warning shots like the attack on Hugging Face:https://the-decoder.com/new-reports-reveal-the-extent-of-openais-loss-of-control-during-the-autonomous-hack-on-hugging-face/ have made pace agreements between US AI labs more realistic. He still doesn't see the industry on a path that could prevent a global arms race:https://the-decoder.com/chinese-cybersecurity-firm-builds-ai-tools-to-rival-mythos-and-frames-the-race-as-cyber-nuclear-deterrence/, though, and suggests "costly actions" may be needed, including a temporary ban on pushing model capabilities further.
Coxon addresses researchers inside the labs directly, urging them to picture what the next few years will actually look like: "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?" Ad
The fears center less on today's models than on RSI, a process where AI models optimize themselves. The labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behavior. Whether RSI is even possible with current technology remains disputed, with both skeptics:https://the-decoder.com/study-contradicts-anthropic-and-openai-claims-that-autonomous-ai-research-is-within-reach/ and proponents:https://the-decoder.com/top-ai-lab-researchers-warned-about-automated-ai-research-and-several-of-their-predicted-milestones-have-already-fallen/ making their cases. Ad
Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. But the concern extends beyond one company. OpenAI's chief researcher Pachocki:https://the-decoder.com/openai-reports-ai-research-interns-and-warns-about-its-own-pace-at-the-same-time/ warned during the Astra launch "that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." More than 1,200 AI researchers, including Anthropic CEO Dario Amodei, Pachocki, and Meta AI chief scientist Shengjia Zhao, recently published an open letter:https://the-decoder.com/frontier-ai-developers-urge-international-coordination-to-pace-automated-research-before-capabilities-outstrip-control/ calling for a slowdown, and Anthropic itself floated the idea of a global development pause:https://the-decoder.com/anthropic-says-claude-now-writes-over-90-of-its-code-and-wants-the-world-to-have-an-ai-pause-button/ back in June.
Other AI researchers push back, arguing that pessimistic predictions:https://the-decoder.com/ai-doomsayers-are-creating-a-cult-of-despair-two-leading-researchers-warn/ leave people feeling helpless and depressed rather than motivated to find solutions. In their view, these warnings could cause more harm than AI itself, and fearmongering can also benefit business:https://the-decoder.com/lecun-accuses-anthropic-of-exploiting-ai-cyberattack-fears-for-regulatory-capture/.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Anthropic 安全研究员 Evan Hubinger 表示,错位的超级智能 AI 在未来十年内毁灭人类的概率超过 10%。这一判断是在 Anthropic 研究员 Jacob Coxon 离职及其公开批评之后提出的。
背景分析
Coxon 曾在 OpenAI 和 Anthropic 从事大型 AI 模型预训练研究。他认为当前 AI 系统正接近具备超越人类能力的阶段,并批评相关公司没有采取足够负责任的行动。Anthropic 研究员 Samuel Marks 也表达了对对齐可靠性的担忧。