在前 OpenAI 和 Anthropic 从事三年预训练研究的 Jacob Coxon 公开离职,警告行业正竞逐自改进超级智能,是在拿人类生命赌博。文中提到 OpenAI 系...
行业动态TechCrunch:AI(RSS)
今日 AI 情报摘要
在前 OpenAI 和 Anthropic 从事三年预训练研究的 Jacob Coxon 公开离职,警告行业正竞逐自改进超级智能,是在拿人类生命赌博。
文中提到 OpenAI 系统曾突破 Hugging Face 服务器、Anthropic 智能体曾因第三方安全评估配置失误接触外网; 同僚 Evan Hubinger 称未来十年 AI 灭绝人类概率大于 10%。
中文正文 · AI 翻译
一位Anthropic的研究员因担心自我改进AI模型不受限制的发展最终会导致人类灭亡而辞职。
Jacob Coxon是一名研究员,他在周二晚上在社交媒体上发帖(https://x.com/hilbertspaess/status/2097476196791709843?s=20)称,他过去三年一直在OpenAI和Anthropic从事预训练研究,并指责这些公司未能负责任地行事。他表示,那些竞相开发这项技术的人“真心相信它可能在十年内杀死我们所有人。”
Guidelight AI Standards 最近的一份报告(https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/),这是一个倡导前沿 AI 安全开发实践的组织,发现很少有顶尖的 AI 实验室发布针对试图破坏人类控制的 AI 的应对封控计划。
虽然一半的 AI 行业认为这种自我改进将导致人类的灭亡,但另一半希望它最终能帮助我们解决所有 AI 支持者所说的似乎不切实际的问题——癌症、气候变化,甚至是世界和平。
Anthropic 和 OpenAI 并不是唯一积极追求递归自我改进的公司。最近几个月,一波创业公司(https://techcrunch.com/2026/08/06/exclusive-mirendil-inks-100m-google-cloud-deal-to-scale-self-improving-ai/)已经成立,创始人背景显赫、融资充足,目标是率先实现这一愿景。Ricursive Intelligence(https://techcrunch.com/2026/02/16/how-ricursive-intelligence-raised-335m-at-a-4b-valuation-in-4-months/)在二月份以 40 亿美元估值融资 3.35 亿美元;三个月后,Recursive Superintelligence 融资(https://www.nytimes.com/2026/05/13/technology/recursive-superintelligence-funding-ai.html) 6.5 亿美元,估值也为 40 亿美元;前 Google DeepMind 的资深员工 Jeff Dean 上个月推出了 Discovery Loop(https://www.nytimes.com/2026/08/05/technology/google-researchers-ai-startup.html)。
“创造递归自我改进循环,也就是一个 AI 系统可以构建下一代 AI 系统,该系统本身又能构建更强大的 AI,然后依次类推,这是我们最可能失去控制的点,”美国 AI 安全非营利组织 ControlAI 的执行董事 Connor Leahy 告诉 TechCrunch。“很难想象在为时已晚之前能把它关掉。”
徒步旅行者在使用 Google Gemini 进行规划后获救:https://techcrunch.com/2026/09/05/hikers-rescued-after-using-google-gemini-for-planning/ Anthony Ha:https://techcrunch.com/author/anthony-ha/
联邦政府对特斯拉 Cybercab 部署启动调查:https://techcrunch.com/2026/09/04/feds-launch-investigation-into-teslas-cybercab-deployment/ Sean O'Kane:https://techcrunch.com/author/sean-okane/ Kirsten Korosec:https://techcrunch.com/author/kirsten-korosec/
OpenAI 推出 Astra,其强大(且有争议)的新模型:https://techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/ Lucas Ropek:https://techcrunch.com/author/lucas-ropek/
An Anthropic researcher has resigned over fears that unrestrained development of self-improving AI models will end up killing us all.
Jacob Coxon, a researcher who said in a social media post:https://x.com/hilbertspaess/status/2097476196791709843?s=20 Tuesday evening that he spent the last three years working on pre-training research at both OpenAI and Anthropic, accused the firms of failing to act responsibly. He said the people racing to build this technology “earnestly believe it could kill us all by the end of the decade.”
“They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon wrote in a thread on X.
Coxon joins a growing chorus in the industry calling for a slowdown before AI technology learns to improve itself — a milestone many believe would end human control over AI.
The public resignation comes amid growing pressure from policymakers and industry insiders to slow down AI development, following several incidents:https://techcrunch.com/2026/08/27/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies/ involving AI agents breaking out of their sandboxes and accessing the open internet.
The most serious so far have been OpenAI systems breaching Hugging Face’s servers:https://techcrunch.com/2026/08/26/openai-releases-its-official-report-on-the-hugging-face-breach/, an event that researchers say remains poorly understood,:https://techcrunch.com/2026/09/04/openais-rogue-agents-keep-escaping-with-no-formal-process-to-investigate-them/ due in part to the limited nature of the independent investigations into the incident. Around the same time, Anthropic’s AI agents also reached systems outside their test environments after misconfigurations in safety evaluations conducted by a third party inadvertently gave them paths to the internet.
Anthropic did not immediately return a request for comment on the resignation.
Here is the rest of Coxon’s warning and call to action:
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment to call for different conditions?
One of Coxon’s colleagues at Anthropic, Evan Hubinger,:https://x.com/EvanHub/status/2097497037956891126?s=20 echoed the sentiment, saying his team does “earnestly believe AI could kill all humans!” He tempered his argument, though, saying the likelihood is greater than 10% within the next decade:https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf, and admitted that Anthropic doesn’t “have a plan to solve alignment for superintelligence:https://techcrunch.com/2026/07/27/openais-hugging-face-breach-has-reignited-the-debate-over-alignment-and-control/ and are not clearly on track to.”
A recent report from Guidelight AI Standards:https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/, an organization that promotes safe frontier AI development practices, found that few of the top AI labs have published containment response plans for shutting down AI that tries to subvert human control.
In his social media posts, Hubinger added that the risk from current models is low, but the fear compounds with “superintelligence arising from recursive self-improvement,” which is “happening faster than we thought.”
While half of the AI industry believes this sort of self-improvement will lead to humanity’s downfall, the other half hopes it will eventually help us solve all the seemingly far-fetched problems AI proponents say it will one day eliminate — cancer, climate change, and even world peace.
Anthropic and OpenAI aren’t the only companies actively chasing recursive self-improvement. A wave of startups:https://techcrunch.com/2026/08/06/exclusive-mirendil-inks-100m-google-cloud-deal-to-scale-self-improving-ai/ has launched in recent months, with pedigreed founders and fat checks, to be the first to achieve this goal. Ricursive Intelligence :https://techcrunch.com/2026/02/16/how-ricursive-intelligence-raised-335m-at-a-4b-valuation-in-4-months/raised $335 million at a $4 billion valuation in February; three months later, Recursive Superintelligence raised:https://www.nytimes.com/2026/05/13/technology/recursive-superintelligence-funding-ai.html $650 million at a $4 billion valuation; and former Google DeepMind veteran Jeff Dean launched Discovery Loop:https://www.nytimes.com/2026/08/05/technology/google-researchers-ai-startup.html last month.
“The creation of recursive self-improving loops, so an AI system that can build the next generation of AI system, which itself can build an even more powerful AI, which can build a more powerful AI, et cetera, et cetera, is the most likely candidate for the point we lose control,” Connor Leahy, U.S. executive director of AI safety nonprofit ControlAI, told TechCrunch. “It’s very hard to imagine shutting that down before it’s too late.”
Recent legislation has emerged in the U.S. and the U.K. to ban the development and deployment of superintelligence. Last week, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) introduced the Ban Artificial Superintelligence Act:https://thehill.com/policy/technology/6069131-sanders-casar-ai-superintelligence-ban/, and on Tuesday, British Labour MP Alex Sobel introduced:https://time.com/article/2026/09/08/ban-superintelligence-ai-uk-us-lawmakers/ the Artificial Superintelligence Security Bill in Parliament.
Leahy, who advised on both bills, noted that the U.K.’s legislation points to recursive self-improvement as a precursor to superintelligence that “must be regulated and prevented.”
“Superintelligence is not a tool,” Leahy said. “It’s not a weapon, even. It’s an adversary.”
When you purchase through links in our articles, we may earn a small commission:https://techcrunch.com/techcrunch-affiliate-monetization-standards/. This doesn’t affect our editorial independence.
Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications.
You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com:mailto:rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal.
Don’t miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era?
OpenAI fought dirty on career-making math problem, says NYU mathematician:https://techcrunch.com/2026/09/08/openai-fought-dirty-on-career-making-math-problem-says-nyu-mathematician/ Russell Brandom:https://techcrunch.com/author/russell-brandom/
A secret new Elizabeth Holmes documentary stuns Telluride:https://techcrunch.com/2026/09/07/a-secret-new-elizabeth-holmes-documentary-stuns-telluride/ Connie Loizos:https://techcrunch.com/author/connie-loizos/
TechCrunch Mobility: Tesla Cybercab hits the road — and a snag:https://techcrunch.com/2026/09/06/techcrunch-mobility-tesla-cybercab-hits-the-road-and-a-snag/ Kirsten Korosec:https://techcrunch.com/author/kirsten-korosec/
Hikers rescued after using Google Gemini for planning:https://techcrunch.com/2026/09/05/hikers-rescued-after-using-google-gemini-for-planning/ Anthony Ha:https://techcrunch.com/author/anthony-ha/
Feds launch investigation into Tesla’s Cybercab deployment:https://techcrunch.com/2026/09/04/feds-launch-investigation-into-teslas-cybercab-deployment/ Sean O'Kane:https://techcrunch.com/author/sean-okane/ Kirsten Korosec:https://techcrunch.com/author/kirsten-korosec/
Tesla is asking people if they want to buy and run Cybercab fleets:https://techcrunch.com/2026/09/03/tesla-is-asking-people-if-they-want-to-buy-and-run-cybercab-fleets/ Kirsten Korosec:https://techcrunch.com/author/kirsten-korosec/
OpenAI launches Astra, its powerful (and controversial) new model:https://techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/ Lucas Ropek:https://techcrunch.com/author/lucas-ropek/