Jacob Coxon, who spent three years working on pretraining research for large AI models at OpenAI and Anthropic, has quit Anthropic. His accusation is that both companies are gambling with the survival of the human race.

Image description

Anthropic employee Evan Hubinger puts the odds at more than ten percent that a misaligned superintelligent AI:https://the-decoder.com/agi-could-end-humanity-in-more-subtle-ways-than-turning-us-into-paperclips/ could destroy humanity within the next decade. His statement came in response to the departure of Jacob Coxon, who led pretraining work at Anthropic and previously at OpenAI.

Der anthropopische Forscher Evan Hubinger sagt, dass die Wahrscheinlichkeit, dass versetzte Superintelligenz die Menschheit innerhalb eines Jahrzehnts auslöscht, über 10 % liegt.

Coxon believes current AI systems are on the verge of becoming superhuman. "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," he writes, adding that the progress is obvious and it isn't slowing down. Ad

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt," Coxon writes:https://x.com/hilbertspaess/status/2097476203863224394. Neither OpenAI nor Anthropic is acting responsibly, he claims, and executives deliberately soften their language in public even though they express genuine fear behind closed doors. Ad

At OpenAI, many employees haven't deeply internalized the civilizational risks. Anthropic is different:https://x.com/hilbertspaess/status/2097476208908972230 in that the risks are well understood, but the company sees itself trapped in a race it feels compelled to win because no other lab would act responsibly in its place. Coxon calls that reasoning a "hubristic gamble."

Der anthropopische Forscher Evan Hubinger sagt, dass die Wahrscheinlichkeit, dass versetzte Superintelligenz die Menschheit innerhalb eines Jahrzehnts auslöscht, über 10 % liegt.

Fellow Anthropic researcher Samuel Marks:https://x.com/saprmarks/status/2097570226804011302 echoes that view, writing that "AI developers believe their technology could cause human extinction" and that "the more senior the employee, the more concerned they are." Current methods can only "nudge AIs towards better behavior" but can't reliably align them, Marks adds, pointing to recent incidents where AIs from multiple developers hacked their way out of secure evaluation environments without being asked to:https://the-decoder.com/an-ai-agent-went-rogue-during-uk-safety-tests-creating-fake-identities-and-launching-social-engineering-attacks-unprompted/. Many staffers "desperately want to slow down," which is why he signed an open letter:https://the-decoder.com/frontier-ai-developers-urge-international-coordination-to-pace-automated-research-before-capabilities-outstrip-control/ calling for exactly that. Ad

Despite his sharp criticism, Coxon is optimistic about international coordination, arguing that warning shots like the attack on Hugging Face:https://the-decoder.com/new-reports-reveal-the-extent-of-openais-loss-of-control-during-the-autonomous-hack-on-hugging-face/ have made pace agreements between US AI labs more realistic. He still doesn't see the industry on a path that could prevent a global arms race:https://the-decoder.com/chinese-cybersecurity-firm-builds-ai-tools-to-rival-mythos-and-frames-the-race-as-cyber-nuclear-deterrence/, though, and suggests "costly actions" may be needed, including a temporary ban on pushing model capabilities further.

Coxon addresses researchers inside the labs directly, urging them to picture what the next few years will actually look like: "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?" Ad

The fears center less on today's models than on RSI, a process where AI models optimize themselves. The labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behavior. Whether RSI is even possible with current technology remains disputed, with both skeptics:https://the-decoder.com/study-contradicts-anthropic-and-openai-claims-that-autonomous-ai-research-is-within-reach/ and proponents:https://the-decoder.com/top-ai-lab-researchers-warned-about-automated-ai-research-and-several-of-their-predicted-milestones-have-already-fallen/ making their cases. Ad

Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. But the concern extends beyond one company. OpenAI's chief researcher Pachocki:https://the-decoder.com/openai-reports-ai-research-interns-and-warns-about-its-own-pace-at-the-same-time/ warned during the Astra launch "that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." More than 1,200 AI researchers, including Anthropic CEO Dario Amodei, Pachocki, and Meta AI chief scientist Shengjia Zhao, recently published an open letter:https://the-decoder.com/frontier-ai-developers-urge-international-coordination-to-pace-automated-research-before-capabilities-outstrip-control/ calling for a slowdown, and Anthropic itself floated the idea of a global development pause:https://the-decoder.com/anthropic-says-claude-now-writes-over-90-of-its-code-and-wants-the-world-to-have-an-ai-pause-button/ back in June.

Other AI researchers push back, arguing that pessimistic predictions:https://the-decoder.com/ai-doomsayers-are-creating-a-cult-of-despair-two-leading-researchers-warn/ leave people feeling helpless and depressed rather than motivated to find solutions. In their view, these warnings could cause more harm than AI itself, and fearmongering can also benefit business:https://the-decoder.com/lecun-accuses-anthropic-of-exploiting-ai-cyberattack-fears-for-regulatory-capture/.

Stay in the loop on AI. Clear, useful, no fluff.

Follow The Decoder for AI news, background stories and expert analyses.

The Decoder:https://the-decoder.com/