《The Information》刚刚爆料,OpenAI 正在尝试一种新技术,这种技术会让模型透露更少的“思考”,从而使其更难以监控:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns?rc=dcf9pt&shared=04ff1d10f5a19606。
正如 Zack Korman 和我几天前在这里所论述的,更好的监控可能是防止 Hugging Face 事件发生的因素之一,根据 OpenAI 自己的说法:
他们正在探索的新技术可能会使这种监控变得困难甚至不可能。
去年,一群顶尖学者撰写了一篇非常有趣的论文,现在看来非常相关,题为《思维链可监控性:AI 安全的一个新而脆弱的机会》:https://arxiv.org/abs/2507.11473,我完全同意其中强调的部分:
他们说得完全正确。思维链监控是不完美的(正如 Subbarao Kambhampati:https://rakaposhi.eas.asu.edu/ 和其他人所展示的),但这是我们监控大型黑盒 LLM 的最佳线索之一。虽然线索很细,但为了(小?)性能提升而牺牲它,感觉像是在玩一场危险的游戏。
今晚早些时候,来自 Guidelight.ai 并且是多位离开 OpenAI 安全团队的研究人员之一的 Steven Adler:https://www.dailymail.com/news/article-14335985/OpenAI-steven-adler-quits-warning-risky-gamble-huge-downside-AI-safety.html 说了这句话,并呼应了 Nathan Calvin 的观点:
分享:https://garymarcus.substack.com/p/red-alert-openai-is-poised-to-cross?utm_source=substack&utm_medium=email&utm_content=share&action=share
附注。给喜欢《终结者》梗的朋友们:
你曾说会有一场黑天鹅事件打破 AI 狂潮。它来了。
“不可预测、不可对齐的 LLM 因 OpenAI 需要某个关键而导致主要银行/公共设施倒塌”。
很明显,我们已经彻底完蛋了。
The Information just broke the scoop that OpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns?rc=dcf9pt&shared=04ff1d10f5a19606 .
As Zack Korman and I argued here a few days ago, better monitoring is one of the things that might have prevented the Hugging Face incident, by OpenAI’s own admission:
The new techniques they are exploring may make such monitoring difficult or impossible.
Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, called Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety:https://arxiv.org/abs/2507.11473 , I fully agree with the highlighted bit:
They are exactly right. CoT monitoring is imperfect (as Subbarao Kambhampati:https://rakaposhi.eas.asu.edu/ and others have shown), but it is one of the best threads we have for monitoring the giant black boxes that we call LLM. It is a slender thread, but sacrificing it thread for (small?) performance gain feels like a dangerous game.
Earlier tonight Steven Adler, of Guidelight.ai and one of the many researchers to have departed from OpenAI’s safety teams:https://www.dailymail.com/news/article-14335985/OpenAI-steven-adler-quits-warning-risky-gamble-huge-downside-AI-safety.html , said this, echoing Nathan Calvin:
Share :https://garymarcus.substack.com/p/red-alert-openai-is-poised-to-cross?utm_source=substack&utm_medium=email&utm_content=share&action=share
P.S. Bonus those who prefer Terminator references:
You said there would be a black swan event that would scupper the AI craze. Here it comes.
"Unpredictable, unalignable LLM knocks over major bank/utility because OpenAI needed some key to jangle".
It’s quite clear we are quite fucked