The Information just broke the scoop that OpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns?rc=dcf9pt&shared=04ff1d10f5a19606 .

As Zack Korman and I argued here a few days ago, better monitoring is one of the things that might have prevented the Hugging Face incident, by OpenAI’s own admission:

The new techniques they are exploring may make such monitoring difficult or impossible.

Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, called Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety:https://arxiv.org/abs/2507.11473 , I fully agree with the highlighted bit:

They are exactly right. CoT monitoring is imperfect (as Subbarao Kambhampati:https://rakaposhi.eas.asu.edu/ and others have shown), but it is one of the best threads we have for monitoring the giant black boxes that we call LLM. It is a slender thread, but sacrificing it thread for (small?) performance gain feels like a dangerous game.

Earlier tonight Steven Adler, of Guidelight.ai and one of the many researchers to have departed from OpenAI’s safety teams:https://www.dailymail.com/news/article-14335985/OpenAI-steven-adler-quits-warning-risky-gamble-huge-downside-AI-safety.html , said this, echoing Nathan Calvin:

Share :https://garymarcus.substack.com/p/red-alert-openai-is-poised-to-cross?utm_source=substack&utm_medium=email&utm_content=share&action=share

P.S. Bonus those who prefer Terminator references:

You said there would be a black swan event that would scupper the AI craze. Here it comes.

"Unpredictable, unalignable LLM knocks over major bank/utility because OpenAI needed some key to jangle".

It’s quite clear we are quite fucked