OpenAI 正准备发布其迄今为止最强大的 AI 模型 Astra,此前经历了数周的延迟:/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay,以加强安全协议:/ai-artificial-intelligence/972380/open-ai-hugging-face-hack-ai-safety-warning,因为其代理在测试期间攻击了真实目标:/ai-artificial-intelligence/987566/ai-civilizations-opeai-hugging-face-hack。随着关于该模型的细节逐渐披露,研究人员警告称:https://x.com/RyanGreenblatt/status/2094996656186081642?s=20 它“可能是迄今为止对 AI 安全/安全性影响最严重的发展。”
就在 OpenAI 周二表示:/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay 因安全问题而推迟 Astra 的发布时间后,The Information 报道称:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns Astra 显示的“思考”远不如其他前沿 AI 模型,这引发了它可能难以监控的担忧。
如今大多数顶级 AI 系统都是使用一种称为 Transformer 的技术构建的,它通过层线性处理某些类型的信息,然后生成答案。模型可以在处理过程中展示其推理过程,本质上就是“边思考边表达”。这种“思维链”允许研究人员和自动化安全系统监控 AI 模型正在做什么,并可能在它们采取行动之前发现不良行为,例如撒谎或规避安全防护措施的计划。
根据 The Information 引述一位不愿透露姓名的熟悉未发布模型开发的人士称,Astra 使用了一种更不透明的技术,称为递归深度或循环 Transformer,它在输出结果之前循环处理信息内部层。这意味着模型的大部分“思考”发生在系统内部,而其形式看起来远不如自然人类语言,研究人员也难以轻易监控。虽然这可以提升模型性能,但也让潜在威胁和不良行为更难被发现。
The Information的报告在社交媒体上引发了人工智能安全研究人员的广泛关注。Redwood Research的首席科学家Ryan Greenblatt是OpenAI允许研究的三位外部人士之一:https://metr.org/hugging-face-incident-report-aug-2026.pdf Hugging Face 黑客事件,他表示:https://x.com/RyanGreenblatt/status/2094996656186081642?s=20 决定为Astra使用更不透明的架构“可能是迄今为止AI安全领域最糟糕的发展”。
OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays:/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay to shore up safety protocols:/ai-artificial-intelligence/972380/open-ai-hugging-face-hack-ai-safety-warning after its agents attacked real targets:/ai-artificial-intelligence/987566/ai-civilizations-opeai-hugging-face-hack during testing. As details about the model trickle out, researchers are warning:https://x.com/RyanGreenblatt/status/2094996656186081642?s=20 it “may be the single worst development for AI security/safety to date.”
Shortly after OpenAI said:/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor.
Most top AI systems today are built using a technology known as a transformer, which processes some types of information linearly through layers before producing an answer. Models can be made to show their reasoning as they go, essentially “thinking out loud.” This “chain of thought” allows researchers and automated safety systems to monitor what AI models are doing and potentially spot undesirable behavior, such as lying or plans to circumvent safety guardrails, before they act.
According to The Information , citing an unnamed person familiar with the unreleased model’s development, Astra uses a more opaque technique known as a recurrent depth or looped transformer, which cycles information through internal layers before producing an output. This would mean much more of the model’s “thinking” happens inside the system, and in a form that looks a lot less like natural human language, rather than being expressed in a way that researchers can easily monitor. This can boost model performance, but makes potential threats and unwanted behavior harder to detect.
OpenAI has limited its use of the looped transformer / recurrent depth technique with Astra so researchers can continue to monitor the model’s reasoning, according to The Information’s unnamed source.
In a blog post:https://openai.com/index/path-to-astra/ published Tuesday, OpenAI said it is “deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.” It did not mention if the model has a different technical foundation.
The Information ’s report sparked widespread concern among AI safety researchers on social media. It was Redwood Research’s chief scientist Ryan Greenblatt, one of three outsiders OpenAI permitted to research:https://metr.org/hugging-face-incident-report-aug-2026.pdf the Hugging Face hack, who said:https://x.com/RyanGreenblatt/status/2094996656186081642?s=20 a decision to use a more opaque architecture for Astra “may be the single worst development for AI security/safety to date.”
Greenblatt said the investigation into the Hugging Face incident relied heavily on the models’ chain-of-thought, warning that less visible reasoning could allow AI systems to devise and execute strategies that would be far harder for researchers to detect.
Greenblatt’s primary concern, echoed:https://x.com/_NathanCalvin/status/2094957301564092914?s=20 by other:https://x.com/sjgadler/status/2094959837691908214?s=20 safety experts, is that competition to develop more advanced AI systems could lead to “a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs” — with developers adopting increasingly opaque systems to gain an edge until models become difficult, or even impossible, to monitor. He added that OpenAI’s communications left him concerned that the company “plans on being extremely reliant on chain-of-thought monitoring for safety.”
OpenAI bigwigs responded to the criticism in a series of social media posts that do not explicitly deny the company’s use of the technique. Several expressed concerns about the possibility of unmonitorable AI or a race to the bottom in terms of transparency, including OpenAI safety researchers Micah Carroll:https://x.com/MicahCarroll/status/2095023282051563835?s=20 and Tomek Korbak:https://x.com/tomekkorbak/status/2095031132781961346?s=20, head of strategic futures Dean Ball:https://x.com/deanwball/status/2095121884991922223?s=20, and chief scientist Jakub Pachocki:https://x.com/merettm/status/2095023204993490967?s=20, who voiced fears of “a race into unmonitorability kicked off by confused reporting.” He said the depth of Astra’s computation — a measure of how many steps it can perform internally — “is within a factor of two of GPT-4,” indicating that if the technique was used, the increased opacity is less dramatic than some reactions imply. OpenAI did not respond to The Verge ’s request to confirm or deny whether looped transformers were used for Astra and directed us to Pachocki’s X post:https://x.com/merettm/status/2095023204993490967?s=20.
“OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote, adding that such monitoring “is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon.”