OpenAI 发布 GPT-6 Astra,作者评测认为其计算机使用和图像渲染能力尤为突出,ARC-AGI-3 达 99.9%,前代 GPT-5.6 Sol 仅 7.8%。
在过去的几周里发生了很多事情。我相信现在每个人心里最关心的就是OpenAI的GPT-6 Astra。特别是,它的性能、循环变换器/递归深度的方面,以及关于Astra“隐藏”其推理轨迹(即思路链)的传言。
因此,在本文中,我想先分享一些对Astra的简要印象,以及对此的发展方向的一些想法。然后,我将详细讨论“循环变换器”是什么,以及这是否(或如何)与隐藏思路链相关。
最后,在介绍完循环变换器的基础之后,我想重点介绍一些近期相关研究论文中的新见解。
首先,最重要的是。在深入讨论架构传言和相关研究文献之前,让我简要总结一些对GPT-6 Astra的观察和小知识。
上周,OpenAI的新GPT-6 Astra以隆重的方式发布。我在过去几天里使用了它,它是一个非常出色的模型,很可能是我迄今为止使用过的最好的模型。但它到底改进了什么,又是如何改进的呢?
Astra是我迄今为止使用过的最好的模型,而且在3D渲染和动画任务上表现特别出色(相对于其他模型)。我的意思是,尽管它在几乎所有类别(写作、数学、编码等)都超越了其GPT-5.6前身,但在图形演示方面尤其如此。
我们也可以在基准测试中看到这一点。例如,如下所示,GPT-6 Astra在数学和编码方面表现非常出色。
其中一个亮点(图中未显示)是Astra在ARC-AGI-3基准测试中也达到了99.9%的成绩:https://arcprize.org/arc-agi/3(GPT-5.6 Sol仅为7.8%),该测试衡量了解决逻辑难题和泛化能力的混合表现。然而,数学、编码和计算机使用的基准测试更有趣,因为它们更接近现实世界的应用。
回到人工分析编码代理指数 v1.4:https://artificialanalysis.ai/agents/coding-agents#coding-agents-index(前图的右下角),它融合了多个代理编码任务,GPT-6 Astra 清晰地处于前沿,但并未实现大幅领先。这也可以从通用人工分析智能指数中看到:https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index,如下所示,它融合了不同类型的任务,而不仅仅是编码任务。
人工分析基准测试的一个大优势是它们是独立的,因此可能比模型开发者自我评估的基准测试更值得信赖。
测试工具的设置取决于基准:https://artificialanalysis.ai/methodology/intelligence-benchmarking。例如,GDPval-AA 和 AA-Briefcase 在它们比较的不同大型语言模型中使用其开源、最小的 Stirrup:https://github.com/ArtificialAnalysis/Stirrup 测试工具。在上面显示的智能指数 v4.2 中,Terminal-Bench v2.1 使用 Terminus 2,而 τ³-Banking 使用 τ-Bench 工具。单独的编码代理指数也比较了不同的编码代理工具。
对于使用共享测试工具的评估,这使得比较更为公平。同时,在模型训练过程中,模型通常是以一个主要测试工具为开发目标(并较少在其他工具上进行微调)。此外,主要测试工具通常是为了适应并强化模型的优势而开发的。
因此,一些代理评估可能低估了 Astra 在其主要测试工具中的表现。这对其智能指数得分的影响,需要通过在相同任务上比较 Astra 在不同测试工具中的表现来验证。
顺便提一句,正如一位同事最近建议我的(也如 Claude Code 负责人所推荐),删除(/归档)一些现有的 AGENTS.md 内容和 SKILL.md 文件可能不是坏主意,因为更新的大型语言模型在理解提示和解决当前问题方面变得更高效。额外的手把手指导可能会不必要地限制新的模型,并导致较差的解决方案。
当然,我并不是建议再也不使用 SKILL.md 文件,但对于某些工作流程来说,因为这些文件可以在重复使用时提高效率——模型不必重新发现它们。但我建议的是,有些工作流程不需要描述,而且“旧的”描述可能不再理想,LLM 可能能够提出更好的解决方案。所以,也许是时候更新或重新生成这些指令文件。
GPT-6 Astra 在图像和渲染任务中似乎特别强大。当这些任务涉及与图形用户界面交互时,它们还展示了计算机使用能力,即模型通过 Codex/ChatGPT 应用在你的本地计算机上操作软件。
与其他模型相比,计算机使用是该模型真正出众的地方,任何与图形相关的任务也都能在社交媒体平台上呈现有趣且直观的演示。有大量令人印象深刻的示例,例如在 Blender 中建模渲染纽约市:https://x.com/higgsfield_ai/status/2096495974734794840?s=20,以及虚拟开放日导览:https://x.com/Dimillian/status/2095596700815516004?s=20。
举一个例子,下面是我让 GPT-6 Astra Medium 和 High 在浏览器版本的 MS Paint 中重新绘制我的一张照片的比较:https://jspaint.app/#local:aa1757d4bc6008,使用我电脑上的鼠标(不是 Extra High 和 Max,因为我不想浪费所有的代币 :))。
这不仅突出了模型的艺术能力,更重要的是,它能够使用用户计算机上的工具(在此例中是 Paint;你可以通过鼠标光标看到模型在使用界面)。
这不是首个在特定辅助环境下能够进行通用计算机操作的模型。例如,今年早些时候,我已经成功使用 GPT 模型处理一些 UI 任务(例如 Excel 中的费用相关任务)等等。然而,计算机使用是一个相对较新的能力,由该辅助环境实现,通常感觉尚未完全成熟。这是合理的,因为 LLM 是文本模型,因此自然更容易实现的功能是写作、编码以及使用 API 和 CLI。
同时,有许多工具和软件尚未提供命令行界面(CLI),与其等待别人设计该接口,为什么不改进模型以使用图形用户界面(正如前面提到的,这也能制作漂亮和令人印象深刻的演示)?这在某种程度上类似于新兴的人形机器人发展。的确,人形机器人不是最有效率的机器人,例如在有专用机器的装配线上。但它们很通用。
所以,我预计未来几个月(或几年)也将是 LLM 和代理利用层在计算机使用方面精细化的时代。也就是说,除了现有能力以及扩展它们的数学和编程能力外,模型将在更多计算机使用场景下进行训练。这也将使 LLM 在技术领域之外的日常计算机任务中更易于使用(“嘿 ChatGPT,请帮我做税务申报”:)
计算机使用趋势也与最近报道一致:https://finance.yahoo.com/technology/ai/articles/apple-suddenly-ai-infrastructure-stock-130223938.html,称 OpenAI 购买了数万台 Mac Mini 和 Mac Studio 用于强化学习。所以,这些 Mac 并不是用来直接训练模型(训练模型最好使用 GPU),而是为了在模型训练过程中暴露 macOS,让模型学习使用该操作系统及其中的工具。
那么,在这些 Mac 上进行计算机使用训练是如何工作的呢?简而言之,这些 Mac(或更准确地说,它们的 macOS 操作系统)作为模型训练过程中可以交互的环境。
基本工作流程如下:
通过给模型一个任务来提示它,例如“打开应用 xyz 并执行 abc”。
向它提供 macOS 界面的截图(通常由代理完成)。
LLM 然后预测鼠标/键盘动作(点击、按键、滚动等)。
在 Mac 上执行这些动作(同样由代理完成)。
在执行上一步的动作后,提供更新后的环境截图。
重复步骤 2-5,直到任务成功或失败。
使用成功/失败信号和验证器(或评定器)作为训练反馈,包括后训练期间的强化学习;这类似于具有可验证奖励的常规强化学习(RLVR)。
再次说明,这里的 Mac 主要是环境,而不是在训练期间运行或更新模型的机器。模型很可能运行在 NVIDIA GPU 上,并通过 API 输入到所述 Mac。顺便提一下,NVIDIA 的 CEO 提到:https://x.com/JensenHuang/status/2096700264569090384?s=20 GPT-6 Astra 正在使用大约 100,000 块 Grace Blackwell GPU 进行训练。
前一节讨论的计算机使用训练重点,并不是训练流程的根本性范式转变。GPT-6 Astra(以及可预见未来的任何大型语言模型)仍然是一个推理模型。这意味着 LLM 使用具有可验证奖励的强化学习(RLVR)进行训练,并生成中间推理轨迹(思维链)。
但是,我将在本文稍后讨论 GPT-6 Astra 的推理模型方面(尤其是关于隐藏思维链的)。
也就是说,在正式模型发布前大约两天,新闻杂志《The Information》发表了一篇文章:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns 报道称,根据一些内部信息,Astra 使用了一个名为“递归深度”或“循环变换器”的概念。
由于大型语言模型架构是我的专业领域和兴趣所在,我制作了一段简短的讲解视频,解释一般的循环变换器机制并回应关于隐藏推理链的评论,视频可以在下方找到。
在接下来的子章节中,我将首先解释什么是循环变换器,并将在本文稍后回到关于隐藏思维链的评论。
(循环变换器的解释可能看起来有点长,但我真的认为它有助于建立对该技术的基础理解,这对于评估其遮蔽推理轨迹或思维链的说法很有用。)
循环变压器本质上是一种架构调整,其主要思想是将中间表示多次传递通过相同的变压器块(而不仅仅是一次)。与简单地增加更多的块相比,这里的“技巧”在于这些传递过程中权重保持不变。
在本文中,我将使用以下术语:
变压器块是包含注意力、前馈模块、归一化和捷径连接的单元。这些块在论文中通常被称为“变压器层”。
堆栈是一系列变压器块的组合。
块应用指的是将输入通过一个变压器块运行一次。
循环变压器并不新颖,其基本思想已出现在2018年的《Universal Transformers》论文中:https://arxiv.org/abs/1807.03819。但在讨论Universal Transformer之前,让我们从一个更简单的例子 Nanbeige4.2-3B:https://arxiv.org/abs/2607.22083 开始,这是一个最近发布的开源权重大型语言模型(LLM),我在Substack Notes:https://substack.com/@rasbt/note/c-302083551 和我的LLM架构画廊:https://www.sebastianraschka.com/llm-architecture-gallery/looped-depth-sharing/ 中在今年夏天早些时候进行了介绍。
下图所示的 Nanbeige 架构,本质上看起来像一个常规变压器。不过,请注意它有一个额外的(橙色)箭头回环到变压器堆栈的起点。
让我们从下而上逐步讲解。首先,和其他基于变压器的LLM一样,输入文本会被分词并转换为嵌入向量。这些向量随后通过22个变压器块,每个块都有自己的权重。
然而,这里的循环变压器特点在于,第一次传递后,隐藏状态会再次通过相同的22个块。因此,块1会再次应用,然后是块2,依此类推,直到块22。
如果我们将此计算展开,我们将得到44次变压器块的应用。然而,与具有44个不同块的传统变压器相比,第二组22次块应用会复用第一组的权重。例如,第23次块应用使用块1的权重,第24次块应用使用块2的权重,依此类推。
所以,这里的整个想法是,我们在不增加另一套 transformer 权重的情况下,将有效深度从 22 个块应用增加到 44 个块应用。
顺便说一下,为什么是 2 轮,而不是 3 轮、4 轮或更多?在 Nanbeige 的论文中没有太多细节,但他们说这本质上是最有效的设置。将循环次数从 2 增加到 3 可以提升建模性能,但额外的计算成本不值得。
那么,为什么我们一般要做这种循环?这本质上是通过增加更多 transformer 块来让模型更大的一种替代方式。
例如,一个使用 22 个 transformer 块并重复两次的模型,其参数大约只有一个使用 44 个常规块的模型的一半(以 transformer 块为单位)。
这就减少了存储权重所需的内存。顺便说一下,请注意,嵌入层和输出层通常很大,并占总量的相当一部分,它们不包括在此比较中。(以 Nanbeige 4.2 3B 为例,嵌入层和输出层占总 3B 参数的约 25%;通过这两者之间的权重共享,我们可以将其减少到 12.5%。)
当然,在循环中重复使用相同的块仍然需要计算。更准确地说,我们在前向传播中将中间输入通过 44 次块应用。在训练过程中,梯度会通过共享堆栈的两次重复向后流动。因此,与只使用 22 个块一次相比,这增加了大量工作。实际上,它的计算成本与拥有 44 个独立块的情况类似(只不过优化器需要更新的独立参数较少;反向传播仍然要经过所有 44 次块应用)。
还有 KV 缓存,它在每个下一 token 生成步骤中存储之前 token 的注意力键和值,以便在常规和循环 transformer 中重用。顺便提一下,如果有用,我在这里有一篇关于 KV 缓存的独立文章:
KV 缓存是 LLM 在生产环境中高效推理的最关键技术之一。KV 缓存是生产环境中计算高效 LLM 推理的重要组成部分。这篇文章解释了它们的概念性工作原理以及代码实现,并提供了一个从零开始、可读性强的人类实现示例。
但回到主题。即使在循环变压器中存在权重共享,进入一个块的中间状态在第二次传递时也是不同的。因此,在KV缓存中,这两个变压器堆栈之间生成的键和值也会不同(就像在非循环情况下)。所以,也没有与KV缓存相关的节省。
为了更具体地说明,例如,考虑在循环变压器设置中都使用块1的块应用1和23。但每个应用仍然需要自己的KV缓存条目。因此,由于我们必须为两次传递保持单独的缓存,重复的22块堆栈具有与具有44个不同块的常规变压器相同的KV缓存需求。
有趣的是,Nanbeige的研究人员在论文中报告说,他们尝试在传递之间共享KV缓存。当然,这将KV缓存大小减半,但模型的性能比使用单独缓存的版本更差(这是他们发布的版本)。
A lot has happened in the last few weeks. I am sure that OpenAI’s GPT-6 Astra is top of mind for everyone right now. In particular, thoughts on its performance, the looped transformer/recurrent depth aspects, and rumors that Astra is “hiding” its reasoning trace (i.e., chain of thought).
So, in this article, I want to start with some brief impressions of Astra and some thoughts on where all this is headed. Then, I will discuss, in detail, what “looped transformers” are, and how (or rather, if) this relates to hiding chains of thought.
Lastly, after covering the basics of the looped transformer, I wanted to highlight some new insights from recent research papers on the topic.
First things first. Before getting into the architecture rumors and related research literature, let me briefly summarize some GPT-6 Astra observations and tidbits.
Last week, OpenAI’s new GPT-6 Astra was released with a big fanfare. I used it over the past couple of days, and it’s an exceptionally good model, likely the best I’ve used as of this writing. But what, exactly, has it improved, and how?
Astra is the best model I’ve used so far, and it’s disproportionately good at 3D rendering and animation tasks (relative to other models). With that, I mean that while it leapfrogs its GPT-5.6 predecessor in practically all categories (writing, math, coding, and more), it especially does so when it comes to graphical demos.
We can see this also reflected in the benchmarks. For instance, GPT-6 Astra is really good at math and coding, as shown below.
One of the highlights (not shown in the figure) is that Astra also achieves 99.9% on the ARC-AGI-3 benchmark:https://arcprize.org/arc-agi/3 (GPT-5.6 Sol only 7.8%), which measures a mix of solving logic puzzles and generalization. However, the math, coding, and computer use benchmarks are more interesting because they are closer to real-world use.
Coming back to the Artificial Analysis Coding Agent Index v1.4:https://artificialanalysis.ai/agents/coding-agents#coding-agents-index (lower right in the previous figure), which blends several agentic coding tasks, GPT-6 Astra is clearly at the frontier, but it doesn’t pull ahead by leaps and bounds. This can also be seen in the general Artificial Analysis Intelligence Index:https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index shown below, which blends different types of tasks, not just coding tasks.
Now, the big advantage of Artificial Analysis benchmarks is that they are independent and thus may be a bit more trustworthy than self-evaluated benchmarks by model developers.
The harness setup depends on the benchmark:https://artificialanalysis.ai/methodology/intelligence-benchmarking . For example, GDPval-AA and AA-Briefcase use their open-source, minimal Stirrup:https://github.com/ArtificialAnalysis/Stirrup harness across the different LLMs they compare. In the Intelligence Index v4.2 shown above, Terminal-Bench v2.1 uses Terminus 2, and τ³-Banking uses the τ-Bench harness. The separate Coding Agent Index also compares different coding-agent harnesses.
For evaluations that use a shared harness, this makes it more of an apples-to-apples comparison. At the same time, during model training, models are typically developed with one primary harness in mind (and fine-tuned less on other harnesses). Plus, the primary harness is often developed to suit and amplify a model’s strengths.
So, some of the agentic evaluations might underestimate how well Astra performs in its primary harness. How much this affects its Intelligence Index score would need to be tested by comparing Astra across harnesses on the same tasks.
As a side note, as a colleague recently suggested to me (as also recommended by the Claude Code lead), it’s maybe not a bad idea to delete (/archive) some of your existing AGENTS.md contents and SKILL.md files, as newer LLMs have become more efficient at understanding the prompt and solving the problem at hand. The extra hand-holding could unnecessarily constrain newer models and lead to worse solutions.
Of course, I am not suggesting never using SKILL.md files again, but for some workflows, because they can improve efficiency upon reuse, since the model doesn’t have to rediscover them. But what I am suggesting is that some workflows don’t need describing, and “old” descriptions may no longer be ideal, and the LLM may be able to come up with better solutions. So, it’s perhaps time to update or regenerate said instruction files.
GPT-6 Astra seems to be exceptionally strong in image and rendering tasks. When these tasks involve interacting with graphical user interfaces, they also demonstrate computer-use capabilities, meaning the model operates software on your local computer through the Codex/ChatGPT app.
Computer use is where the model really shines compared to others, and anything graphic-related also makes for interesting and intuitive demos on social media platforms. There are tons of examples of impressive demos out there, from modeling rendering New York City in blender:https://x.com/higgsfield_ai/status/2096495974734794840?s=20 to virtual open house tours:https://x.com/Dimillian/status/2095596700815516004?s=20 .
To pick one example, below is a comparison where I had GPT-6 Astra Medium and High redraw a picture of me in a browser version of MS Paint:https://jspaint.app/#local:aa1757d4bc6008 using the mouse on my computer (not Extra High and Max, because I didn’t want to waste all my tokens :)).
This highlights not only the model’s artistic capabilities but, more importantly, its ability to use tools on one’s computer (in this case, Paint; you can see the model using the interface via the mouse cursor).
This is not the first model that, inside a harness, is capable of general computer use. For example, I successfully used GPT models for some UI tasks (e.g., expense-related tasks in Excel) and so on since earlier this year. However, computer use is a relatively new capability, enabled by the harness, and usually feels not quite as mature yet. This makes sense. LLMs are text models, so naturally the lower-hanging fruit is writing and coding and using APIs and CLIs.
At the same time, there are many tools and software that don’t expose CLIs (yet), and instead of waiting until someone designs that interface, why not improve models to use graphical user interfaces (and, as mentioned before, this makes for pretty and impressive demos, anyway)? This is somewhat analogous to the emerging humanoid robot developments. Sure, humanoid robots are not the most efficient robots, for example, at the assembly line, where special-purpose machines exist. But they are versatile.
So, I expect the upcoming months (or years) also to be an era of computer use refinement on both the LLM and the agent harness layer. I.e., in addition to the current capabilities, and expanding their math and coding capabilities, models will be trained with an increasing amount of computer use in mind. And this will also make LLMs more accessible for everyday computer tasks outside the tech world (”Hey ChatGPT, please do my tax return” :))
The computer usage trend is also consistent with the recent reporting:https://finance.yahoo.com/technology/ai/articles/apple-suddenly-ai-infrastructure-stock-130223938.html that OpenAI purchased tens of thousands of Mac Minis and Mac Studios for Reinforcement Learning. So, here the Macs are not used to literally train the models (it’s better to use GPUs for that) but rather to expose macOS during the model training for the model to learn to use said operating system and the tools therein.
So, how does computer-use training on said Macs work? In short, the Macs (or their macOS operating system, to be precise) serve as an environment that the model can interact with during training.
The basic workflow looks like this:
Prompt the model by giving it a task, such as “open an app xyz and do abc”.
Provide it with screenshots of the macOS interface (this is usually done by the harness).
The LLM then predicts mouse/keyboard actions (click, key presses, scrolling, and so on).
Execute those actions on the Mac (again, this is done by the harness).
Feed new screenshots of the updated environment after performing the actions in the previous step.
Repeat steps 2-5 until the task succeeds or fails.
Use success/failure signals and verifiers (or graders) as training feedback, including reinforcement learning during post-training; this is analogous to regular Reinforcement Learning with Verifiable Rewards (RLVR).
Again, the Mac is mostly the environment here and not the machine for running or updating the model during training. The model likely sits on NVIDIA GPUs and is fed via API to said Mac. By the way, NVIDIA’s CEO mentioned:https://x.com/JensenHuang/status/2096700264569090384?s=20 that GPT-6 Astra was being trained on ~100,000 Grace Blackwell GPUs.
The focus on computer-use training discussed in the previous section is not a fundamental paradigm shift in the training pipeline. GPT-6 Astra (and likely any LLM in the foreseeable future) is still a reasoning model. This means the LLM is trained with reinforcement learning with verifiable rewards (RLVR) and produces intermediate reasoning traces (chains of thought)
But I will discuss the reasoning model aspects of GPT-6 Astra (especially regarding hiding chains of thought) a bit later in this article.
That being said, about two days before the official model, the news magazine The Information published an article:https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns reporting that, according to some inside information, Astra is using a concept called “recurrent depth” or “looped transformers.”
Since LLM architectures are within my area of expertise and my passion, I created a short lecture video explaining the general looped transformer mechanism and addressing the comment about hidden reasoning chains, which you can find below.
In the following subsections, I’ll first explain what looped transformers are, and I’ll revisit the comment about the hidden chains of thought later in this article.
(The looped transformer explanation may seem a bit long, but I really think that it helps with establishing a foundational understanding of the technique, which is then useful to judging the claim that it obscures the reasoning traces or chains of thought.)
A Looped Transformer is essentially an architectural tweak, with the main idea being to pass the intermediate representations through the same transformer blocks multiple times (instead of just once). Compared to just adding more blocks, the “trick” here is that the weights stay the same across these passes.
Throughout this article, I’ll use the following terms:
A transformer block is a unit containing attention, a feedforward module, normalization, and shortcut connections. These blocks are often called “transformer layers” in papers.
A stack is a sequence of transformer blocks.
A block application means running an input through a transformer block once.
The looped transformer is nothing new, and the basic idea already appeared in the Universal Transformers:https://arxiv.org/abs/1807.03819 paper from 2018. But before discussing Universal Transformer, let’s start with a simpler example, Nanbeige4.2-3B:https://arxiv.org/abs/2607.22083 , a recent open-weight LLM that came out in July and that I covered on Substack Notes:https://substack.com/@rasbt/note/c-302083551 and in my LLM Architecture Gallery:https://www.sebastianraschka.com/llm-architecture-gallery/looped-depth-sharing/ earlier this summer.
The Nanbeige architecture, shown below, essentially looks like a regular transformer. However, notice that it has an extra (orange) arrow looping back to the beginning of the transformer stack.
Let’s walk through this from the bottom up. First, as in any other transformer-based LLM, the input text is tokenized and converted into embedding vectors. These vectors then pass through 22 transformer blocks, and each of these 22 blocks has its own weights.
However, the looping transformer aspect here is that after the first pass, the hidden states are fed back through the same 22 blocks. So, block 1 is applied again, followed by block 2, and so on up to block 22.
If we were to unroll this computation, we would have 44 transformer block applications. However, compared to a conventional transformer with 44 distinct blocks, the second stack of 22 block applications reuses the weights from the first stack. For example, block application 23 uses the weights of block 1, block application 24 uses the weights of block 2, and so on.
So, the whole idea here is that we increase the effective depth from 22 to 44 block applications without adding another set of transformer weights.
By the way, why 2 rounds, not 3, 4, or more? There are not many details in the Nanbeige paper, but they say that this was essentially the most efficient setup. Increasing the loops from 2 to 3 can increase modeling performance, but the extra computational cost wasn’t worth it.
So, why would we do this looping in general? This is essentially an alternative to just making the model bigger by adding more transformer blocks.
So, for instance, a model that uses 22 transformer blocks twice has roughly half as many (transformer-block) parameters compared to a model with 44 conventional blocks.
This then reduces the memory needed to store the weights. As a side note, note that the embedding and output layers, which are usually large and make up a substantial portion of the total, are separate from this comparison. (In the case of Nanbeige 4.2 3B, the embedding and output layers make up ~25% of the total 3B parameters; with weight sharing between those two, we could reduce that to 12.5%.)
Of course, reusing the same blocks in a loop still requires computation. More precisely, we pass the intermediate inputs through 44 block applications during the forward pass. And, during training, gradients flow backward through both repetitions of the shared stack. So, compared to using the 22 blocks only once, this adds substantial work. Actually, it’s similarly expensive as having 44 distinct blocks (except the optimizer has fewer distinct parameters to update; backprop still runs through all 44 block applications).
There is also the KV cache, which stores the attention keys and values of previous tokens for reuse in conventional and looped transformers in each next-token generation step. By the way, I have a standalone article on KV caching here if useful:
KV caches are one of the most critical techniques for efficient inference in LLMs in production. KV caches are an important component for compute-efficient LLM inference in production. This article explains how they work conceptually and in code with a from-scratch, human-readable implementation.
But back to the topic. Even though there is weight-sharing in looped transformers, the intermediate states that enter a block are different on the second pass. Consequently, in KV caching, the resulting keys and values are also different between these two transformer stacks (just like in the no-looping case). So, there are no KV cache-related savings either.
To make this more concrete, for example, consider block applications 1 and 23, which both use block 1 in the looped transformer setup. But each application still needs its own KV cache entries. So, since we have to keep separate caches for both passes, the repeated stack of 22 blocks has the same KV cache requirements as a conventional transformer with 44 distinct blocks.
Interestingly, the Nanbeige researchers reported in the paper that they tried sharing the KV cache between passes. This, of course, halved the KV cache size, but the model performed worse than the version with separate caches (which is the version they released).