AI 算力现货价格自 2 月低点已上涨 40% 以上,Google 和 Anthropic 从 SpaceX 租用 11 万块 GPU 的月租金达 9 亿美元,约为现货价格的 2 倍。
我想尝试写一篇非常快速的博客文章,把写作时间限定为两小时。我可能无法解决很多重要的子问题,但另一种选择就是在我感兴趣的许多不同主题上完全没有进展。
今天我想谈谈未来几年实验室的计算情况。
Anthropic 的收入同比增长了 10 倍。Anthropic 今年可能以大约 100–150 亿美元的收入结束年底。要让这一趋势继续,Anthropic 到明年年底的收入必须达到 1 万亿美元。没有深层次的理由表明这一趋势必须继续,而它很可能不会——归根结底,这是一个关于人工智能能力的问题。但假设它会继续。那个世界必须满足什么条件?
实验室计算资源每年增长 3 倍:https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute。要让一个实验室在收入增长 10 倍的同时计算资源仅增长 3 倍,必须发生以下三种可能组合之一:1. 实验室利润率必须增加,2. 计算价格必须上涨,3. 实验室必须将更多的计算资源用于推理。
据我了解,这三件事基本上都在发生:1. Anthropic 的利润率从 2025 年的 40% 提升到今年 Fable 推理 1 的可能超过 80% #footnote-1(虽然边际计算可能没有达到这个水平——见下文更多内容)。2. 计算的现货价格比二月份的低谷上涨了 40% 以上,这很可能低估了实验室需要支付的费用(下文会详细说明)。3. 根据 Epoch 的数据,OpenAI 2024 年大约四分之一的计算支出用于推理:https://epoch.ai/data-insights/openai-compute-spend,现在这个比例肯定接近 50%,如果不是更高的话。
实验室通常不愿意做第3点(将越来越多的计算资源用于推理)。正如人们开玩笑所说的,推理收入的意义在于说服投资者给你更多资金,以购买更多计算资源来训练更大的模型。如果你把大部分计算资源都花在推理上,你实际上是在宣告人工智能进展停滞了,因为再投资训练已经不值得,而你的业务基本上就是云服务提供商。实验室并不认为这是真的——他们认为自己服务的模型可能在一年内看起来非常糟糕,但这是为了建立继续训练更智能模型的商业理由。
所以这就剩下第1点:实验室利润率会提高,或者第2点:计算成本会变高。如果领先的1至2家实验室明显领先于竞争对手,那么更多可能是前者。在市场中,你的利润率取决于你比下一个最佳选择好多少。要让效应1占主导,到明年年底,利润率必须达到中高90%的水平。这对我来说听起来相当疯狂。但我认为,人工智能实验室的收入持续惊人地快速增长是可能的。
所以还剩下一个效应可以解释明年年底疯狂的1万亿美元收入世界如何可能实现:计算成本大幅上升。如我所提,这已经开始发生。当我们看实验室所需的这部分计算资源时,价格上涨更为明显——他们显然不能依赖现货实例——他们需要保证权重和客户信息的安全,并且需要足够的规模以获得良好的利用率和灵活性。为了看清这一部分计算市场的疯狂程度,可以看看谷歌和Anthropic从SpaceX租赁计算资源的价格。报道称,谷歌每月支付9亿美元租用11万块GPU:https://finance.yahoo.com/sectors/technology/articles/google-paying-spacex-over-900-151816933.html,这些GPU是GB200和GB300的混合。这大约是这些GPU每小时现货价格的2倍。而且当前现货价格本身比2月份高出了40%。
我想在这里强调一个关键结论:随着人工智能模型变得更智能,它们将能更好地利用同等计算资源赚钱。如果有一个真正的人类级软件工程师,能够在等效于 H100 的硬件上运行,按照目前软件工程师的市场薪资,这台 H100 的年租金应超过 25 万美元。这是目前现货价格的 15 倍。
当然,你可能会认为如果我们增加 1000 万名额外的软件工程师,软件工程师的边际价值会下降,那么那台 H100 不一定能产生比现在多 15 倍的收入。但我实际上并不确定这是否为真。如果我们把这个论点应用到人类身上,而不是 AI,那么这将是经典的劳动总量谬误:https://en.wikipedia.org/wiki/Lump_of_labour_fallacy。经济学家通常认为,高技能移民不会在长期内降低工资,因为专业化和创新会提升劳动价值。也许这次劳动力供给的冲击如此巨大且迅速,以至于这个一般性启发法不再适用。但如果我们相信标准经济学对劳动的看法,那么计算资源的边际价值(因此计算的边际价格)应当变得异常高。
随着顶尖模型在变现计算资源方面越来越出色,追赶变得越来越困难。如果到了 2028 年我们实现了软件工程自动化,而计算价格比现在高 15 倍,那么对于你来说,没有收入,要在计算资源上与前沿实验室竞争将变得更加困难。
如果你拥有最好的模型,那么你将能够收取比目前高得多的利润。这就是经济学中的阿尔钦–艾伦效应(Alchian–Allen effect):https://en.wikipedia.org/wiki/Alchian%E2%80%93Allen_effect,但简单来说,如果每小时 H100 的成本不是 2 美元,而是 20 美元,那么在这台计算资源上运行一个较简单的模型就是愚蠢的,因为可能会消耗更多的 token(从而浪费这个昂贵的计算资源)来得到相同结果。当然,如果你通过 API 使用,那么你是间接支付这些计算成本的,但逻辑仍然相同。如果你必须为基础计算付出如此高的费用,那你不如直接支付使用运行在相同计算量上的最优、最高效模型的费用。
很多当前流行的人工智能应用会被价格排除在外。人工智能目前相对便宜的原因,至少相比于人工劳动来说,部分原因是它无法做到许多顶尖人类可以做到的事情。到某个时候,这种情况将不再存在。因此,用GPU制作短视频“垃圾内容”也会被价格排除在外。
与此同时,这种预测确实与人们过去在稀缺性问题上的错误认识有一定的模式匹配。我想到的是西蒙–厄尔里希打赌(Simon–Ehrlich wager),保罗·厄尔里希(Paul Ehrlich)打赌,在1990年前的十年里,一篮子商品的成本会上升而非下降。这个打赌在流行经济学讨论中很有名,因为它被认为说明了厄尔里希的“马尔萨斯世界观”被证伪——他低估了市场信号和人类智慧找到更好方式节约稀缺资源的能力。(不过其他分析:https://en.wikipedia.org/wiki/Simon%E2%80%93Ehrlich_wager#Analysis 显示,如果这个赌局是在另外一个十年进行,厄尔里希可能会赢)。
我猜算力的参考类别不是西蒙–厄尔里希的一篮子商品,因为算力供应的弹性要小得多,也远不如不同金属的开采能够吸收大规模需求冲击。每年算力容量增长3倍来自将以下数字相乘得出;:1.4倍来源于摩尔定律,1.2倍来自新晶圆厂建设(至少到2030年受EUV设备供应限制:https://www.dwarkesh.com/p/dylan-patel),1.8倍来自人工智能占用其他设备的先进晶圆分配(到2027年底,这将开始碰壁,当时人工智能占N3从60%升至86%)。
我想澄清的是,未来某个时候算力将再次变得便宜。到那时,机器人将能够将硅砂海滩和铜矿变成电脑。此时,算力的价格应该更接近原材料和工具的成本。我这里讲的是当前这个阶段,人工智能算力每年只增长3倍,这还不够抵消人工智能年复一年变得更有用所带来的价格效应。
真的很喜欢这个。我有不同的观点,但与其反驳,不如为讨论增加一个维度。
软件工程是一个特殊的“光锥”领域——没有原子,没有人类尺度的循环延迟。这正是每小时25万美元的H100机制最强的地方,因为智能确实是瓶颈,因此提升它可以将整个劳动循环变现。
但是,一旦循环触及物理世界,瓶颈就会从智能转移——并且不会再与劳动价格挂钩。AI可以将药物设计压缩到计算时间,但第一个人体剂量仍然是实验:真实世界的推断,而不是AI推断。因此,计算的变现能力只会在循环完全可模拟的情况下以替代劳动的价格体现。25万美元可能是真的,作为第一个人类→AI转换率,但它是弹性的——一旦编程代理接管,它可以降到5万美元。
供应墙点无论如何都成立——这是一个独立于能力的稀缺性故事,并且可能在未来几年中更为稳健的一半。
很多有趣的内容,但关于专有前沿模型是最有效使用计算资源的说法是一个重要的子问题,需要更多审查。非前沿操作系统模型和中国模型在智能/价格方面往往更占优势(智能/价格可以作为计算的代理,也通常更小、更轻量)。
事实上,如果我们假设对更智能系统的效用存在某种其他限制或递减收益,那么运行10个GLM-5.5实例实际上可能比运行1个GPT-6实例更合理地使用计算资源。
在操作系统模型紧跟专有模型的情况下,我不会自动假设最新的模型就是最有效的。
I want to experiment with very quick blog post where I time-box writing for 2 hours. I won’t be able to nail down a lot of important sub-questions, but the alternative is just not making any progress on a lot of different topics I’m curious about.
Today I want to talk about the compute situation of the labs over the coming years.
Anthropic revenue has 10xed year over year. Anthropic likely ends the year with ~$100–150B of revenue. For this trend to continue, Anthropic would have to make $1T in revenue by the end of next year. There’s no deep reason why the trend needs to continue, and it very well might not - it’s ultimately a question about AI capabilities. But suppose it does. What would have to be true about that world?
Lab compute 3x-es :https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute year over year. For a lab to 10x revenue while continuing to only 3x compute, some combination of the following 3 things has to happen: 1. Lab margins have to increase, 2. The price of compute has to increase, 3. Labs have to spend a greater fraction of their compute on inference.
My understanding is that basically all 3 of these things have been happening: 1. Anthropic went from 40% margins in 2025 to probably >80% this year for Fable inference 1:#footnote-1 (though perhaps not on the marginal compute - see more below). 2. Spot prices for compute are up 40%+ from the February trough, and that likely understates how much more the labs have to pay (again, more below). 3. Roughly a quarter of OpenAI’s 2024 compute spend went to inference according to Epoch :https://epoch.ai/data-insights/openai-compute-spend , and it’s certainly closer to 50% if not higher now.
Labs would prefer not to do 3 (spend greater and greater shares of compute on inference). As has been said jokingly, the point of inference revenue is to convince investors to give you more money to buy more compute to train bigger models. If you’re spending most of your compute on inference, you’re kind of declaring that AI progress has stalled, because it’s not worth investing more in training, and your business is basically that of a cloud provider. The labs do not think this is true - they think they are serving models that will look extremely shitty within a year in order to build up the business case to continue training smarter models.
So that leaves 1. lab margins will increase, or 2. compute gets more expensive. It will be more of the former if the leading 1 to 2 labs are significantly ahead of the competition. In a market, your margins are set by how much better you are than the next best alternative. For effect 1 to dominate, margins would have to be in the mid-90s percent by the end of next year. That sounds quite crazy to me. But it does sound plausible to me that AI lab revenues will keep growing astonishingly fast.
So that leaves one more effect to explain how this world of crazy $1T revenue by end of next year might come to be: compute gets a lot more expensive. As I mentioned, this is already starting to happen. The price increase is even stronger when we look at the tranche of compute that the labs need - they obviously can’t rely on spot instances - they need security for their weights and customer info, and enough scale to get good utilization and flexibility. To look at how crazy the compute market is in that tranche, consider the price at which Google and Anthropic are renting compute from SpaceX. Google is reportedly paying $900 million a month for 110K GPUs :https://finance.yahoo.com/sectors/technology/articles/google-paying-spacex-over-900-151816933.html that are a blend of GB200s and GB300s. That’s roughly 2x the spot price per hour for those GPUs. And the current spot price is itself 40% higher than it was in February.
I want to emphasize the key conclusion here: as AI models become smarter, they’ll better monetize the same amount of compute. If a true human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot prices.
Of course you might expect that if we have 10 million extra software engineers, the marginal value of a software engineer would decrease and so that H100 wouldn’t necessarily be able to produce 15x more revenue than it currently does. But I actually don’t know if that’s true. If we applied this argument to people instead of AI, then this would be the classic lump of labor fallacy :https://en.wikipedia.org/wiki/Lump_of_labour_fallacy . Economists generally believe that high-skilled immigration does not decrease wages in the long run because of how specialization and innovation increase the value of labor. Maybe this labor supply shock is so big and so fast that this general heuristic no longer applies. But if we believe what standard economics says about labor, then the marginal value of compute (and thus the marginal price of compute) should become astonishingly high.
As the top models get better and better at monetizing compute, it becomes harder and harder to catch up. If by 2028 we’ve automated software engineering and the price of compute is 15x higher than it is right now, then it’s going to be much more difficult for you, with no revenue, to compete for compute against the frontier labs.
If you have the best model, then you’ll be able to charge much higher margins than you currently can. This is the Alchian–Allen effect :https://en.wikipedia.org/wiki/Alchian%E2%80%93Allen_effect in economics, but to explain very simply, if instead of $2 per H100 hour, you’re spending $20, then it would be stupid to run a dumber model on that compute, which may take more tokens (thus wasting this expensive compute) to get you the same result. Of course if you’re using an API then you’re paying for that compute indirectly, but the logic is still the same. If you’re gonna have to pay so much for the underlying compute in the first place, you might as well pay to use the very best, most efficient model that runs on that amount of compute.
A lot of current popular applications of AI get priced out. The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.
At the same time, this kind of prediction does pattern match onto the ways people have been wrong about scarcity in the past. I’m thinking of the Simon–Ehrlich wager, where Paul Ehrlich bet that the cost of a basket of commodities would increase rather than decrease in the decade leading up to 1990. This wager is famous in popular economics discussion, because it’s supposed to illustrate how Ehrlich’s Malthusian worldview was falsified — he underestimated the way in which market signals and human ingenuity would find ways to better economize scarce inputs. (Though other analysis :https://en.wikipedia.org/wiki/Simon%E2%80%93Ehrlich_wager#Analysis shows that if this bet had been made in a different decade, Ehrlich would have won).
I’m guessing the Simon-Ehrlich basket of commodities is not the correct reference class for compute, because compute supply is much less elastic, and much less capable of absorbing large demand shocks, than the extraction of different metals is. That 3x in compute capacity per year comes from multiplying the following numbers together ; : 1.4x from Moore’s Law, 1.2x from new fab construction ( bottlenecked through at least 2030 by EUV tool supply :https://www.dwarkesh.com/p/dylan-patel ), 1.8x from AI taking leading edge wafer allocation from other devices (which will start to hit a wall by end of 2027, when AI will go from 60% of N3 to 86%).
I want to clarify that at some point in the future compute becomes cheap again. At some point robots will be able to turn shores of silica sand and mines of copper into computers. At which point the price of compute should fall closer to the cost of raw inputs and tools. I’m just talking about this current regime where AI compute merely 3x-es year over year, which is not fast enough to offset the price effect of how much more useful AI is becoming year over year.
Really enjoyed this. I have a different view, but rather than pushing back, I'd like to add a dimension to the discussion.
Software engineering is a special "light cone" domain — no atoms, no human-scale latency in the loop. That's exactly where the $250k/H100 mechanism is strongest, because intelligence really is the bottleneck, so improving it monetizes the whole labor loop.
But once the loop touches the physical world, the bottleneck moves off intelligence — and stops tracking labor price. AI can collapse drug design toward compute time, but the first dose in a human is still a trial: real-world inference, not AI inference. So compute only monetizes at displaced-labor rates where the loop is fully simulable. The $250k might be real as the first human→AI conversion rate, but it's elastic — once the coding agent takes over, it can fall to $50k.
The supply-wall point stands either way — a scarcity story independent of capability, and probably the more robust half for the next few years.
Lots of interesting stuff but the claim that the proprietary frontier models are the most efficient use of compute is an important sub-question that requires more scrutiny. The non-frontier OS and Chinese models tend to win on intelligence/price (a proxy for compute and also tend to be smaller and more lightweight).
In fact, if we assume there is some other constraint or diminishing returns on the utility of more intelligence systems it might actually be the case that running 10 instances of GLM-5.5 is a better use of computational resources than 1 instance of GPT-6.
I wouldn't automatically assume that the newest models are the most efficient in a world where the OS models are so hot on the heels of the proprietary ones.