According to Wccftech, Google’s Senior Director of Supply Chain Infrastructure Nikil Cherian
revealed at a summit forum that the AI industry has shifted from being compute-constrained to memory-constrained, with high-performance memory accounting for about 75% of the material cost of a single AI server. To address this, Google dismantles decommissioned servers to recover DDR4 memory modules and designs dedicated adapter cards to use DDR4 in some scenarios as a substitute for the originally planned DDR5 in new AI servers. Google is also optimizing underlying libraries, model architectures, and KV cache compression algorithms to reduce memory usage.
IT之家:https://www.ithome.com/ 9 月 2 日消息,据 Wccftech 报道,DRAM 领域的供应持续吃紧,迫使 AI 服务商不得不各出奇招,只为抢购极度紧缺的内存资源。
谷歌供应链基础设施高级总监尼基尔 · 切里安(Nikhil Cherian)近日在出席峰会论坛时透露, 人工智能产业已迅速由“算力受限”转变为“内存受限” 。在单台典型的 AI 服务器中,高性能内存占其物料清单的成本比例已高达约 75%。
切里安表示,为了“突破内存瓶颈, 谷歌正同时从软件和硬件两端协同发力 ”。为此,谷歌甚至拆解了退役服务器以回收尚具利用价值的元器件,从而构建出一套内部循环利用的供应链体系。
在某些应用场景下,谷歌似乎已被迫妥协 —— 不得不重新采用规格较旧的 DDR4 内存模组,以替代其新型 AI 服务器原本规划的 DDR5 模组。
事实上,谷歌设计了专用的硬件转接卡,使得 DDR4 等上一代内存方案能够接入其新一代 AI 服务器。切里安坦言,谷歌正是特意将这批退役服务器调回机房, 目的就是拆解并回收其中的 DDR4 内存模组 。
IT之家注:为模型推理任务进行针对性优化的 TPU8i 芯片采用专用分层存储体系,在整机层面依托高速 DDR5 内存体系结构,专门负责处理数据预处理等任务。此外,每颗 TPU 8i 芯片均搭载 288 GB 的 HBM3e 高带宽内存。
与此同时,这家科技巨头还在持续优化底层函数库、模型架构以及 KV 缓存压缩算法,竭力降低单位算力下的内存开销。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.