如果该芯片兑现其承诺,它仍可能成为重大的竞争优势。在 AI 业务中,公司对推理成本的优化程度正日益决定其利润率。谷歌可以使用 Frozen v2 以更低的价格运行强大的模型,并从 OpenAI 和 Anthropic 手中争取市场份额。
保持对 AI 的关注。清晰、有用,无废话。
关注 The Decoder 获取 AI 新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
Google is building a new server chip internally called "Frozen v2" that embeds the Gemini AI model's architecture directly into silicon.
The chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips, according to sources cited by The Information:https://www.theinformation.com/articles/google-plans-new-frozen-chip-run-ai-models-efficiently. Google plans to deploy it starting in 2028 and sees Frozen v2 as a test run for specialized chips, with a smaller production volume than its TPU line.
Unlike Google's TPUs, which work with many models, Frozen v2 has parts of Gemini's model structure built right into the hardware. The name follows the same logic as "freezing" parameters in AI models, where you lock values so they stop changing. With Frozen v2, a portion of the model gets permanently frozen into the chip itself, which cuts down on compute steps and speeds up responses. Ad
The original idea reportedly came from Jeff Dean, Google Deepmind's chief scientist. His first Frozen design called for embedding the model weights directly into the chip. Weights are the specific settings that determine how an AI model responds to queries. Google scrapped that approach because the chip would have only worked with a single Gemini version and would have become outdated too quickly. Ad DEC_D_Incontent-1
Frozen v2 takes a more flexible path by embedding the model architecture instead of weights, meaning the underlying blueprint rather than the tuned parameters. New weights can still be loaded onto the chip. How much of the architecture will actually be hardcoded hasn't been decided yet, according to The Information.
Because the chip only works as long as Google sticks with the same model architecture, it probably won't become a product for outside customers. Google already leases its TPUs to Meta:https://the-decoder.com/meta-signs-multi-billion-dollar-deal-to-rent-googles-tpus-in-a-direct-challenge-to-nvidias-ai-chip-dominance/, offers them to external cloud customers:https://the-decoder.com/google-unveils-8th-gen-tpus-agent-platform-and-workspace-ai-layer-at-cloud-next-26/, and positions them through its "TPU@Premises" program as an alternative to Nvidia:https://the-decoder.com/the-mere-existence-of-google-tpus-reportedly-saved-openai-30-on-nvidia-chips/ with an internal goal of capturing ten percent of Nvidia's annual revenue:https://the-decoder.com/google-cloud-aims-to-capture-ten-percent-of-nvidias-annual-revenue-with-tpus/. Frozen v2, by contrast, is meant to ease Google's internal crunch on AI compute capacity:https://www.theinformation.com/articles/inside-balancing-act-googles-compute-crunch. Ad
If the chip delivers on its promise, it could still become a major competitive edge. In the AI business, how well companies optimize inference costs increasingly determines their margins. Google could use Frozen v2 to run powerful models at lower prices and take market share from OpenAI and Anthropic.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.