Moonshot AI 发布 Kimi K3,早期评估显示其性能接近 Anthropic Opus 4.8,但仍落后于 Fable 5 和 GPT-5.6 Sol。
该模型每任务平均成本 0.94 美元,低于 Opus 4.8 的 1.80 美元。 Google Deepmind 研究员称其"好得离谱",认为仅靠知识蒸馏无法解释其表现。
Moonshot AI 发布了 Kimi K3,这款模型据称已接近与西方顶级模型匹敌的水平。这一发布再次引发人们对美国出口管制是否真正奏效的怀疑。即便是 OpenAI 的战略师也感到震惊。
Somaia 认为,整个西方共识——从出口管制到超级大厂上百亿美元的投资竞赛:https://the-decoder.com/big-techs-ai-spending-balloons-to-725-billion-this-year/ 以及“计算护城河”投资理论——都建立在一个假设之上:计算能力决定能力。
但稀缺性推动了创新。Somaia 表示,Moonshot AI 自研的 Mooncake 堆栈:https://kvcache-ai.github.io/Mooncake/ 用于 AI 训练,正是因为这家初创公司 GPU 不够才开发的:https://x.com/AnikaSomaia/status/2077892561386299664。“一个有品味的小型实验室可以压缩构建前沿模型所需的计算,即使它负担不起服务整个模型。”
硬件分析公司 SemiAnalysis 的创始人 Dylan Patel 也同意这一观点。“他们凭借一支极具天赋的小团队、在 RL、架构、数据方面的强大研究,弥补了大量计算上的不足,”他写道:https://x.com/dylan522p/status/2078084636719435959。但 Patel 也指出,中国公司可以轻松在中国境外租用 GPU,这使得部分出口限制变得无效。
西方 AI 实验室常指责中国公司通过蒸馏进行某种形式的数据窃取:https://the-decoder.com/google-and-openai-complain-about-distillation-attacks-that-clone-their-ai-models-on-the-cheap/,即较小的 AI 模型从较大的模型输出中学习,本质上是搭便车,这威胁了西方 AI 实验室的商业模式。直到现在,蒸馏一直是解释中国实验室如何在计算资源较少的情况下仍能保持竞争力的主要原因。
对于Kimi K3来说,这个解释显然站不住脚。MIT和谷歌Deepmind的人工智能研究员Michiel Bakker写道:“这些结果似乎无法仅靠蒸馏来解释。”httpshttps://x.com/bakkermichiel/status/2077857476574052730,他称该模型“非常出色。”据彭博社报道,谷歌自家旗舰机型Gemini 3.5 Pro已推迟数月:https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals,因其未能达到性能目标,尤其是在主要应用场景编程方面。该公司的人工智能战略再次受到批评,谷歌在人工智能搜索领域也面临监管阻力,尤其是来自德国的挑战:https://the-decoder.com/germany-puts-googles-ai-overviews-and-perplexity-under-media-law-in-first-of-its-kind-ruling/。
Dean W. Ball:https://x.com/deanwball/status/2078133895766114412,OpenAI战略未来主管、前政府顾问称Kimi为“非常好的模型”,在基于代理的编码会话中,它与“2026年第一季度最好的公开模型”相匹配。但他也指出,这款模型“非常渴望代币化”,因此“我并不明显这个型号运行成本真的那么低。”
他说得没错。根据Artificial Analysis:https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/#artificial-analysis-confirms-k3s-strong-performance-but-flags-higher-hallucination-rate,Kimi K3每项任务平均成本为0.94美元。这接近GPT 5.6 Sol的1.04美元,但价格大约是Opus 4.8的一半,1.80美元。它仍然比西方顶级车型便宜,但差距比之前的版本缩小,价格也比早期的空重中国车型要高得多。
不过,鲍尔表示他对中国政府允许如此强大的模型开源发布感到惊讶。他将其中75%归因于战略性盲目,称中共在评估人工智能风险的方式上“非常像Yann LeCun”,并未看到任何生存威胁:https://the-decoder.com/yann-lecun-warns-ai-labs-like-openai-and-anthropic-face-a-big-bubble-explosion/。
剩下的问题归结为客户侧推理的计算能力不足,这使得开放权重策略成为美国出口管制的意外副产品:https://the-decoder.com/chinas-semiconductor-independence-push-is-turning-us-export-controls-into-a-domestic-boom/。Ball称,这些公司也知道,几乎没有人会为低于前沿水平的中国模型付费。
Ball认为,开放权重模型“本质上是减速型”的,因为它们会减缓进一步的人工智能投资。在一个由这类模型主导的世界中,可能的结果之一是“全面人工智能共产主义”,由国家作为数字基础设施提供人工智能作为公共产品。据Ball称,这正是中国所提出的方案,他称这种情景为“反乌托邦地狱景象”。
当然,OpenAI的一位策略师如此严厉地批评开放权重模型并非没有私利。他的公司依赖封闭的商业模式,并面临来自Moonshot AI和Deepseek等供应商日益增加的价格压力:https://the-decoder.com/ai-startup-lindy-ditched-claude-entirely-for-deepseek-saving-millions-as-cost-pressure-mounts-on-anthropic/。
Ball预测特朗普政府将围绕使用中国开放权重模型制造监管风险。他认为没有必要禁止开源,将其称为“AI政策讨论中最愚蠢的主题之一”。当局只需通过“软法”制造足够的不确定性,例如让美联储对中国AI模型可能存在的后门发出警告。其理由甚至无需有充分依据。
目标是寻找一种中间立场,制造足够的风险以阻止受监管公司使用中国模型,同时又不会吓到超级计算中心,以致初创公司迁移到信誉较差的供应商。Ball预计政府会推出某种版本的这一策略:https://the-decoder.com/anthropic-frames-ai-competition-with-china-as-a-now-or-never-moment-for-washington/。
Kimi的进展并不一定意味着需要更少的计算能力,但如果真是这样,美国科技公司的大规模基础设施建设就显得不必要,可能会引发股市崩盘。不过,更可能的情况恰恰相反。杰文斯悖论:https://en.wikipedia.org/wiki/Jevons_paradox 表明,更高效的模型会推动更多AI部署,这实际上可能进一步增加对计算能力的需求。
根据SemiAnalysis:https://x.com/SemiAnalysis_/status/2077966560447074689,Kimi K3拥有2.8万亿参数,如此庞大以至于单个Nvidia DGX B200无法容纳,即使进行了FP4量化。它需要更强大的系统,如GB300 NVL72或B300,每个GPU具有288 GB内存。
再次强调,与Deepseek的相似之处不容忽视:https://the-decoder.com/deepseek-v3-emerges-as-chinas-most-powerful-open-source-language-model-to-date/。当时,怀疑者预测会出现计算冗余,并短暂震动了市场。相反,随着推理模型逐渐受到欢迎,对计算能力的需求增加,讽刺的是,这在一定程度上是由Deepseek自己的模型推动的。或者如谷歌DeepMind CEO Demis Hassabis所说:“世界上没有人知道接下来会发生什么。”:https://the-decoder.com/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/
保持对AI的关注。清晰、有用,无废话。
关注The Decoder,获取AI新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
Moonshot AI has released Kimi K3, a model reportedly close to matching top Western models. The launch raises fresh doubts about whether U.S. export controls are actually working. Even an OpenAI strategist is impressed.
Somaia argues that the entire Western consensus, from export controls to the hyperscalers' hundreds-of-billions investment race:https://the-decoder.com/big-techs-ai-spending-balloons-to-725-billion-this-year/ to the "Compute Moat" investment thesis, rests on a single assumption: that computing power determines capability.
But scarcity has forced innovation. Moonshot AI's in-house Mooncake stack:https://kvcache-ai.github.io/Mooncake/ for AI training was built precisely because the startup didn't have enough GPUs, Somaia says:https://x.com/AnikaSomaia/status/2077892561386299664. "A small lab with taste can compress the compute needed to make a frontier model, even if it can't afford to serve one."
Dylan Patel, founder of hardware analysis firm SemiAnalysis, agrees. "What they did with an extremely talented small team, strong research in RL, arch, data helps make up for lot of the compute deficit," he writes:https://x.com/dylan522p/status/2078084636719435959. But Patel also points out that Chinese companies can easily rent GPUs outside of China, which makes a portion of the export restrictions pointless.
Western AI labs often accuse Chinese companies of a form of data theft through distillation:https://the-decoder.com/google-and-openai-complain-about-distillation-attacks-that-clone-their-ai-models-on-the-cheap/, where a smaller AI model learns from the output of a larger one and essentially free-rides, threatening Western AI labs' business models. Until now, distillation has been the go-to explanation for how Chinese labs stay competitive despite having less compute.
For Kimi K3, that explanation apparently doesn't hold up. "These results seem impossible to explain through distillation alone," writes Michiel Bakker:https://x.com/bakkermichiel/status/2077857476574052730, an AI researcher at MIT and Google Deepmind, calling the model "insanely good." Google's own flagship model, Gemini 3.5 Pro, meanwhile, has been delayed for months according to Bloomberg:https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals because it isn't hitting performance targets, especially in coding, its main use case. The company's AI strategy is drawing criticism again, and Google is also facing regulatory headwinds in AI search, particularly from Germany:https://the-decoder.com/germany-puts-googles-ai-overviews-and-perplexity-under-media-law-in-first-of-its-kind-ruling/.
Dean W. Ball:https://x.com/deanwball/status/2078133895766114412, Head of Strategic Futures at OpenAI and a former government advisor, calls Kimi a "very good model" that in agent-based coding sessions matches "the best public models from Q1 2026." But he also notes that it seemed "very token hungry," making it "not obvious to me that this model is actually that cheap to run."
He's not wrong. According to Artificial Analysis:https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/#artificial-analysis-confirms-k3s-strong-performance-but-flags-higher-hallucination-rate, Kimi K3 costs an average of $0.94 per task. That's close to GPT 5.6 Sol at $1.04 but roughly half the cost of Opus 4.8 at $1.80. It's still cheaper than the top Western models, but the gap has narrowed compared to the previous version, and it's much pricier than earlier open-weight Chinese models.
Still, Ball says he's surprised the Chinese government allows such powerful models to be released as open-source. He attributes 75 percent of it to strategic blindness, saying the CCP is "very Yann LeCun-y" in how it assesses AI risks and doesn't see any existential threats:https://the-decoder.com/yann-lecun-warns-ai-labs-like-openai-and-anthropic-face-a-big-bubble-explosion/.
The rest comes down to a lack of computing capacity for client-side inference, which makes the open-weight strategy an unintended byproduct of U.S. export controls:https://the-decoder.com/chinas-semiconductor-independence-push-is-turning-us-export-controls-into-a-domestic-boom/. The companies also know that hardly anyone would pay for Chinese models below the frontier, Ball claims.
Open-weight models are "inherently decelerationist," Ball argues, because they slow down further AI investment. One possible outcome of a world dominated by them would be "full AI communism," with AI as a public good provided by the state as digital infrastructure. That's what China is proposing, according to Ball, who calls this scenario a "dystopian hellscape."
That an OpenAI strategist is criticizing open-weight models this sharply is, of course, not without self-interest. His company relies on a closed business model and faces growing price pressure:https://the-decoder.com/ai-startup-lindy-ditched-claude-entirely-for-deepseek-saving-millions-as-cost-pressure-mounts-on-anthropic/ from providers like Moonshot AI and Deepseek.
Ball predicts the Trump administration will create regulatory risk around using Chinese open-weight models. There's no need to ban open source, he argues, calling it "one of the dumber motifs of AI policy discussion." Authorities would only need to create enough uncertainty through "soft law," like having the Federal Reserve issue warnings about potential backdoors in Chinese AI models. The rationale wouldn't even need to be well-founded.
The goal is a middle ground with enough risk to deter regulated companies from using Chinese models, without spooking the hyperscalers so badly that startups migrate to less reputable providers. Ball expects the government to roll out some version of this strategy:https://the-decoder.com/anthropic-frames-ai-competition-with-china-as-a-now-or-never-moment-for-washington/.
Kimi's progress doesn't necessarily mean less computing power is needed, but if it did, U.S. tech companies' massive infrastructure buildouts would look unnecessary, likely triggering a stock market crash. The opposite is more likely, though. The Jevons paradox:https://en.wikipedia.org/wiki/Jevons_paradox suggests that more efficient models lead to more AI being deployed, which could actually drive even more demand for computing power.
According to SemiAnalysis:https://x.com/SemiAnalysis_/status/2077966560447074689, Kimi K3 has 2.8 trillion parameters and is so large it doesn't fit on a single Nvidia DGX B200, even with FP4 quantization. It needs more powerful systems like the GB300 NVL72 or B300, each with 288 GB of memory per GPU.
Again, the parallels to Deepseek are hard to miss:https://the-decoder.com/deepseek-v3-emerges-as-chinas-most-powerful-open-source-language-model-to-date/. Back then, skeptics predicted a compute surplus and briefly rattled the markets. Instead, demand for computing power climbed as reasoning models gained traction, ironically driven in part by Deepseek's own models. Or as Google Deepmind CEO Demis Hassabis puts it, "Nobody in the world knows what happens next.":https://the-decoder.com/deepmind-ceo-hassabis-says-nobody-in-the-world-knows-what-happens-next-so-cautious-optimism-means-building-guardrails-now/
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/