新架构声称每单位算力的智能度提升 2.5 倍。 该模型在基准测试中接近 Fable 5 和 GPT-5.6 Sol 等西方前沿模型,但英国网络研究所的独立测试显示其网络能力和数学技能仍落后于前沿模型。
中国人工智能公司Moonshot AI已经发布了Kimi K3的模型权重和技术报告。除了在Hugging Face上提供的模型权重:https://huggingface.co/moonshotai/Kimi-K3,该公司还在开源其部分基础设施,包括高性能注意力内核、MoE通信库,以及用于大规模运行AI代理的工具。Moonshot AI声称,该新架构每单位计算提供2.5倍的智能。技术报告:https://github.com/MoonshotAI/Kimi-K3/blob/master/k3_tech_report.pdf 可在GitHub上获取。
自2026年7月中旬首次宣布以来:https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/,Kimi K3引起了轰动:https://the-decoder.com/just-like-deepseek-chinas-kimi-k3-is-forcing-western-ai-labs-to-question-their-compute-advantage/,在流行基准测试中,其得分接近西方前沿模型如Fable 5和GPT-5.6 Sol,但成本略低,并且现在开放了权重。然而,英国网络研究所进行的独立测试:https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/发现,该模型在网络能力方面远远落后于前沿模型。其数学能力也是如此:https://the-decoder.com/moonshots-kimi-k3-outperforms-fable-5-in-frontend-code-but-lags-far-behind-in-complex-math/。
这两个差距都可能表明Kimi K3依赖于蒸馏技术,这是一种由小型模型从更强大的模型输出中学习的技术。中国模型经常面临这种指控:https://the-decoder.com/google-and-openai-complain-about-distillation-attacks-that-clone-their-ai-models-on-the-cheap/。与此同时,美国的开源权重支持者越来越将蒸馏视为一种合法技术:https://the-decoder.com/nadella-calls-out-ai-labs-like-openai-and-anthropic-for-banning-distillation-while-training-on-everyone-elses-data/。Ad DEC_D_Incontent-1 Ad
保持对AI的关注。清晰、有用,无废话。
关注The Decoder获取AI新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
Chinese AI company Moonshot AI has released the model weights and technical report for Kimi K3. Along with the model weights on Hugging Face:https://huggingface.co/moonshotai/Kimi-K3, the company is open-sourcing parts of its infrastructure, including high-performance attention kernels, an MoE communication library, and tools for running AI agents at scale. Moonshot AI claims the new architecture delivers 2.5 times more intelligence per unit of compute. The technical report:https://github.com/MoonshotAI/Kimi-K3/blob/master/k3_tech_report.pdf is available on GitHub.
Since its initial announcement:https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/ in mid-July 2026, Kimi K3 has caused a stir:https://the-decoder.com/just-like-deepseek-chinas-kimi-k3-is-forcing-western-ai-labs-to-question-their-compute-advantage/ by scoring close to Western frontier models such as Fable 5 and GPT-5.6 Sol on popular benchmarks, but at a slightly lower cost and now with open weights. However, an independent test by the UK's Cyber Institute:https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/ found that the model's cyber capabilities lag far behind those of frontier models. The same is true of its math skills:https://the-decoder.com/moonshots-kimi-k3-outperforms-fable-5-in-frontend-code-but-lags-far-behind-in-complex-math/.
Both gaps could suggest that Kimi K3 relies on distillation, a technique in which a smaller model learns from the outputs of a more capable one. Chinese models often face this accusation:https://the-decoder.com/google-and-openai-complain-about-distillation-attacks-that-clone-their-ai-models-on-the-cheap/. At the same time, American open-weight advocates increasingly view distillation as a legitimate technique:https://the-decoder.com/nadella-calls-out-ai-labs-like-openai-and-anthropic-for-banning-distillation-while-training-on-everyone-elses-data/. Ad DEC_D_Incontent-1 Ad
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/