Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
The Decoder:AI News(RSS)Aioga 编辑团队2026-07-16T19:49:39.000Z热度
Kimi is launching K3, a multimodal open-weight model with 2.8 trillion parameters and on...
AI资讯The Decoder:AI News(RSS)
今日 AI 情报摘要
Kimi is launching K3, a multimodal open-weight model with 2.8 trillion parameters and one million
tokens of context. In the company's own benchmarks, it comes close to Claude Fable 5 and GPT 5.6 Sol while beating Opus 4.8 and GLM 5.2, in some cases by a wide margin. The model is also significantly pricier than its predecessor. Full weights are scheduled for release by July 27.
中文正文 · AI 翻译
Kimi 正在推出 K3,这是一款参数量为 2.8 万亿、具有一百万令牌上下文窗口的多模态模型。在公司自己的基准测试中,它的表现与领先的专有模型相当。
在 Kimi 自己的基准测试中:https://www.kimi.com/blog/kimi-k3,K3 仍然落后于顶级专有模型 Claude Fable 5 和 GPT 5.6 Sol,但击败了所有其他测试系统,包括 Claude Opus 系列模型和中国竞争对手 GLM-5.2。根据公司介绍,所有结果均来自 Kimi,并在最大或高强度思考下获得。
根据 Kimi 的说法,该模型的主要使用场景是长期软件开发,且几乎无需人工监督。K3 旨在分析大型代码库、协调终端工具,并在多个工作步骤中保持对任务的专注。Ad
该模型将编程与视觉反馈相结合:它会检查屏幕截图、修改代码,然后检查可见输出。Kimi 将这种闭环系统称为“Vision in the Loop”,并将其定位为游戏开发、UI 设计和 CAD 的基础。
作为演示,Kimi 展示了一个程序生成的 3D 开放世界游戏:https://horseback-open-world.ok.kimi.link/,据称 K3 完全在浏览器中使用 Three.js、WebGPU 和 GPU Compute 构建,以及一个交互式黑洞可视化:https://blackhole-visualizer.ok.kimi.link/。对于开放世界演示,K3 程序生成了环境,并使用外部工具创建了 3D 骑手和马模型。其他演示:https://www.kimi.com/blog/kimi-k3 包括长征十号火箭发射与返回的模拟,以及一个 Game Boy Advance 模拟器。
K3 已可通过 Kimi.com:http://kimi.com/,iOS、Android 和 HarmonyOS 移动应用,Kimi Work 桌面客户端:https://www.kimi.com/zh-cn/products/kimi-work(版本 3.1.0 及以上),以及 Kimi Code:https://www.kimi.com/code 获取。在 OpenRouter 上,该模型的标识为 "moonshotai/kimi-k3",但目前只通过 Moonshot 本身提供服务。预计开源权重将在七月底前发布。
保持对 AI 的了解,内容清晰、有用,无冗余。
关注 The Decoder 获取 AI 新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
Kimi is launching K3, a multimodal model with 2.8 trillion parameters and a context window of one million tokens. In the company's own benchmarks, it performs on par with leading proprietary models.
According to Kimi, the new flagship model K3 has 2.8 trillion total parameters, processes images and video natively, and supports a context window of one million tokens. Kimi calls K3 the first open model in the roughly 3 trillion parameter range. Full model weights are scheduled for release by July 27. The model targets long-running programming tasks, knowledge work, and complex reasoning.
In Kimi's own benchmarks:https://www.kimi.com/blog/kimi-k3, K3 still trails the top proprietary models Claude Fable 5 and GPT 5.6 Sol but beats every other system tested, including the Claude Opus models and Chinese rival GLM-5.2. All results come from Kimi and were achieved at maximum or high thinking intensity, according to the company. Ad
Across all 35 tests, K3 took first place about seven times and landed second or third in most of the rest. Fable 5 won the most individual tests. In nearly every benchmark, K3 beat Opus 4.8, GPT 5.5, and GLM 5.2 by a wide margin. Depending on the benchmark, one of three agent systems was used: KimiCode, Claude Code, or Codex. That means the results weren't all collected under identical conditions. Ad DEC_D_Incontent-1
Independent testing lab Artificial Analysis:https://x.com/ArtificialAnlys/status/2077832874183860404 has published its first evaluation of Kimi K3. The model scores 57 on the Artificial Analysis Intelligence Index, putting it on par with Opus 4.8 and GPT-5.5 but still behind Fable 5 and GPT-5.6 Sol. That largely lines up with Kimi's own claims.
On agentic tasks, K3 reaches an Elo rating of 1,668 on GDPval v2, a big jump from K2.6's 1,190. It beats GLM-5.2 (1,514), GPT-5.5 (1,494), and Claude Opus 4.8 (1,600), though it still falls short of Claude Fable 5 (1,760). K3 also takes the top spot on AutomationBench-AA, Artificial Analysis's version of Zapier's agentic SaaS workflow evaluation, with a score of 53 percent. Ad
On AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reaches an overall Elo of 1,547, up 732 points from K2.6. Only Claude Fable 5 scores higher. Artificial Analysis calls K3 well-rounded, with rubric scoring and analytical quality close to Fable 5's level. GPT-5.6 Sol still leads on presentation quality, though.
K3's accuracy rate improved from 33 percent to 46 percent on the AA-Omniscience Index, pushing the overall score from +6 to +18. But its hallucination rate climbed from 39 percent to 51 percent, meaning K3 fabricates more answers even as it gets more questions right. Ad DEC_D_Incontent-2
According to Kimi, the model's primary use case is long-running software development with minimal human oversight. K3 is built to analyze large codebases, coordinate terminal tools, and stay focused on a task across many work steps. Ad
The model pairs programming with visual feedback: it examines screen captures, modifies code, then checks the visible output. Kimi calls this closed-loop system "Vision in the Loop" and positions it as a foundation for game development, UI design, and CAD.
As demos, Kimi shows off a procedurally generated 3D open-world game:https://horseback-open-world.ok.kimi.link/ that K3 reportedly built entirely in the browser using Three.js, WebGPU, and GPU Compute, along with an interactive black hole visualization:https://blackhole-visualizer.ok.kimi.link/. For the open-world demo, K3 procedurally generated the environment and used an external tool to create the 3D rider and horse models. Other demos:https://www.kimi.com/blog/kimi-k3 include a simulation of the Long March 10 rocket launch and return, plus a Game Boy Advance emulator.
K3 uses a mixture-of-experts architecture that activates only 16 of 896 experts at a time. It's paired with a new attention architecture called Kimi Delta Attention:https://arxiv.org/abs/2510.26692, which Kimi says enables up to 6.3x faster decoding for million-token contexts. "Attention residuals":https://arxiv.org/abs/2603.15031 reportedly boost training efficiency by about 25 percent while adding less than 2 percent in extra compute overhead.
According to the Kimi API docs:https://platform.kimi.com/, one million input tokens cost $0.30 with a cache hit and $3.00 without. One million output tokens, including reasoning, cost $15.00. These prices apply regardless of context length. Caching happens automatically, which makes unmodified long prefixes especially useful for agents and large codebases.
That puts K3 well above the price level of its predecessor K2.6, which officially costs $0.16 per million tokens with a cache hit, $0.95 without, and $4.00 for output. Chinese providers aren't offering their frontier models at rock-bottom prices anymore either.
Still, K3 is much cheaper than the top Western models and sits more in the upper midrange. Anthropic's new Sonnet 5:https://the-decoder.com/claude-sonnet-5-continues-anthropics-pattern-of-hiding-price-increases-behind-unchanged-token-rates/, for example, also costs $3 per million input tokens and $15 for output but delivers lower performance.
According to Artificial Analysis:https://x.com/ArtificialAnlys/status/2077832885021835289, K3 averages $0.94 per task on the Intelligence Index, close to GPT-5.6 Sol at $1.04 and about half the price of Opus 4.8 at $1.80. It's well above open-weight peers like GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04), though.
K3 also uses fewer tokens than its predecessor. It needed about 132 million output tokens to complete all nine evaluations, down from roughly 166 million for K2.6, a 21 percent reduction while scoring 13 points higher. Because of the much higher token prices, K3 will likely still cost more per task than K2.6 in most cases.
K3 is already available through Kimi.com:http://kimi.com/, the mobile app for iOS, Android, and HarmonyOS, the Kimi Work desktop:https://www.kimi.com/zh-cn/products/kimi-work client (version 3.1.0 and later), and Kimi Code:https://www.kimi.com/code. On OpenRouter, the model is listed under the identifier "moonshotai/kimi-k3," though it's currently served there only through Moonshot itself. The open weights are expected by the end of July.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI。 Aioga 将其归入「AI资讯」方向,重点关注它对真实使用和行业竞争的影响。