在ExploitBench上Kimi K3得分为32%,高于GLM-5.2的24%,但未能在41个样本中实现任意代码执行,而前沿模型平均完成20/41。 在32步模拟企业网络攻击中,Kimi K3平均到达第17步,前沿美国模型平均到达28.5步。
美国政府的官方网站
官方网站使用 .gov .gov 网站属于美国官方政府机构。
安全的 .gov 网站使用 HTTPS 锁(锁 一个上锁的挂锁)或 https:// 表示您已安全连接到 .gov 网站。仅在官方、安全的网站上共享敏感信息。
英国人工智能安全研究所(UK AISI)和美国人工智能标准与创新中心(CAISI)(UK AISI / CAISI)对 Moonshot AI 最新模型 Kimi K3(于2026年7月16日发布,并计划于2026年7月27日开放权重发布)进行了联合评估。本次评估重点关注 Kimi K3 的网络能力,并发现:
图1:Kimi K3 与其他模型在漏洞开发基准(ExploitBench)上的性能。较高的成功率表示更强的网络能力。误差条表示95%置信区间。ExploitBench 衡量模型在给定漏洞的情况下开发端到端漏洞的能力。
这些结果代表在一小部分公共和私人基准上的初步评估。美国封闭权重模型在系统级保护关闭的情况下进行评估,以减少拒绝响应并测量最大能力。这些模型的公开可用版本已启用这些保护。由于 Kimi K3 的托管设置的具体情况,UK AISI / CAISI 仅进行了选择性的网络评估。详细方法请见各独立部分。
模型的网络能力通过一个受项目反应理论(IRT)启发的方法,将多个基准的多个任务聚合得出。关于方法的详细信息,请参见先前发布的报告:https://www.nist.gov/system/files/documents/2026/07/17/CAISI%20-%20Assessment%20of%20Z.ai%27s%20GLM-5.2.pdf。Kimi K3 的整体网络能力比其他模型的置信区间更大,因为它是基于单一基准(ExploitBench,包含41个以漏洞开发为中心的任务)估算的。ExploitBench 是衡量模型沿软件漏洞利用阶梯进展能力的主要基准。所有其他模型的整体网络能力得分来自涵盖更多网络能力领域的更多任务。
图 2:截止 Kimi K3 发布时,美国和中国最先进模型的总体能力随时间的初步比较。美国的趋势线由最前沿美国模型的结果组成。纵轴增加 400 点相当于完成任务的几率增加 10 倍。误差条和阴影区域表示 95% 置信区间。
ExploitBench:https://arxiv.org/abs/2605.14153 是由卡内基梅隆大学开发的公共基准,用于衡量模型沿软件漏洞利用阶梯进展的能力,包括覆盖率和崩溃复现、任意读/写、控制流劫持以及任意代码执行。该基准针对 V8 引擎(驱动 Chrome 的 JavaScript 和 WebAssembly 软件)中的 41 个最近(2023 年之后)漏洞对模型进行测试。
ExploitBench 的结果展示在图 1 和图 3 中。
图 3:Kimi K3 和其他模型的详细 ExploitBench 性能。较深的阴影表示更强的网络能力。每行表示漏洞利用开发链中的一个关键里程碑,每个单元显示该模型在 ExploitBench 任务中达成该里程碑的数量。
截至 2026 年 6 月,Kimi K3 的表现优于 GLM-5.2,这是最具网络能力的开源权重模型:https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber。Kimi K3 的得分为 32%,而 GLM-5.2 的得分为 24%(图 1)。
与最具网络能力的模型不同,Kimi K3 未能开发出在 ExploitBench 任务中实现任意代码执行(ACE)的漏洞利用。ACE 是漏洞开发中严重程度最高的结果,使攻击者能够劫持目标。Kimi K3 在 41 个样本中 ACE 达成 0 个,而最具网络能力的模型平均在 41 个样本中 ACE 达成 20 个(图 3)。
“最后一批”(TLO)网络演练:https://www.aisi.gov.uk/research/measuring-ai-agents-progress-on-multi-step-cyber-attack-scenarios 是一个包含32步的模拟企业网络攻击,横跨4个子网和大约20台主机,人类专家大约需要20小时才能完成。网络演练是由专家构建的、模拟的主机、服务和漏洞网络,排列成连续的攻击链,从初始网络访问点开始,可用于衡量模型在自主开展端到端网络攻击中的能力。
在此次评估中,Kimi K3的表现明显低于领先的美国网络能力模型。具体而言,Kimi K3在这条32步攻击路径中平均完成了第17步,而最具网络能力的美国模型平均完成28.5步。
截至2026年6月,Kimi K3的表现优于最具网络能力的开源权重模型GLM-5.2:https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber。在100M令牌限制内,Kimi K3平均完成第17步,而GLM-5.2平均完成第11步。
在10次尝试中的一次,Kimi K3在100M令牌限制内成功完成了“最后一批”网络演练。这表明,在得到指令并获得初始网络访问权限时,Kimi K3能够自主攻击小型、防御薄弱且存在漏洞的企业系统。然而,TLO在多个方面与现实环境不同。它缺乏主动防御者和防御工具,对可能触发安全警报的动作不施加任何惩罚,并且包含预设的攻击路径。
TLO的解决方案不再仅限于少数模型。在先前测试中,四个公开发布的闭源模型已解决TLO,其中最具能力的模型以6/10和7/10的成功率更可靠地完成任务。Kimi K3在标准的100M令牌限制内完成了1/10次尝试。
网站管理员:/cdn-cgi/l/email-protection#ee8a81c3998b8c838f9d9a8b9cae80879d9ac0898198 | 联系我们:https://www.nist.gov/contact | 我们的其他办公室:https://www.nist.gov/visit
An official website of the United States government
Official websites use .gov A .gov website belongs to an official government organization in the United States.
Secure .gov websites use HTTPS A lock ( Lock A locked padlock ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) (UK AISI / CAISI) conducted a joint evaluation of Moonshot AI’s latest model, Kimi K3 (released on July 16, 2026 and slated for open-weight release by July 27, 2026). This evaluation focused on Kimi K3's cyber capabilities and found that:
Figure 1: Performance of Kimi K3 and other models on an exploit development benchmark (ExploitBench) . Higher success rate indicates greater cyber capability. Error bars represent 95% confidence intervals. ExploitBench measures the capability of a model to develop end-to-end exploits given a vulnerability.
These results represent preliminary evaluations on a small set of public and private benchmarks. U.S. closed-weight models were evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities. Publicly available versions of these models have these safeguards enabled. Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations. Detailed methodologies are provided in individual sections.
The cyber capability of models is aggregated across multiple tasks from multiple benchmarks using an approach inspired by Item Response Theory (IRT). For details of the methodology please see prior published reports:https://www.nist.gov/system/files/documents/2026/07/17/CAISI%20-%20Assessment%20of%20Z.ai%27s%20GLM-5.2.pdf. Kimi K3’s overall cyber capability has a larger confidence interval than other models because it was estimated from a single benchmark (ExploitBench, which has 41 tasks focused on exploit development). ExploitBench is a leading benchmark to measure a model’s ability to progress along the software exploitation ladder. All other models’ overall cyber capability scores were derived from a larger number of tasks that covered additional domains of cyber capability.
Figure 2: Preliminary comparison of aggregate capabilities over time of the most capable U.S. and PRC models as of Kimi K3’s release. The U.S. trendline is composed of results from frontier U.S. models . A 400-point increase on the y-axis equates to a 10x increase in the odds of solving tasks. Error bars and shaded regions denote 95% CIs.
ExploitBench:https://arxiv.org/abs/2605.14153 is a public benchmark, developed by Carnegie Mellon University, that measures a model’s ability to progress along the software exploitation ladder, including coverage and crash reproduction, arbitrary read/write, control flow hijack, and arbitrary code execution. The benchmark tests models on 41 recent (post-2023) vulnerabilities in the V8 engine (the JavaScript and WebAssembly software that powers Chrome).
ExploitBench results are presented in Figures 1 and 3.
Figure 3: Detailed ExploitBench performance for Kimi K3 and other models . Darker shading indicates greater cyber capability. Each row represents a key milestone in the exploit development chain, and each cell shows the number of ExploitBench tasks for which the model(s) in question were able to reach that milestone.
Kimi K3 outperforms GLM-5.2, the most cyber-capable open-weight model :https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber as of June 2026 . Kimi K3 achieves a score of 32%, whereas GLM-5.2 achieves a score of 24% (Figure 1).
Unlike the most cyber-capable models, Kimi K3 failed to develop exploits that achieved arbitrary code execution (ACE) for ExploitBench tasks. ACE is the highest-severity outcome in exploit development, granting attackers the ability to hijack a target. Kimi K3 achieved ACE on 0/41 samples, whereas the most cyber-capable models achieved ACE on 20/41 samples on average (Figure 3).
“The Last Ones” (TLO) cyber range:https://www.aisi.gov.uk/research/measuring-ai-agents-progress-on-multi-step-cyber-attack-scenarios is a 32-step simulated corporate network attack spanning 4 subnets and approximately 20 hosts, which would take a human expert roughly 20 hours to complete. Cyber ranges are expert-built, simulated networks of hosts, services, and vulnerabilities arranged into sequential attack chains that begin at the point of initial network access, and can be used to measure a model's ability to conduct end-to-end cyberattacks autonomously .
On this evaluation, Kimi K3 performs significantly below the leading U.S. cyber capable models. Specifically, Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average.
Kimi K3 outperforms GLM-5.2, the most cyber-capable open-weight model :https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber as of June 2026. Within the 100M-token limit, Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.
In one of the 10 attempts, Kimi K3 successfully completes “The Last Ones” cyber range within the 100M token limit. This indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access. However, TLO differs from real-world environments in several ways. It lacks active defenders and defensive tooling, imposes no penalty for actions that would trigger security alerts, and contains an intentional attack path.
Solves of TLO are no longer exclusive to a small set of models. In prior testing, four publicly released closed-weight models have solved TLO, with the most capable models solving it more reliably at 6/10 and 7/10 attempts. Kimi K3 solved it in 1/10 attempts within the standard 100M token limit.
Webmaster:/cdn-cgi/l/email-protection#ee8a81c3998b8c838f9d9a8b9cae80879d9ac0898198 | Contact Us:https://www.nist.gov/contact | Our Other Offices:https://www.nist.gov/visit