{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T07:00:52.885Z","headline":"Kimi K3 在网络安全漏洞利用测试中大幅落后美国前沿模型，知识蒸馏或为原因","description":"英国AI安全研究所与美国AI标准与创新中心联合评估显示，月之暗面的Kimi K3在ExploitBench基准上得分32.2%，远低于美国领先模型的76.2%，但优于智谱GLM-5.2的24.4%。","url":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/","mainEntityOfPage":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/","datePublished":"2026-07-24T09:48:32.000Z","dateModified":"2026-07-24T09:48:32.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why","https://aihot.virxact.com/items/cmryrih7804c9rolge6wdk3v8"],"canonicalUrl":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/","directAnswer":{"@type":"Answer","text":"英国AI安全研究所与美国AI标准与创新中心的联合评估显示，Kimi K3在ExploitBench上的得分为32.2%，低于美国领先模型的76.2%，但高于GLM-5.2的24.4%。","url":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/","dateCreated":"2026-07-24T09:48:32.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why","datePublished":"2026-07-24T09:48:32.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmryrih7804c9rolge6wdk3v8","datePublished":"2026-07-24T09:48:32.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmryrih7804c9rolge6wdk3v8"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why"},"article":{"id":"cmryrih7804c9rolge6wdk3v8","slug":"cmryrih7804c9rolge6wdk3v8","url":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/","title":"Kimi K3 在网络安全漏洞利用测试中大幅落后美国前沿模型，知识蒸馏或为原因","title_en":"Kimi K3 trails frontier US models by a wide margin on cyber exploits， and distillation may explain why","summary":"英国AI安全研究所与美国AI标准与创新中心联合评估显示，月之暗面的Kimi K3在ExploitBench基准上得分32.2%，远低于美国领先模型的76.2%，但优于智谱GLM-5.2的24.4%。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why","aiHotUrl":"https://aihot.virxact.com/items/cmryrih7804c9rolge6wdk3v8","publishedAt":"2026-07-24T09:48:32.000Z","category":"行业动态","score":73,"selected":true,"articleBody":["The British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) jointly evaluated Moonshot AI's latest model, Kimi K3.","Kimi K3 trails the leading U.S. frontier models by a wide margin on offensive cyber tasks but outperforms China's GLM-5.2, setting a new benchmark among open-weight models. Its safeguards didn't block exploit development or offensive cyber operations, and the model assisted with both without pushback.","The institutes used ExploitBench：https://the-decoder.com/new-benchmark-shows-claude-mythos-and-gpt-5-5-can-develop-real-browser-exploits-autonomously/, a benchmark developed by Carnegie Mellon University, to test exploit development skills. It uses 41 vulnerabilities found in Chrome's V8 engine after 2023 to track how far a model advances through the software exploitation process. The leading U.S. models averaged 76.2 percent, compared with 32.2 percent for Kimi K3 and 24.4 percent for GLM-5.2. Ad","Kimi K3 didn't reach the highest level, known as Arbitrary Code Execution (ACE), on any of the 41 tasks. ACE is the most severe exploit level because it gives attackers full control over a target system. The leading U.S. models achieved ACE in 20 of the 41 tasks. Ad DEC_D_Incontent-1","The institutes tested the U.S. closed-weight models with their system-level safeguards disabled to measure their maximum capabilities. Those safeguards are enabled in the publicly available versions.","The second test, \"The Last Ones\" (TLO), simulates a corporate network attack with a 32-step attack path across four subnets and about 20 hosts. A human expert would need roughly 20 hours to complete it, according to the institutes. Only a small group of models can solve TLO at all. Four publicly available closed-weight models have passed the test so far, with the strongest succeeding six or seven times out of ten. Ad","Kimi K3 reached step 17 out of 32 on average, compared with 28.5 steps for the leading U.S. models and just 11 for GLM-5.2. It completed the entire attack path in one of ten attempts while staying within the 100 million token limit, showing that it has the capability but can't call on it reliably. \"Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access\", the institute writes.","TLO doesn't account for active defense, so it isn't fully realistic. But the results would raise red flags in real-world scenarios. A fresh example showed up this week when OpenAI models tried to autonomously hack into Hugging Face：https://the-decoder.com/hugging-face-says-an-ai-agent-hacked-its-infrastructure-and-it-used-ai-to-fight-back/. Hugging Face fended off the attack, though it took real effort and the use of open-weight models：https://the-decoder.com/hugging-face-says-an-ai-agent-hacked-its-infrastructure-and-it-used-ai-to-fight-back/. Ad DEC_D_Incontent-2","A time-series analysis by CAISI tracks the cyber capabilities of U.S. and Chinese models since early 2025 on an Elo-based scale. Both trend lines are climbing, but Chinese models consistently remain behind their U.S. counterparts. Ad","In a previous analysis, the British institute pegged the performance gap for open models at four to seven months：https://the-decoder.com/open-weight-models-now-match-frontier-cyber-performance-from-just-four-months-ago-at-a-fraction-of-the-cost/, compared with six to ten months at the start of 2025. The new results fit this pattern. Chinese open-weight models are getting stronger, but they remain well behind leading U.S. systems.","AISI warns that this gap shouldn't breed complacency. The growing cyber capabilities of open models create \"a persistent and irreversible risk of misuse.\"","The Kimi findings also lend support to distillation allegations against Chinese model developers. U.S. science advisor Michael Kratsios：https://x.com/mkratsios47/status/2079933645888880708 recently accused Moonshot AI of \"distilling\" Anthropic's Fable：https://the-decoder.com/nadella-calls-out-ai-labs-like-openai-and-anthropic-for-banning-distillation-while-training-on-everyone-elses-data/ by using Fable's best outputs as training data to boost Kimi K3's performance. Kratsios also alleged that Moonshot AI had access to Nvidia's GB300s：https://the-decoder.com/nvidia-sets-new-mlperf-records-with-288-gpus-while-amd-and-intel-focus-on-different-battles/, which are subject to U.S. export controls.","One explanation for the gap between strong general benchmarks and weak cyber scores is that Kimi K3 may have been trained mostly on Claude outputs covering general knowledge, programming, and agent tasks. Anthropic's safety classifiers specifically block advanced offensive cyber queries：https://www.anthropic.com/news/fable-safeguards-jailbreak-framework, so those outputs would be underrepresented in a distillation dataset built from Claude responses. Kimi K3 could therefore match leading Western models on standard benchmarks without picking up their deeper exploit capabilities.","The AISI results support this reading. The institute disabled system-level safeguards on the U.S. models, revealing cyber capabilities that are nearly impossible to access through public interfaces and therefore largely unavailable for distillation.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/aisi_kimik3_cyber_eval-1.png","alt":"","afterParagraph":2,"url":"/media/articles/cmryrih7804c9rolge6wdk3v8/e45c8e834027412c.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/aisi_kimik3_cyber_eval-3.png","alt":"","afterParagraph":6,"url":"/media/articles/cmryrih7804c9rolge6wdk3v8/a83ffa8fc46ffcf8.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/aisi_kimik3_cyber_eval-2-scaled-1.png","alt":"","afterParagraph":8,"url":"/media/articles/cmryrih7804c9rolge6wdk3v8/7dd0eb5316024f28.png"}],"mediaStatus":"ok","articleBodyZh":["英国人工智能安全研究所（UK AISI）和美国人工智能标准与创新中心（CAISI）联合评估了Moonshot AI的最新模型Kimi K3。","在攻击性网络任务中，Kimi K3远远落后于美国的领先前沿模型，但它的表现优于中国的GLM-5.2，在开源权重模型中树立了新标杆。其安全防护未能阻止漏洞开发或攻击性网络操作，模型在这两方面都提供了辅助，没有任何反制。","两家研究所使用了由卡内基梅隆大学开发的ExploitBench基准：https://the-decoder.com/new-benchmark-shows-claude-mythos-and-gpt-5-5-can-develop-real-browser-exploits-autonomously/，来测试模型的漏洞开发技能。该基准使用了2023年后在Chrome的V8引擎中发现的41个漏洞，用于跟踪模型在软件漏洞利用过程中的进展程度。美国领先模型的平均得分为76.2%，而Kimi K3为32.2%，GLM-5.2为24.4%。","Kimi K3在这41个任务中未达到最高级别——任意代码执行（ACE）。ACE是最严重的漏洞等级，因为它允许攻击者完全控制目标系统。美国领先模型在41个任务中有20个实现了ACE。","研究所测试了美国的闭源模型，并在禁用其系统级安全防护的情况下测量其最大能力。这些安全防护在公开可用版本中都是启用的。","第二个测试，“最后者”（TLO），模拟了一个包含四个子网、约20台主机的32步企业网络攻击路径。据研究所称，一个人类专家大约需要20小时才能完成。只有少数模型能够完全解决TLO。目前已有四个公开可用的闭源模型通过了该测试，其中最强的模型十次中成功六到七次。","Kimi K3 平均达到32步中的第17步，而领先的美国模型平均为28.5步，GLM-5.2仅为11步。它在十次尝试中完成了整个攻击路径的一次，同时保持在一亿令牌限制内，这表明它具备能力，但不能可靠地调用。研究所写道：“当被指示并获得初始网络访问权限时，Kimi K3 有能力自主攻击小型、防御薄弱且易受攻击的企业系统。”","TLO 并未考虑主动防御，因此其结果并不完全现实。但这些结果在现实场景中会引发警示。本周出现了一个新的例子，当 OpenAI 模型试图自主攻击 Hugging Face 时：https://the-decoder.com/hugging-face-says-an-ai-agent-hacked-its-infrastructure-and-it-used-ai-to-fight-back/。Hugging Face 成功防御了此次攻击，尽管这需要实际的努力和使用开放权重模型：https://the-decoder.com/hugging-face-says-an-ai-agent-hacked-its-infrastructure-and-it-used-ai-to-fight-back/。","CAISI 的时间序列分析跟踪了美国和中国模型自2025年初以来的网络能力，基于 Elo 评分。两条趋势线都在上升，但中国模型始终落后于美国同类模型。","在之前的一次分析中，英国研究所将开放模型的性能差距评估为四到七个月：https://the-decoder.com/open-weight-models-now-match-frontier-cyber-performance-from-just-four-months-ago-at-a-fraction-of-the-cost/，而在2025年初这一差距为六到十个月。新的结果符合这一模式。中国的开放权重模型正在变强，但仍远落后于领先的美国系统。","AISI 警告称，这一差距不应让人自满。开放模型日益增强的网络能力带来了“持续且不可逆的误用风险。”","Kimi的研究结果也支持针对中国模型开发者的蒸馏指控。美国科学顾问Michael Kratsios：https://x.com/mkratsios47/status/2079933645888880708 最近指控Moonshot AI通过使用Anthropic的Fable：https://the-decoder.com/nadella-calls-out-ai-labs-like-openai-and-anthropic-for-banning-distillation-while-training-on-everyone-elses-data/的最佳输出作为训练数据来提升Kimi K3的性能，从而进行“蒸馏”。Kratsios还指称Moonshot AI能够访问Nvidia的GB300s：https://the-decoder.com/nvidia-sets-new-mlperf-records-with-288-gpus-while-amd-and-intel-focus-on-different-battles/，而这些设备受到美国出口管制。","强大的通用基准与较弱的网络安全评分之间的差距，一个解释是Kimi K3可能主要在Claude的输出上进行训练，这些输出涵盖通用知识、编程和代理任务。Anthropic的安全分类器会专门阻挡高级攻击性网络查询：https://www.anthropic.com/news/fable-safeguards-jailbreak-framework，因此这些输出在由Claude响应构建的蒸馏数据集中将被低估。因此，Kimi K3可以在标准基准上与领先的西方模型匹配，而不会掌握它们更深层的利用能力。","AISI的结果支持这一解读。该研究所禁用了美国模型的系统级安全防护，从而揭示了几乎无法通过公共接口访问的网络能力，这些能力因此在蒸馏中大多不可用。","保持对AI的关注。清晰、有用，不卖弄。","关注The Decoder，获取AI新闻、背景故事和专家分析。","The Decoder：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"英国AI安全研究所与美国AI标准与创新中心的联合评估显示，Kimi K3在ExploitBench上的得分为32.2%，低于美国领先模型的76.2%，但高于GLM-5.2的24.4%。","background":"评估采用ExploitBench和“The Last Ones”两项测试。前者覆盖41个2023年后发现的Chrome V8漏洞，后者模拟跨四个子网、约20台主机的32步企业网络攻击路径。","viewpoint":"Aioga判断，Kimi K3已表现出一定自主漏洞利用与网络攻击能力，但在测试中的稳定性和最高水平推进能力明显有限。标题提到的知识蒸馏可能是原因之一，材料未提供直接验证。","implications":"值得关注的是，Kimi K3在ExploitBench的41项任务中未达到任意代码执行级别，而美国领先模型在其中20项达到该级别。TLO结果也显示，其平均完成17步，低于美国领先模型的28.5步。","nextStep":"后续应关注评估方法、模型版本、测试配置及防护开关对结果的影响，并进一步核实知识蒸馏与性能差异之间是否存在可重复的证据链。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-26T07:49:34.717Z","sourceHash":"01353b753a223279","review":{"approved":true,"groundedness":97,"clarity":93,"duplicationRisk":18,"blockingIssues":[],"notes":["候选内容准确复述了ExploitBench与TLO的主要数据，并明确区分了材料事实与“Aioga判断”。","对知识蒸馏成因使用了“可能”并注明材料未提供直接验证，避免将标题中的推测表述为已证实事实。","可选措辞优化：首次出现“Aioga”时可简要说明其身份或将其改为“编辑判断”，以提升读者理解，但不构成阻断问题。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"Kimi K3 在网络安全漏洞利用测试中大幅落后美国前沿模型，知识蒸馏或为原因","summary":"英国AI安全研究所与美国AI标准与创新中心联合评估显示，月之暗面的Kimi K3在ExploitBench基准上得分32.2%，远低于美国领先模型的76.2%，但优于智谱GLM-5.2的24.4%。","category":"行业动态","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 在网络安全漏洞利用测试中大幅落后美国前沿模型，知识蒸馏或为原因 - Aioga AI资讯","description":"英国AI安全研究所与美国AI标准与创新中心联合评估显示，月之暗面的Kimi K3在ExploitBench基准上得分32.2%，远低于美国领先模型的76.2%，但优于智谱GLM-5.2的24.4%。","url":"https://www.aioga.com/news/cmryrih7804c9rolge6wdk3v8/"},"en":{"title":"Kimi K3 lags far behind leading American models in cybersecurity vulnerability exploitation tests, with knowledge distillation possibly being the reason","summary":"A joint assessment by the UK AI Safety Institute and the US Center for AI Standards and Innovation shows that Kimi K3 from Moon's Dark Side scored 32.2% on the ExploitBench benchmark, far below the 76.2% of leading US models, but higher than Zhishu GLM-5.2's 24.4%.","category":"Industry","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 lags far behind leading American models in cybersecurity vulnerability exploitation tests, with knowledge distillation possibly being the reason - Aioga AI News","description":"A joint assessment by the UK AI Safety Institute and the US Center for AI Standards and Innovation shows that Kimi K3 from Moon's Dark Side scored 32.2% on the ExploitBench benchma...","url":"https://www.aioga.com/en/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:25:11.599Z"},"ja":{"title":"Kimi K3はサイバーセキュリティの脆弱性利用テストでアメリカの最先端モデルに大きく遅れ、知識蒸留が原因かもしれない","summary":"英国AI安全研究所と米国AI標準・イノベーションセンターの共同評価によると、月の裏側のKimi K3はExploitBenchベンチマークで32.2%のスコアを記録し、米国の先進モデルの76.2%を大きく下回ったが、Zhizhi GLM-5.2の24.4%よりは優れていた。","category":"業界動向","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3はサイバーセキュリティの脆弱性利用テストでアメリカの最先端モデルに大きく遅れ、知識蒸留が原因かもしれない - Aioga AIニュース","description":"英国AI安全研究所と米国AI標準・イノベーションセンターの共同評価によると、月の裏側のKimi K3はExploitBenchベンチマークで32.2%のスコアを記録し、米国の先進モデルの76.2%を大きく下回ったが、Zhizhi GLM-5.2の24.4%よりは優れていた。","url":"https://www.aioga.com/ja/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:25:21.980Z"},"ko":{"title":"Kimi K3는 사이버 보안 취약점 이용 테스트에서 미국 최첨단 모델에 크게 뒤처졌으며, 지식 증류가 그 원인일 수 있다","summary":"영국 AI 안전 연구소와 미국 AI 표준 및 혁신 센터의 공동 평가에 따르면, 월의 암면 Kimi K3는 ExploitBench 벤치마크에서 32.2%의 점수를 기록하여 미국 선도 모델의 76.2%보다 훨씬 낮았지만, 지휘 GLM-5.2의 24.4%보다는 높았다.","category":"업계 동향","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3는 사이버 보안 취약점 이용 테스트에서 미국 최첨단 모델에 크게 뒤처졌으며, 지식 증류가 그 원인일 수 있다 - Aioga AI 뉴스","description":"영국 AI 안전 연구소와 미국 AI 표준 및 혁신 센터의 공동 평가에 따르면, 월의 암면 Kimi K3는 ExploitBench 벤치마크에서 32.2%의 점수를 기록하여 미국 선도 모델의 76.2%보다 훨씬 낮았지만, 지휘 GLM-5.2의 24.4%보다는 높았다.","url":"https://www.aioga.com/ko/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:25:58.673Z"},"es":{"title":"Kimi K3 está muy por detrás de los modelos avanzados de Estados Unidos en las pruebas de explotación de vulnerabilidades de ciberseguridad, la destilación de conocimientos podría ser la razón","summary":"La evaluación conjunta del Instituto de Seguridad de IA del Reino Unido y el Centro de Estándares e Innovación en IA de Estados Unidos muestra que Kimi K3 de 'La cara oculta de la luna' obtuvo un 32,2 % en el benchmark ExploitBench, muy por debajo del 76,2 % del modelo líder estadounidense, pero superior al 24,4 % de Zhipu GLM-5.2.","category":"Industria","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 está muy por detrás de los modelos avanzados de Estados Unidos en las pruebas de explotación de vulnerabilidades de ciberseguridad, la destilación de conocimientos podría ser la razón - Aioga Noticias de IA","description":"La evaluación conjunta del Instituto de Seguridad de IA del Reino Unido y el Centro de Estándares e Innovación en IA de Estados Unidos muestra que Kimi K3 de 'La cara oculta de la...","url":"https://www.aioga.com/es/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:25:57.475Z"},"fr":{"title":"Kimi K3 accuse un retard considérable par rapport aux modèles de pointe américains dans les tests d'exploitation des vulnérabilités en cybersécurité, la distillation des connaissances pourrait en être la cause","summary":"Une évaluation conjointe de l'Institut britannique de sécurité en IA et du Centre américain pour les normes et l'innovation en IA montre que le Kimi K3 de Moon's Dark Side a obtenu un score de 32,2 % sur le benchmark ExploitBench, bien inférieur aux 76,2 % des modèles américains de pointe, mais supérieur aux 24,4 % de Zhipu GLM-5.2.","category":"Industrie","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 accuse un retard considérable par rapport aux modèles de pointe américains dans les tests d'exploitation des vulnérabilités en cybersécurité, la distillation des connaissances pourrait en être la cause - Aioga Actualités IA","description":"Une évaluation conjointe de l'Institut britannique de sécurité en IA et du Centre américain pour les normes et l'innovation en IA montre que le Kimi K3 de Moon's Dark Side a obtenu...","url":"https://www.aioga.com/fr/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:26:34.246Z"},"de":{"title":"Kimi K3 liegt beim Testen von Sicherheitslücken im Bereich der Cybersicherheit deutlich hinter den führenden US-Modellen zurück, Wissensdestillation könnte der Grund sein","summary":"Eine gemeinsame Bewertung des britischen AI-Sicherheitsinstituts und des US-amerikanischen AI-Standards- und Innovationszentrums zeigt, dass Kimi K3 auf der dunklen Seite des Mondes im ExploitBench-Benchmark 32,2 % erreicht, weit unter den 76,2 % des führenden US-Modells, aber besser als die 24,4 % von Zhipu GLM-5.2.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 liegt beim Testen von Sicherheitslücken im Bereich der Cybersicherheit deutlich hinter den führenden US-Modellen zurück, Wissensdestillation könnte der Grund sein - Aioga KI-News","description":"Eine gemeinsame Bewertung des britischen AI-Sicherheitsinstituts und des US-amerikanischen AI-Standards- und Innovationszentrums zeigt, dass Kimi K3 auf der dunklen Seite des Monde...","url":"https://www.aioga.com/de/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:26:34.504Z"},"pt-BR":{"title":"Kimi K3 está muito atrás dos modelos avançados dos EUA em testes de exploração de vulnerabilidades de segurança na rede, e a destilação de conhecimento pode ser a causa","summary":"Uma avaliação conjunta do Instituto de Segurança de IA do Reino Unido e do Centro de Padrões e Inovação de IA dos EUA mostrou que o Kimi K3 do lado obscuro da lua obteve 32,2% na referência ExploitBench, muito abaixo dos 76,2% dos principais modelos dos EUA, mas acima dos 24,4% do Zhipu GLM-5.2.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 está muito atrás dos modelos avançados dos EUA em testes de exploração de vulnerabilidades de segurança na rede, e a destilação de conhecimento pode ser a causa - Aioga Notícias de IA","description":"Uma avaliação conjunta do Instituto de Segurança de IA do Reino Unido e do Centro de Padrões e Inovação de IA dos EUA mostrou que o Kimi K3 do lado obscuro da lua obteve 32,2% na r...","url":"https://www.aioga.com/pt-BR/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:27:09.612Z"},"ru":{"title":"Kimi K3 значительно отстает от передовых моделей США в тестировании эксплуатации уязвимостей кибербезопасности, причиной может быть дистилляция знаний","summary":"Совместная оценка Британского института безопасности ИИ и Американского центра стандартов и инноваций в области ИИ показала, что Kimi K3 из серии \"Темная сторона Луны\" набрал 32,2% по эталонному тесту ExploitBench, что значительно ниже 76,2% ведущей модели США, но выше 24,4% модели Zhipu GLM-5.2.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 значительно отстает от передовых моделей США в тестировании эксплуатации уязвимостей кибербезопасности, причиной может быть дистилляция знаний - Aioga Новости ИИ","description":"Совместная оценка Британского института безопасности ИИ и Американского центра стандартов и инноваций в области ИИ показала, что Kimi K3 из серии \"Темная сторона Луны\" набрал 32,2%...","url":"https://www.aioga.com/ru/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:27:17.454Z"},"ar":{"title":"كيمي K3 متأخرة بشكل كبير عن النماذج المتقدمة الأمريكية في اختبار استغلال الثغرات الأمنية على الإنترنت، وقد يكون تقطير المعرفة هو السبب","summary":"أظهر تقييم مشترك بين معهد أبحاث أمان الذكاء الاصطناعي في المملكة المتحدة ومركز الولايات المتحدة لمعايير وابتكار الذكاء الاصطناعي أن Kimi K3 من Moon's Dark Side حصل على 32.2% في معيار ExploitBench، وهو أدنى بكثير من النموذج الرائد في الولايات المتحدة بنسبة 76.2%، ولكنه أعلى من Zhipu GLM-5.2 بنسبة 24.4%.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"كيمي K3 متأخرة بشكل كبير عن النماذج المتقدمة الأمريكية في اختبار استغلال الثغرات الأمنية على الإنترنت، وقد يكون تقطير المعرفة هو السبب - Aioga أخبار الذكاء الاصطناعي","description":"أظهر تقييم مشترك بين معهد أبحاث أمان الذكاء الاصطناعي في المملكة المتحدة ومركز الولايات المتحدة لمعايير وابتكار الذكاء الاصطناعي أن Kimi K3 من Moon's Dark Side حصل على 32.2% في معي...","url":"https://www.aioga.com/ar/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:27:58.536Z"},"hi":{"title":"Kimi K3 साइबर सुरक्षा भिन्नताओं के परीक्षण में अमेरिका के अग्रणी मॉडलों से काफी पीछे है, ज्ञान आसवन संभवतः इसका कारण है","summary":"ब्रिटेन के एआई सुरक्षा संस्थान और अमेरिका के एआई मानक और नवाचार केंद्र के संयुक्त मूल्यांकन से पता चलता है कि चंद्रमा के अंधेरे पक्ष का Kimi K3 ExploitBench बेंचमार्क पर 32.2% अंक प्राप्त करता है, जो अमेरिकी अग्रणी मॉडलों के 76.2% की तुलना में काफी कम है, लेकिन Zhipu GLM-5.2 के 24.4% से बेहतर है।","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 साइबर सुरक्षा भिन्नताओं के परीक्षण में अमेरिका के अग्रणी मॉडलों से काफी पीछे है, ज्ञान आसवन संभवतः इसका कारण है - Aioga AI समाचार","description":"ब्रिटेन के एआई सुरक्षा संस्थान और अमेरिका के एआई मानक और नवाचार केंद्र के संयुक्त मूल्यांकन से पता चलता है कि चंद्रमा के अंधेरे पक्ष का Kimi K3 ExploitBench बेंचमार्क पर 32.2% अंक...","url":"https://www.aioga.com/hi/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:28:01.884Z"},"it":{"title":"Kimi K3 è molto indietro rispetto ai modelli all'avanguardia degli Stati Uniti nei test di sfruttamento delle vulnerabilità della sicurezza informatica, la distillazione della conoscenza potrebbe essere la causa","summary":"La valutazione congiunta dell'Istituto britannico per la sicurezza dell'IA e del Centro americano per gli standard e l'innovazione dell'IA mostra che il Kimi K3 del lato oscuro della luna ha ottenuto il 32,2% nel benchmark ExploitBench, ben al di sotto del 76,2% dei modelli leader statunitensi, ma superiore al 24,4% del Zhipu GLM-5.2.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 è molto indietro rispetto ai modelli all'avanguardia degli Stati Uniti nei test di sfruttamento delle vulnerabilità della sicurezza informatica, la distillazione della conoscenza potrebbe essere la causa - Aioga Notizie IA","description":"La valutazione congiunta dell'Istituto britannico per la sicurezza dell'IA e del Centro americano per gli standard e l'innovazione dell'IA mostra che il Kimi K3 del lato oscuro del...","url":"https://www.aioga.com/it/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:28:38.661Z"},"nl":{"title":"Kimi K3 loopt sterk achter op geavanceerde Amerikaanse modellen bij tests voor het misbruiken van beveiligingslekken op het gebied van cybersecurity; kennisdistillatie kan de oorzaak zijn","summary":"Een gezamenlijke beoordeling door het Britse AI Security Institute en het Amerikaanse AI Standards and Innovation Center toonde aan dat Kimi K3 van de donkere zijde van de maan een score van 32,2% behaalde op de ExploitBench-benchmark, ver onder het Amerikaanse toonaangevende model van 76,2%, maar beter dan Zhizhu GLM-5.2's 24,4%.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 loopt sterk achter op geavanceerde Amerikaanse modellen bij tests voor het misbruiken van beveiligingslekken op het gebied van cybersecurity; kennisdistillatie kan de oorzaak zijn - Aioga AI-nieuws","description":"Een gezamenlijke beoordeling door het Britse AI Security Institute en het Amerikaanse AI Standards and Innovation Center toonde aan dat Kimi K3 van de donkere zijde van de maan een...","url":"https://www.aioga.com/nl/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:28:38.420Z"},"tr":{"title":"Kimi K3, siber güvenlik açıklarından faydalanma testlerinde ABD'nin ileri düzey modellerinin oldukça gerisinde kaldı; bilgi damıtımı bunun nedeni olabilir","summary":"İngiltere AI Güvenliği Enstitüsü ile ABD AI Standartları ve Yenilik Merkezi'nin ortak değerlendirmesi, Kimi K3'ün ExploitBench kıyaslamasında %32,2 puan aldığını ve ABD'nin önde gelen modellerinin %76,2'sinin çok altında olduğunu, ancak Zhizhu GLM-5.2'nin %24,4'ünden daha yüksek olduğunu gösteriyor.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3, siber güvenlik açıklarından faydalanma testlerinde ABD'nin ileri düzey modellerinin oldukça gerisinde kaldı; bilgi damıtımı bunun nedeni olabilir - Aioga AI Haberleri","description":"İngiltere AI Güvenliği Enstitüsü ile ABD AI Standartları ve Yenilik Merkezi'nin ortak değerlendirmesi, Kimi K3'ün ExploitBench kıyaslamasında %32,2 puan aldığını ve ABD'nin önde ge...","url":"https://www.aioga.com/tr/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:29:17.187Z"},"vi":{"title":"Kimi K3 tụt lại phía sau đáng kể so với các mô hình tiên tiến của Mỹ trong thử nghiệm khai thác lỗ hổng bảo mật mạng, việc chưng cất kiến thức có thể là nguyên nhân","summary":"Viện Nghiên cứu An toàn AI của Anh và Trung tâm Tiêu chuẩn và Đổi mới AI của Mỹ cùng đánh giá cho thấy, Kimi K3 của Mặt Trăng Đen đạt 32,2% trên tiêu chuẩn ExploitBench, thấp hơn nhiều so với mô hình hàng đầu của Mỹ là 76,2%, nhưng cao hơn GLM-5.2 của Zhìpǔ là 24,4%.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 tụt lại phía sau đáng kể so với các mô hình tiên tiến của Mỹ trong thử nghiệm khai thác lỗ hổng bảo mật mạng, việc chưng cất kiến thức có thể là nguyên nhân - Tin tức AI Aioga","description":"Viện Nghiên cứu An toàn AI của Anh và Trung tâm Tiêu chuẩn và Đổi mới AI của Mỹ cùng đánh giá cho thấy, Kimi K3 của Mặt Trăng Đen đạt 32,2% trên tiêu chuẩn ExploitBench, thấp hơn n...","url":"https://www.aioga.com/vi/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:29:16.896Z"},"id":{"title":"Kimi K3 tertinggal jauh di depan model Amerika dalam pengujian eksploitasi kerentanan keamanan siber, distilasi pengetahuan mungkin menjadi penyebabnya","summary":"Penilaian gabungan oleh Institut Keamanan AI Inggris dan Pusat Standar dan Inovasi AI Amerika Serikat menunjukkan bahwa Kimi K3 dari Bulan Gelap memperoleh skor 32,2% pada tolok ukur ExploitBench, jauh lebih rendah daripada model unggulan Amerika Serikat yang mencapai 76,2%, tetapi lebih tinggi daripada Zhipu GLM-5.2 yang memperoleh 24,4%.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 tertinggal jauh di depan model Amerika dalam pengujian eksploitasi kerentanan keamanan siber, distilasi pengetahuan mungkin menjadi penyebabnya - Berita AI Aioga","description":"Penilaian gabungan oleh Institut Keamanan AI Inggris dan Pusat Standar dan Inovasi AI Amerika Serikat menunjukkan bahwa Kimi K3 dari Bulan Gelap memperoleh skor 32,2% pada tolok uk...","url":"https://www.aioga.com/id/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:29:54.655Z"},"th":{"title":"Kimi K3 ล้าหลังกว่ารุ่นชั้นนำของสหรัฐฯ อย่างมากในการทดสอบการใช้ประโยชน์จากช่องโหว่ด้านความปลอดภัยทางไซเบอร์ การกลั่นความรู้อาจเป็นสาเหตุ","summary":"การประเมินร่วมกันระหว่างสถาบันวิจัยความปลอดภัย AI ของสหราชอาณาจักรและศูนย์มาตรฐานและนวัตกรรม AI ของสหรัฐอเมริกาแสดงให้เห็นว่า Kimi K3 ของ Moon's Dark Side ได้คะแนน 32.2% ในมาตรฐาน ExploitBench ซึ่งต่ำกว่ารุ่นชั้นนำของสหรัฐฯ ที่ได้ 76.2% แต่สูงกว่าของ Zhizhu GLM-5.2 ซึ่งได้ 24.4%","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 ล้าหลังกว่ารุ่นชั้นนำของสหรัฐฯ อย่างมากในการทดสอบการใช้ประโยชน์จากช่องโหว่ด้านความปลอดภัยทางไซเบอร์ การกลั่นความรู้อาจเป็นสาเหตุ - ข่าว AI Aioga","description":"การประเมินร่วมกันระหว่างสถาบันวิจัยความปลอดภัย AI ของสหราชอาณาจักรและศูนย์มาตรฐานและนวัตกรรม AI ของสหรัฐอเมริกาแสดงให้เห็นว่า Kimi K3 ของ Moon's Dark Side ได้คะแนน 32.2% ในมาตรฐาน...","url":"https://www.aioga.com/th/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:30:06.355Z"},"pl":{"title":"Kimi K3 w testach wykorzystania luk w cyberbezpieczeństwie znacznie odstaje od czołowych modeli amerykańskich, a powodem może być destylacja wiedzy","summary":"Wspólna ocena brytyjskiego Instytutu Bezpieczeństwa AI i amerykańskiego Centrum Standardów i Innowacji AI wykazała, że Kimi K3 z ciemnej strony Księżyca uzyskał wynik 32,2% w benchmarku ExploitBench, znacznie poniżej 76,2% wiodącego amerykańskiego modelu, ale przewyższa 24,4% modelu Zhipu GLM-5.2.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kimi K3 w testach wykorzystania luk w cyberbezpieczeństwie znacznie odstaje od czołowych modeli amerykańskich, a powodem może być destylacja wiedzy - Aioga Wiadomości AI","description":"Wspólna ocena brytyjskiego Instytutu Bezpieczeństwa AI i amerykańskiego Centrum Standardów i Innowacji AI wykazała, że Kimi K3 z ciemnej strony Księżyca uzyskał wynik 32,2% w bench...","url":"https://www.aioga.com/pl/news/cmryrih7804c9rolge6wdk3v8/","contentTranslated":true,"sourceHash":"77048f2199081b62","translatedAt":"2026-07-26T02:30:48.827Z"}}}}