{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T08:01:28.298Z","headline":"Moonshot AI 发布 PerceptionBench：多模态模型视觉感知能力诊断基准","description":"Moonshot AI 发布 PerceptionBench，一个从 40 多个现有基准中模型实际失败案例归纳出的视觉感知基准，包含 10 项原子感知能力和 3000 道验证题。所有测试模型准确率均未超过 60%，且大量正确答案在重复提问时无法复现，表明模型更多是猜测而非真正感知。PerceptionBench 旨在精确诊断多模态 AI 的视觉感知断裂点，推动其实现忠实、一致的视觉理解。","url":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/","mainEntityOfPage":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/","datePublished":"2026-07-15T16:00:00.000Z","dateModified":"2026-07-15T16:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.kimi.com/blog/perception-bench","https://aihot.virxact.com/items/cmrnvwztt01bebixyznbje716"],"canonicalUrl":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/","directAnswer":{"@type":"Answer","text":"Moonshot AI 发布 PerceptionBench，将前沿模型在40多个既有基准中的失败追溯至最早视觉原因，归纳10项原子感知能力，并形成3000道经过验证的问题。","url":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/","dateCreated":"2026-07-15T16:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Moonshot AI：Kimi Blog source article","url":"https://www.kimi.com/blog/perception-bench","datePublished":"2026-07-15T16:00:00.000Z","provider":{"@type":"Organization","name":"Moonshot AI：Kimi Blog","url":"https://www.kimi.com/blog/perception-bench"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrnvwztt01bebixyznbje716","datePublished":"2026-07-15T16:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrnvwztt01bebixyznbje716"}}],"aggregationSource":"Moonshot AI：Kimi Blog","originalPublisher":{"name":"Moonshot AI：Kimi Blog","url":"https://www.kimi.com/blog/perception-bench"},"article":{"id":"cmrnvwztt01bebixyznbje716","slug":"cmrnvwztt01bebixyznbje716","url":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/","title":"Moonshot AI 发布 PerceptionBench：多模态模型视觉感知能力诊断基准","title_en":"PerceptionBench","summary":"Moonshot AI 发布 PerceptionBench，一个从 40 多个现有基准中模型实际失败案例归纳出的视觉感知基准，包含 10 项原子感知能力和 3000 道验证题。所有测试模型准确率均未超过 60%，且大量正确答案在重复提问时无法复现，表明模型更多是猜测而非真正感知。PerceptionBench 旨在精确诊断多模态 AI 的视觉感知断裂点，推动其实现忠实、一致的视觉理解。","source":"Moonshot AI：Kimi Blog","sourceUrl":"https://www.kimi.com/blog/perception-bench","aiHotUrl":"https://aihot.virxact.com/items/cmrnvwztt01bebixyznbje716","publishedAt":"2026-07-15T16:00:00.000Z","category":"论文研究","score":70,"selected":true,"articleBody":["Evaluating Atomic Visual Perception in Multimodal Large Language Models","We are releasing PerceptionBench , a benchmark that isolates visual perception and evaluates it as a set of atomic capabilities— discovered from how today's models fail, not defined in advance. By attributing frontier-model failures across 40+ benchmarks to their earliest visual cause, we distill 10 perceptual capabilities and 3,000 verified questions, each answerable by looking, with no reasoning or outside knowledge required.","The result is a sharp diagnosis rather than one more score. No model we evaluate clears 60% accuracy, and models with nearly identical overall scores can exhibit very different perceptual strengths and weaknesses. More strikingly, a large share of correct answers fail to survive a repeated ask—evidence that current models often guess rather than perceive. PerceptionBench is built to expose exactly where perception breaks, and to drive progress toward multimodal AI that sees faithfully and consistently.","Each source benchmark captures a narrow slice of perception errors, and these slices overlap only weakly (mean pairwise weighted Jaccard 0.20). No single benchmark—or small group of them—covers perception as a whole, which motivates a capability-centric benchmark that aggregates and rebalances these fragmented views.","The dataset consists of 3,000 high-quality verified samples. The distribution aims to isolate atomic perceptual capabilities from confounding factors, and distinguishes itself through three core design principles:","Note on Quality : To make the benchmark a reliable gold standard, all samples underwent rigorous verification and difficulty-balancing, keeping only genuinely perceptual failures with a single verifiable answer.","Which part of the mug does the purple line connect to?","How many different colors of hats appear?","How many red dots are there in the picture?","How many equal subintervals is the u₁ axis divided into?","How many staff lines does the rightmost note in the figure touch?","How many faces of these small cubes are in direct contact with the ground?","What is the digit in the lower-right box?","How many animal silhouettes in the picture are exactly the same?","How many black queens are on the chessboard?","How many monkeys have touched the wheel cover?","Cell x lies in the intersection of how many circles?","How many yellow hollow rings appear in the figure?","How many people are inside the truck?","PerceptionBench is a simple but challenging benchmark for evaluating atomic visual perception in frontier models. It measures what multimodal models actually see rather than what they infer, providing a faithful and fine-grained diagnosis of the perceptual capabilities of current and future multimodal models."],"articleImages":[{"sourceUrl":"https://www.kimi.com/images/blog/perception-bench/openbench_jaccard.svg","alt":"Per-benchmark error-type distributions and their weak pairwise overlap across existing benchmarks","afterParagraph":3,"url":"/media/articles/cmrnvwztt01bebixyznbje716/453020717ab10cda.jpg"},{"sourceUrl":"https://kimi-file.moonshot.cn/prod-chat-kimi/kfs/4/2/2026-07-16/d9cf6gff2ena6204rq2g?x-tos-process=image%2Fauto-orient%2C1%2Fstrip%2Fignore-error%2C1","alt":"Visual Localization example","afterParagraph":5,"url":"/media/articles/cmrnvwztt01bebixyznbje716/4eccfd873c7a9954.jpg"},{"sourceUrl":"https://kimi-file.moonshot.cn/prod-chat-kimi/kfs/4/2/2026-07-16/d9cf6gudcmosb3rnfku0?x-tos-process=image%2Fauto-orient%2C1%2Fstrip%2Fignore-error%2C1","alt":"Visual Localization example","afterParagraph":6,"url":"/media/articles/cmrnvwztt01bebixyznbje716/c316e9f5806fff85.jpg"},{"sourceUrl":"https://kimi-file.moonshot.cn/prod-chat-kimi/kfs/4/2/2026-07-16/d9cf6giav1fc645q5gm0?x-tos-process=image%2Fauto-orient%2C1%2Fstrip%2Fignore-error%2C1","alt":"Visual Attribute example","afterParagraph":6,"url":"/media/articles/cmrnvwztt01bebixyznbje716/6d77de2faa6c3b31.jpg"},{"sourceUrl":"https://kimi-file.moonshot.cn/prod-chat-kimi/kfs/4/2/2026-07-16/d9cf6fpl51jas5brmvag?x-tos-process=image%2Fauto-orient%2C1%2Fstrip%2Fignore-error%2C1","alt":"Visual Attribute example","afterParagraph":7,"url":"/media/articles/cmrnvwztt01bebixyznbje716/c47dedf01aaecd0c.png"},{"sourceUrl":"https://kimi-file.moonshot.cn/prod-chat-kimi/kfs/4/2/2026-07-16/d9cf6gvf2ena6204rq3g?x-tos-process=image%2Fauto-orient%2C1%2Fstrip%2Fignore-error%2C1","alt":"Visual Counting example","afterParagraph":8,"url":"/media/articles/cmrnvwztt01bebixyznbje716/865f7151f18fe61a.jpg"}],"mediaStatus":"ok","articleBodyZh":["在多模态大语言模型中评估原子视觉感知能力","我们发布了 PerceptionBench，这是一个将视觉感知孤立出来并以一组原子能力进行评估的基准——这些能力是从当今模型的失败中发现的，而非事先定义的。通过将前沿模型在40个基准上的失败归因于其最早的视觉原因，我们提炼出了10种感知能力和3,000个经过验证的问题，每个问题都可以通过观察回答，无需推理或外部知识。","其结果是精确的诊断，而不是又一个分数。我们评估的没有一个模型能达到60%的准确率，且整体分数几乎相同的模型可能表现出截然不同的感知优劣。更令人惊讶的是，大量正确答案在重复提问时未能维持——这表明当前模型常常是猜测而非真正感知。PerceptionBench旨在暴露感知断裂的具体位置，并推动多模态AI朝着忠实且稳定的视觉理解迈进。","每个原始基准都捕捉了感知错误的一个狭窄切片，而这些切片之间的重叠很弱（平均成对加权Jaccard指数0.20）。没有单一基准或少数几个基准能覆盖感知的整体，这促使我们提出以能力为中心的基准，将这些零散视角进行汇总和平衡。","该数据集包含3,000个高质量的经过验证的样本。其分布旨在从干扰因素中孤立出原子感知能力，并且通过三个核心设计原则自成特色：","质量说明：为了使基准成为可靠的黄金标准，所有样本都经过严格验证和难度平衡，仅保留真正的感知失败样本，并保证每个问题有唯一可验证的答案。","紫色的线连接到杯子的哪一部分？","出现了多少种不同颜色的帽子？","图片中有多少个红点？","u₁轴被分成多少个相等的子区间？","图中最右侧的音符接触了多少条五线谱线？","这些小立方体有多少个面直接接触地面？","右下角的方框内的数字是多少？","图片中有多少个动物剪影完全相同？","棋盘上有多少个黑色皇后？","有多少只猴子碰过方向盘罩？","单元格 x 位于多少个圆的交集中？","图中出现了多少个黄色空心环？","卡车里有多少人？","PerceptionBench 是一个简单但具有挑战性的基准测试，用于评估前沿模型的基本视觉感知能力。它衡量的是多模态模型实际看到的内容，而不是它们推断的内容，从而为当前和未来的多模态模型的感知能力提供了真实而细致的诊断。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Moonshot AI 发布 PerceptionBench，将前沿模型在40多个既有基准中的失败追溯至最早视觉原因，归纳10项原子感知能力，并形成3000道经过验证的问题。","background":"现有基准各自覆盖较窄的感知错误类型，彼此重叠较弱，平均成对加权Jaccard为0.20。该基准因此采用能力导向方式，聚合并重新平衡分散的视觉感知测试。","viewpoint":"Aioga 判断，PerceptionBench的价值不只在于提供总分，而在于区分模型具体的感知强弱。受测模型准确率均未超过60%，且相近总分可能对应明显不同的能力结构。","implications":"值得关注的是，大量正确答案在重复提问时未能保持，材料将其视为模型经常猜测而非真正感知的证据。这可能促使多模态评测进一步重视答案一致性与具体失效位置。","nextStep":"Aioga 判断，后续应关注基准是否公布更细的能力维度结果、重复提问设置与样本验证方法，并观察模型改进能否同时提升准确率和回答稳定性，而非只提高综合得分。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-23T01:08:58.405Z","sourceHash":"9e824a803a6e8acf","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":18,"blockingIssues":[],"notes":["“Aioga 判断”明确标示了观点性质，未将建议冒充来源事实。","nextStep 中关于进一步关注能力维度结果、重复提问设置和样本验证方法的表述属于合理的后续观察建议，但并非来源材料已确认的后续计划。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["论文研究","Moonshot AI：Kimi Blog"],"translations":{"zh-CN":{"title":"Moonshot AI 发布 PerceptionBench：多模态模型视觉感知能力诊断基准","summary":"Moonshot AI 发布 PerceptionBench，一个从 40 多个现有基准中模型实际失败案例归纳出的视觉感知基准，包含 10 项原子感知能力和 3000 道验证题。所有测试模型准确率均未超过 60%，且大量正确答案在重复提问时无法复现，表明模型更多是猜测而非真正感知。PerceptionBench 旨在精确诊断多模态 AI 的视觉感知断裂点，推动其实现忠实、一致的视觉理解。","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI 发布 PerceptionBench：多模态模型视觉感知能力诊断基准 - Aioga AI资讯","description":"Moonshot AI 发布 PerceptionBench，一个从 40 多个现有基准中模型实际失败案例归纳出的视觉感知基准，包含 10 项原子感知能力和 3000 道验证题。所有测试模型准确率均未超过 60%，且大量正确答案在重复提问时无法复现，表明模型更多是猜测而非真正感知。PerceptionBench 旨在精确诊断多模态 AI 的视觉感知断裂点，推","url":"https://www.aioga.com/news/cmrnvwztt01bebixyznbje716/"},"en":{"title":"Moonshot AI releases PerceptionBench: a diagnostic benchmark for multimodal model visual perception capabilities","summary":"Moonshot AI released PerceptionBench, a visual perception benchmark summarized from over 40 existing benchmark cases where models have actually failed, including 10 atomic perception capabilities and 3,000 verification questions. All test models had accuracy rates below 60%, and many correct answers could not be reproduced when asked repeatedly, indicating that the models were more guessing than actual perception. PerceptionBench aims to precisely diagnose the visual perception breakpoints of multimodal AI, driving it to achieve faithful and consistent visual understanding.","category":"Research","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI releases PerceptionBench: a diagnostic benchmark for multimodal model visual perception capabilities - Aioga AI News","description":"Moonshot AI released PerceptionBench, a visual perception benchmark summarized from over 40 existing benchmark cases where models have actually failed, including 10 atomic percepti","url":"https://www.aioga.com/en/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:50.995Z"},"ja":{"title":"Moonshot AIがPerceptionBenchをリリース:マルチモーダルモデルの視覚知覚能力の診断ベンチマーク","summary":"Moonshot AIはPerceptionBenchをリリースしました。これは、実際に失敗したモデルが40件以上の既存のベンチマークケースをまとめた視覚知覚ベンチマークで、10件の原子知覚能力と3,000件の検証質問が含まれています。 すべてのテストモデルは60%未満の精度を持ち、多くの正解は繰り返し尋ねても再現できなかったため、モデルは実際の知覚よりも推測に偏っていたことが示唆されました。 PerceptionBenchは、マルチモーダルAIの視覚知覚のブレークポイントを正確に診断し、忠実かつ一貫した視覚的理解を達成することを目指しています。","category":"論文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AIがPerceptionBenchをリリース:マルチモーダルモデルの視覚知覚能力の診断ベンチマーク - Aioga AIニュース","description":"Moonshot AIはPerceptionBenchをリリースしました。これは、実際に失敗したモデルが40件以上の既存のベンチマークケースをまとめた視覚知覚ベンチマークで、10件の原子知覚能力と3,000件の検証質問が含まれています。 すべてのテストモデルは60%未満の精度を持ち、多くの正解は繰り返し尋ねても再現できなかったため、モデルは実際の知覚よりも推","url":"https://www.aioga.com/ja/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:51.458Z"},"ko":{"title":"Moonshot AI가 PerceptionBench를 출시하다: 다중 모달 모델 시각 인식 기능을 위한 진단 벤치마크","summary":"Moonshot AI는 PerceptionBench를 공개했는데, 이는 실제로 실패한 40개 이상의 기존 벤치마크 사례를 요약한 시각적 인지벤치마크로, 10개의 원자 인식 능력과 3,000개의 검증 질문을 포함합니다. 모든 테스트 모델의 정확도는 60% 미만이었고, 반복 질문을 해도 많은 정답을 재현하지 못해 모델들이 실제 인지보다는 추측에 더 가깝다는 것을 나타냈다. PerceptionBench는 멀티모달 AI의 시각적 인식 중단점을 정밀하게 진단하여 충실하고 일관된 시각적 이해를 달성하는 것을 목표로 합니다.","category":"연구","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI가 PerceptionBench를 출시하다: 다중 모달 모델 시각 인식 기능을 위한 진단 벤치마크 - Aioga AI 뉴스","description":"Moonshot AI는 PerceptionBench를 공개했는데, 이는 실제로 실패한 40개 이상의 기존 벤치마크 사례를 요약한 시각적 인지벤치마크로, 10개의 원자 인식 능력과 3,000개의 검증 질문을 포함합니다. 모든 테스트 모델의 정확도는 60% 미만이었고, 반복 질문을 해도 많은 정답을 재현하지 못해 모델들이 실","url":"https://www.aioga.com/ko/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:51.236Z"},"es":{"title":"Moonshot AI lanza PerceptionBench: un benchmark diagnóstico para las capacidades de percepción visual de modelos multimodales","summary":"Moonshot AI lanzó PerceptionBench, un benchmark visual de percepción resumido a partir de más de 40 casos existentes donde los modelos han fallado realmente, incluyendo 10 capacidades de percepción atómica y 3.000 preguntas de verificación. Todos los modelos de prueba tenían tasas de precisión inferiores al 60%, y muchas respuestas correctas no podían ser replicadas cuando se les preguntaba repetidamente, lo que indica que los modelos eran más adivinaciones que percepciones reales. PerceptionBench tiene como objetivo diagnosticar con precisión los puntos de inconvivencia visual de la IA multimodal, impulsándola a alcanzar una comprensión visual fiel y consistente.","category":"Investigación","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI lanza PerceptionBench: un benchmark diagnóstico para las capacidades de percepción visual de modelos multimodales - Aioga Noticias de IA","description":"Moonshot AI lanzó PerceptionBench, un benchmark visual de percepción resumido a partir de más de 40 casos existentes donde los modelos han fallado realmente, incluyendo 10 capacida","url":"https://www.aioga.com/es/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:51.264Z"},"fr":{"title":"Moonshot AI publie PerceptionBench : un benchmark diagnostique pour les capacités de perception visuelle des modèles multimodaux","summary":"Moonshot AI a publié PerceptionBench, un benchmark visuel de perception résumé à partir de plus de 40 cas existants où les modèles ont effectivement échoué, incluant 10 capacités de perception atomique et 3 000 questions de vérification. Tous les modèles de test avaient des taux de précision inférieurs à 60 %, et de nombreuses bonnes réponses ne pouvaient pas être reproduites lorsqu’on les posait à plusieurs reprises, ce qui indique que les modèles étaient plus des devinettes que des perceptions réelles. PerceptionBench vise à diagnostiquer précisément les points de rupture de perception visuelle de l’IA multimodale, en la poussant à atteindre une compréhension visuelle fidèle et cohérente.","category":"Recherche","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI publie PerceptionBench : un benchmark diagnostique pour les capacités de perception visuelle des modèles multimodaux - Aioga Actualités IA","description":"Moonshot AI a publié PerceptionBench, un benchmark visuel de perception résumé à partir de plus de 40 cas existants où les modèles ont effectivement échoué, incluant 10 capacités d","url":"https://www.aioga.com/fr/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:51.353Z"},"de":{"title":"Moonshot AI veröffentlicht PerceptionBench: einen diagnostischen Benchmark für die visuellen Wahrnehmungsfähigkeiten multimodaler Modelle","summary":"Moonshot AI veröffentlichte PerceptionBench, einen visuellen Wahrnehmungs-Benchmark, der aus über 40 bestehenden Benchmark-Fällen zusammengefasst ist, in denen Modelle tatsächlich versagt haben, darunter 10 atomare Wahrnehmungsfähigkeiten und 3.000 Verifikationsfragen. Alle Testmodelle hatten Genauigkeitsraten unter 60 %, und viele richtige Antworten konnten bei wiederholter Frage nicht wiedergegeben werden, was darauf hindeutet, dass die Modelle eher raten als tatsächliche Wahrnehmung waren. PerceptionBench zielt darauf ab, die visuellen Wahrnehmungs-Breakpoints der multimodalen KI präzise zu diagnostizieren und sie zu einem treuen und konsistenten visuellen Verständnis zu führen.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI veröffentlicht PerceptionBench: einen diagnostischen Benchmark für die visuellen Wahrnehmungsfähigkeiten multimodaler Modelle - Aioga KI-News","description":"Moonshot AI veröffentlichte PerceptionBench, einen visuellen Wahrnehmungs-Benchmark, der aus über 40 bestehenden Benchmark-Fällen zusammengefasst ist, in denen Modelle tatsächlich ","url":"https://www.aioga.com/de/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.358Z"},"pt-BR":{"title":"Moonshot AI lança o PerceptionBench: um benchmark diagnóstico para capacidades de percepção visual em modelos multimodais","summary":"A Moonshot AI lançou o PerceptionBench, um benchmark de percepção visual resumido a partir de mais de 40 casos existentes em que modelos realmente falharam, incluindo 10 capacidades de percepção atômica e 3.000 perguntas de verificação. Todos os modelos de teste apresentaram taxas de precisão abaixo de 60%, e muitas respostas corretas não puderam ser reproduzidas quando perguntadas repetidamente, indicando que os modelos eram mais palpites do que percepção real. O PerceptionBench tem como objetivo diagnosticar com precisão os pontos de quebra da percepção visual da IA multimodal, impulsionando-a a alcançar uma compreensão visual fiel e consistente.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI lança o PerceptionBench: um benchmark diagnóstico para capacidades de percepção visual em modelos multimodais - Aioga Notícias de IA","description":"A Moonshot AI lançou o PerceptionBench, um benchmark de percepção visual resumido a partir de mais de 40 casos existentes em que modelos realmente falharam, incluindo 10 capacidade","url":"https://www.aioga.com/pt-BR/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.819Z"},"ru":{"title":"Moonshot AI выпускает PerceptionBench: диагностический эталонный инструмент для мультимодальных возможностей визуального восприятия","summary":"Moonshot AI выпустила PerceptionBench — бенчмарк визуального восприятия, обобщённый из более чем 40 существующих примеров эталона, где модели действительно не сработали, включая 10 атомарных возможностей восприятия и 3000 вопросов на верификацию. Все тестовые модели имели точность ниже 60%, и многие правильные ответы не могли быть воспроизведены при многократных вопросах, что указывало на то, что модели скорее предполагали, чем фактически восприятие. PerceptionBench стремится точно диагностировать точки визуального восприятия мультимодального ИИ, чтобы достичь точного и последовательного визуального понимания.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI выпускает PerceptionBench: диагностический эталонный инструмент для мультимодальных возможностей визуального восприятия - Aioga Новости ИИ","description":"Moonshot AI выпустила PerceptionBench — бенчмарк визуального восприятия, обобщённый из более чем 40 существующих примеров эталона, где модели действительно не сработали, включая 10","url":"https://www.aioga.com/ru/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:56.701Z"},"ar":{"title":"Moonshot AI تصدر PerceptionBench: معيار تشخيصي لقدرات الإدراك البصري متعدد الوسائط","summary":"أصدرت Moonshot AI مؤشر PerceptionBench، وهو معيار للإدراك البصري ملخص أكثر من 40 حالة مرجعية قائمة حيث فشلت النماذج فعليا، بما في ذلك 10 قدرات إدراك ذرية و3000 سؤال تحقق. جميع نماذج الاختبار كانت معدلات دقتها أقل من 60٪، ولم يكن بالإمكان تكرار العديد من الإجابات الصحيحة عند السؤال مرارا، مما يشير إلى أن النماذج كانت أكثر تخمينا من الإدراك الفعلي. تهدف PerceptionBench إلى تشخيص نقاط الانقطاع في الإدراك البصري للذكاء الاصطناعي متعدد الوسائط بدقة، مما يدفعه لتحقيق فهم بصري دقيق ومتسق.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI تصدر PerceptionBench: معيار تشخيصي لقدرات الإدراك البصري متعدد الوسائط - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت Moonshot AI مؤشر PerceptionBench، وهو معيار للإدراك البصري ملخص أكثر من 40 حالة مرجعية قائمة حيث فشلت النماذج فعليا، بما في ذلك 10 قدرات إدراك ذرية و3000 سؤال تحقق. جميع نماذ","url":"https://www.aioga.com/ar/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:56.688Z"},"hi":{"title":"मूनशॉट एआई ने परसेप्शन बेंच जारी किया: मल्टीमॉडल मॉडल दृश्य धारणा क्षमताओं के लिए एक नैदानिक बेंचमार्क","summary":"मूनशॉट एआई ने पर्सेप्शन बेंच जारी किया, जो 40 से अधिक मौजूदा बेंचमार्क मामलों से संक्षेप में प्रस्तुत किया गया है, जहां मॉडल वास्तव में विफल रहे हैं, जिसमें 10 परमाणु धारणा क्षमताएं और 3,000 सत्यापन प्रश्न शामिल हैं। सभी परीक्षण मॉडल में सटीकता दर 60% से कम थी, और बार-बार पूछे जाने पर कई सही उत्तरों को पुन: प्रस्तुत नहीं किया जा सकता था, यह दर्शाता है कि मॉडल वास्तविक धारणा की तुलना में अधिक अनुमान लगा रहे थे। परसेप्शनबेंच का लक्ष्य मल्टीमॉडल एआई के दृश्य धारणा ब्रेकप्वाइंट का सटीक निदान करना है, जिससे यह वफादार और सुसंगत दृश्य समझ प्राप्त कर सके।","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"मूनशॉट एआई ने परसेप्शन बेंच जारी किया: मल्टीमॉडल मॉडल दृश्य धारणा क्षमताओं के लिए एक नैदानिक बेंचमार्क - Aioga AI समाचार","description":"मूनशॉट एआई ने पर्सेप्शन बेंच जारी किया, जो 40 से अधिक मौजूदा बेंचमार्क मामलों से संक्षेप में प्रस्तुत किया गया है, जहां मॉडल वास्तव में विफल रहे हैं, जिसमें 10 परमाणु धारणा क्षमताए","url":"https://www.aioga.com/hi/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:56.433Z"},"it":{"title":"Moonshot AI rilascia PerceptionBench: un benchmark diagnostico per le capacità di percezione visiva dei modelli multimodali","summary":"Moonshot AI ha pubblicato PerceptionBench, un benchmark di percezione visiva riassunto da oltre 40 casi esistenti in cui i modelli hanno effettivamente fallito, inclusi 10 capacità di percezione atomica e 3.000 domande di verifica. Tutti i modelli di test avevano tassi di accuratezza inferiori al 60%, e molte risposte corrette non potevano essere riprodotte quando ripetutamente richieste, indicando che i modelli erano più un'ipotesi che una percezione reale. PerceptionBench mira a diagnosticare con precisione i punti di rottura della percezione visiva dell'IA multimodale, spingendola a raggiungere una comprensione visiva fedele e coerente.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI rilascia PerceptionBench: un benchmark diagnostico per le capacità di percezione visiva dei modelli multimodali - Aioga Notizie IA","description":"Moonshot AI ha pubblicato PerceptionBench, un benchmark di percezione visiva riassunto da oltre 40 casi esistenti in cui i modelli hanno effettivamente fallito, inclusi 10 capacità","url":"https://www.aioga.com/it/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.786Z"},"nl":{"title":"Moonshot AI brengt PerceptionBench uit: een diagnostische benchmark voor multimodale modelwaarnemingsmogelijkheden","summary":"Moonshot AI bracht PerceptionBench uit, een visuele waarnemingsbenchmark samengevat uit meer dan 40 bestaande benchmarkcases waarbij modellen daadwerkelijk faalden, waaronder 10 atomaire waarnemingsmogelijkheden en 3.000 verificatievragen. Alle testmodellen hadden een nauwkeurigheidsgraad onder de 60%, en veel correcte antwoorden konden niet worden herhaald wanneer ze herhaaldelijk werden gesteld, wat aangeeft dat de modellen meer aan het raden waren dan aan het daadwerkelijk perceptie. PerceptionBench streeft ernaar de breekpunten van de visuele waarneming van multimodale AI nauwkeurig te diagnosticeren, waardoor het wordt gedreven om een getrouw en consistent visueel begrip te bereiken.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI brengt PerceptionBench uit: een diagnostische benchmark voor multimodale modelwaarnemingsmogelijkheden - Aioga AI-nieuws","description":"Moonshot AI bracht PerceptionBench uit, een visuele waarnemingsbenchmark samengevat uit meer dan 40 bestaande benchmarkcases waarbij modellen daadwerkelijk faalden, waaronder 10 at","url":"https://www.aioga.com/nl/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.790Z"},"tr":{"title":"Moonshot AI, çok modlu model görsel algı yetenekleri için tanısal bir kıyaslama olan PerceptionBench'i yayınladı","summary":"Moonshot AI, modellerin gerçekten başarısız olduğu 40'tan fazla mevcut kıyaslama vakasından özet olarak 10 atomik algı yeteneği ve 3.000 doğrulama sorusu içeren bir görsel algı kıstası olan PerceptionBench'i yayınladı. Tüm test modellerinin doğruluk oranları %60'ın altındaydı ve birçok doğru cevap tekrar sorulduğunda üretilemiyordu; bu da modellerin gerçek algıdan çok tahmin yürüttüğünü gösteriyordu. PerceptionBench, çoklu modal yapay zekanın görsel algı kırılma noktalarını kesin olarak teşhis etmeyi ve sadık ve tutarlı görsel anlayışa ulaşmasını amaçlamaktadır.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI, çok modlu model görsel algı yetenekleri için tanısal bir kıyaslama olan PerceptionBench'i yayınladı - Aioga AI Haberleri","description":"Moonshot AI, modellerin gerçekten başarısız olduğu 40'tan fazla mevcut kıyaslama vakasından özet olarak 10 atomik algı yeteneği ve 3.000 doğrulama sorusu içeren bir görsel algı kıs","url":"https://www.aioga.com/tr/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.987Z"},"vi":{"title":"Moonshot AI phát hành PerceptionBench: một điểm chuẩn chẩn đoán cho khả năng nhận thức trực quan của mô hình đa phương thức","summary":"Moonshot AI đã phát hành PerceptionBench, một điểm chuẩn nhận thức trực quan được tóm tắt từ hơn 40 trường hợp điểm chuẩn hiện có mà các mô hình thực sự bị lỗi, bao gồm 10 khả năng nhận thức nguyên tử và 3.000 câu hỏi xác minh. Tất cả các mô hình thử nghiệm đều có tỷ lệ chính xác dưới 60% và nhiều câu trả lời đúng không thể được tái tạo khi được hỏi nhiều lần, cho thấy rằng các mô hình đoán nhiều hơn là nhận thức thực tế. PerceptionBench nhằm mục đích chẩn đoán chính xác các điểm ngắt nhận thức thị giác của AI đa phương thức, thúc đẩy nó đạt được sự hiểu biết trực quan trung thực và nhất quán.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI phát hành PerceptionBench: một điểm chuẩn chẩn đoán cho khả năng nhận thức trực quan của mô hình đa phương thức - Tin tức AI Aioga","description":"Moonshot AI đã phát hành PerceptionBench, một điểm chuẩn nhận thức trực quan được tóm tắt từ hơn 40 trường hợp điểm chuẩn hiện có mà các mô hình thực sự bị lỗi, bao gồm 10 khả năng","url":"https://www.aioga.com/vi/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:56.408Z"},"id":{"title":"Moonshot AI merilis PerceptionBench: tolok ukur diagnostik untuk kemampuan persepsi visual model multimoda","summary":"Moonshot AI merilis PerceptionBench, tolok ukur persepsi visual yang dirangkum dari lebih dari 40 kasus tolok ukur yang ada di mana model benar-benar gagal, termasuk 10 kemampuan persepsi atom dan 3.000 pertanyaan verifikasi. Semua model pengujian memiliki tingkat akurasi di bawah 60%, dan banyak jawaban yang benar tidak dapat direproduksi ketika ditanya berulang kali, menunjukkan bahwa model lebih menebak daripada persepsi aktual. PerceptionBench bertujuan untuk mendiagnosis dengan tepat titik putus persepsi visual dari AI multimodal, mendorongnya untuk mencapai pemahaman visual yang setia dan konsisten.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI merilis PerceptionBench: tolok ukur diagnostik untuk kemampuan persepsi visual model multimoda - Berita AI Aioga","description":"Moonshot AI merilis PerceptionBench, tolok ukur persepsi visual yang dirangkum dari lebih dari 40 kasus tolok ukur yang ada di mana model benar-benar gagal, termasuk 10 kemampuan p","url":"https://www.aioga.com/id/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.444Z"},"th":{"title":"Moonshot AI เปิดตัว PerceptionBench: เกณฑ์มาตรฐานการวินิจฉัยสําหรับความสามารถในการรับรู้ภาพของโมเดลหลายรูปแบบ","summary":"Moonshot AI เปิดตัว PerceptionBench ซึ่งเป็นเกณฑ์มาตรฐานการรับรู้ภาพที่สรุปจากกรณีเกณฑ์มาตรฐานที่มีอยู่กว่า 40 กรณีที่โมเดลล้มเหลวจริง ๆ รวมถึงความสามารถในการรับรู้อะตอม 10 รายการและคําถามตรวจสอบ 3,000 ข้อ แบบจําลองการทดสอบทั้งหมดมีอัตราความแม่นยําต่ํากว่า 60% และไม่สามารถทําซ้ําคําตอบที่ถูกต้องจํานวนมากได้เมื่อถูกถามซ้ําๆ ซึ่งบ่งชี้ว่าแบบจําลองนั้นคาดเดาได้มากกว่าการรับรู้จริง PerceptionBench มีจุดมุ่งหมายเพื่อวินิจฉัยจุดพักการรับรู้ภาพของ AI หลายรูปแบบอย่างแม่นยํา ขับเคลื่อนให้บรรลุความเข้าใจทางสายตาที่ซื่อสัตย์และสม่ําเสมอ","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI เปิดตัว PerceptionBench: เกณฑ์มาตรฐานการวินิจฉัยสําหรับความสามารถในการรับรู้ภาพของโมเดลหลายรูปแบบ - ข่าว AI Aioga","description":"Moonshot AI เปิดตัว PerceptionBench ซึ่งเป็นเกณฑ์มาตรฐานการรับรู้ภาพที่สรุปจากกรณีเกณฑ์มาตรฐานที่มีอยู่กว่า 40 กรณีที่โมเดลล้มเหลวจริง ๆ รวมถึงความสามารถในการรับรู้อะตอม 10 รายการแ","url":"https://www.aioga.com/th/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:55.962Z"},"pl":{"title":"Moonshot AI udostępnia PerceptionBench: diagnostyczny benchmark dla możliwości percepcji wizualnej modeli multimodalnych","summary":"Moonshot AI wypuściło PerceptionBench, benchmark percepcji wizualnej podsumowany z ponad 40 istniejących przypadków, w których modele faktycznie zawiodły, w tym 10 możliwości percepcji atomowej i 3 000 pytań weryfikacyjnych. Wszystkie modele testowe miały wskaźniki dokładności poniżej 60%, a wiele poprawnych odpowiedzi nie można było powtórzyć wielokrotnie, co wskazuje, że modele bardziej zgadywały niż faktycznie postrzegały ich. PerceptionBench dąży do precyzyjnego zdiagnozowania punktów krytycznych percepcji wzrokowej multimodalnej AI, napędzając ją do wiernego i spójnego wizualnego zrozumienia.","category":"论文研究","source":"Moonshot AI：Kimi Blog","pageTitle":"Moonshot AI udostępnia PerceptionBench: diagnostyczny benchmark dla możliwości percepcji wizualnej modeli multimodalnych - Aioga Wiadomości AI","description":"Moonshot AI wypuściło PerceptionBench, benchmark percepcji wizualnej podsumowany z ponad 40 istniejących przypadków, w których modele faktycznie zawiodły, w tym 10 możliwości perce","url":"https://www.aioga.com/pl/news/cmrnvwztt01bebixyznbje716/","contentTranslated":true,"sourceHash":"b426968d438918ad","translatedAt":"2026-07-19T15:57:57.047Z"}}}}