Kimi.ai 发布 PerceptionBench,一个从当前前沿模型在 42 个基准上的失败模式中归纳出的视觉感知基准。
该基准将视觉感知拆解为 10 种原子能力,并构建了 3000 道验证题,每道题只考察单一感知能力,无需推理或外部知识。
Kimi.ai released PerceptionBench, a visual perception benchmark derived from the failure patterns of current state-of-the-art models across 42 benchmarks....
Kimi.ai released PerceptionBench, a visual perception benchmark derived from the failure patterns of
current state-of-the-art models across 42 benchmarks. This benchmark breaks down visual perception into 10 atomic abilities and has created 3,000 validation questions, each testing a single perception ability, requiring no reasoning or external knowledge.
Kimi.ai 发布 PerceptionBench,一个从当前前沿模型在 42 个基准上的失败模式中归纳出的视觉感知基准。
该基准将视觉感知拆解为 10 种原子能力,并构建了 3000 道验证题,每道题只考察单一感知能力,无需推理或外部知识。
Kimi.ai 发布视觉感知基准 PerceptionBench。该基准源自对当前前沿模型在 42 个基准上失败模式的归纳,包含 3000 道验证题。
PerceptionBench 将视觉感知拆解为 10 种原子能力。按照发布材料,每道验证题仅考察一种感知能力,不要求模型进行推理,也不依赖外部知识。
Aioga 判断,这种单项能力拆解方式有助于更清晰地定位模型的视觉感知薄弱环节,并减少推理能力或外部知识对测试结果的干扰。
该基准可能为视觉模型评测提供更细粒度的观察框架。值得关注的是,其题目设计能否稳定区分不同模型在各类原子感知能力上的表现。 下一步值得关注 PerceptionBench 的具体题型、评测方法与使用方式,以及后续是否披露不同前沿模型在 10 种原子能力上的测试结果。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: x.com
Source: X:Kimi.ai (@Kimi_Moonshot)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: social-summary · Updated: 2026-07-27T18:45:20.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.