{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T06:20:51.496Z","headline":"Anthropic Claude Opus 5 在 ARC-AGI-3 基准上以 30.2% 得分大幅领先 GPT-5.6 Sol，展现全新推理行为","description":"Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。","url":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/","mainEntityOfPage":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/","datePublished":"2026-07-26T09:43:02.000Z","dateModified":"2026-07-26T09:43:02.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence","https://aihot.virxact.com/items/cms1mf9bq00qxro05gwqhi1m4"],"canonicalUrl":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/","dateCreated":"2026-07-26T09:43:02.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence","datePublished":"2026-07-26T09:43:02.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cms1mf9bq00qxro05gwqhi1m4","datePublished":"2026-07-26T09:43:02.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cms1mf9bq00qxro05gwqhi1m4"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence"},"article":{"id":"cms1mf9bq00qxro05gwqhi1m4","slug":"cms1mf9bq00qxro05gwqhi1m4","url":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/","title":"Anthropic Claude Opus 5 在 ARC-AGI-3 基准上以 30.2% 得分大幅领先 GPT-5.6 Sol，展现全新推理行为","title_en":"Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence","summary":"Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence","aiHotUrl":"https://aihot.virxact.com/items/cms1mf9bq00qxro05gwqhi1m4","publishedAt":"2026-07-26T09:43:02.000Z","category":"技巧观点","score":55,"selected":false,"articleBody":["The creators of the ARC-AGI benchmark say Anthropic's Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning.","The model scored 30.2 percent on ARC-AGI-3, making it the new leader. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level. That also puts it ahead of Anthropic's \"Fable-class\" models, which hit around 20 percent according to ARC Prize.","ARC Prize's analysis：https://x.com/arcprize/status/2080716561539907928 credits the lead to stronger logical reasoning, \"which enables more autonomous exploration, planning, and execution across unfamiliar environments.\" During testing, Opus 5 also showed behavior that researchers hadn't seen from a model before. It translated tasks into algebraic notation and independently formulated reflection equations for the first time. Ad","Six of the 25 public demo environments have now been solved. The full results：https://arcprize.org/results/anthropic-claude-opus-5, replays：https://arcprize.org/scorecards/model/anthropic-claude-opus-5-high, and benchmarking code：https://github.com/arcprize/arc-agi-3-benchmarking are publicly available. On the older ARC-AGI-2 benchmark：https://arcprize.org/leaderboard, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Both results match previous top scores, though at slightly higher costs, according to ARC Prize. Ad DEC_D_Incontent-1","ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training, including ones humans can usually handle with ease. The current version works like a game. The model must infer the rules of an interactive environment, plan its actions, and carry them out step by step. This tests general reasoning rather than stored knowledge.","Some AI systems may have：https://news.ycombinator.com/item?id=48935905 already passed the benchmark, but they rely on extra software known as a harness：https://the-decoder.com/frontier-radar-1-from-chatbots-to-problem-solvers-the-state-of-ai-agents-in-2026/. Official scores count only the language model's own performance. ARC Prize argues that future AGI systems shouldn't need outside help to solve new tasks. Opus 5 would likely score even higher if used within Claude Code. Ad","Anthropic hasn't explained the gain, but targeted data labeling and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after ARC-AGI-3 and its format became public. That may have let Anthropic target the benchmark's skills and puzzle formats, though it doesn't show the company trained on the exact tasks. Annotators could have labeled reasoning traces, useful actions, failed attempts, and recovery steps from similar puzzles. Reinforcement learning could then reward exploration, planning, rule discovery, and self-correction.","Tests on Witness：https://x.com/quietnning/status/2080786711861407883, Guanghan Ning's private benchmark for interactive puzzle games, point to narrower gains. Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3. It identified a conventional puzzle's hidden rules before taking any action but trailed Opus 4.8 on a game with less familiar mechanics. Ning says that pattern fits training on genre-specific data, though Witness can't identify what data Anthropic used. Ad DEC_D_Incontent-2","Greg Kamradt, one of the researchers behind ARC-AGI-3：https://x.com/GregKamradt/status/2081031602596200614, said the results don't rule out broader reasoning gains. A game based on familiar mechanics doesn't test adaptation to novelty, while one weak result doesn't outweigh the model's overall improvement without detailed scores for that task. Witness was also designed around ARC-AGI-3-style puzzles, so better performance could reflect real transfer rather than memorization. Ad","Ning later clarified：https://x.com/quietnning/status/2081073990614061084 that Opus 5 did generalize to Witness, just far less than on ARC-AGI-3. He compared the process to the evolution of coding benchmarks. As a major target for interactive reasoning, ARC-AGI-3 will likely attract the most training effort first. Covering more edge cases could then help models generalize to a wider range of abstract reasoning tasks. Ning said coding followed a similar path, moving from saturated benchmarks such as HumanEval to frequently updated competitions and today's coding agents.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/arc_agi_3_opus_5.png","alt":"Image description","afterParagraph":0,"url":"/media/articles/cms1mf9bq00qxro05gwqhi1m4/1765568052a4e627.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/opus_5_arc-agi-3.jpg","alt":"","afterParagraph":2,"url":"/media/articles/cms1mf9bq00qxro05gwqhi1m4/fdfda7c283f978a6.jpg"}],"mediaStatus":"ok","articleBodyZh":["ARC-AGI 基准的创建者表示，Anthropic 的 Claude Opus 5 在 ARC-AGI-3 上取得巨大领先是由于其真正更强的推理能力。","该模型在 ARC-AGI-3 上得分为 30.2%，成为新的领先者。前一纪录为 7.8%，由 OpenAI 的 GPT-5.6 Sol（Max）创造。Opus 5 解决了五个之前未解决的环境，其中四个达到或超过了人类水平。这也使它领先于 Anthropic 的“Fable 级”模型，据 ARC Prize 数据，这些模型得分约为 20%。","ARC Prize 的分析：https://x.com/arcprize/status/2080716561539907928 将其领先归因于更强的逻辑推理能力，“这使得在不熟悉的环境中能够进行更自主的探索、规划和执行。”在测试过程中，Opus 5 还显示了研究人员此前从模型中未见过的行为。它首次将任务转换为代数符号表示，并独立制定了反射方程。Ad","目前 25 个公开演示环境中已有六个被解决。完整结果：https://arcprize.org/results/anthropic-claude-opus-5，回放：https://arcprize.org/scorecards/model/anthropic-claude-opus-5-high，以及基准测试代码：https://github.com/arcprize/arc-agi-3-benchmarking 均已公开。在旧的 ARC-AGI-2 基准测试中：https://arcprize.org/leaderboard，Opus 5 得分为 90.4%，在 ARC-AGI-1 上达 97.5%。根据 ARC Prize 的说法，这两个结果与之前的最高得分相匹配，但成本略高。Ad DEC_D_Incontent-1","ARC-AGI-3 测量 AI 模型解决训练中未遇到的新任务的能力，包括人类通常可以轻松处理的任务。当前版本像玩游戏一样运行。模型必须推断交互环境的规则，规划行动，并逐步执行。这测试的是通用推理能力，而不是储存的知识。","一些 AI 系统可能已经通过了该基准测试：https://news.ycombinator.com/item?id=48935905，但它们依赖于一种称为 harness 的额外软件：https://the-decoder.com/frontier-radar-1-from-chatbots-to-problem-solvers-the-state-of-ai-agents-in-2026/。官方成绩只计算语言模型自身的表现。ARC Prize 认为，未来的 AGI 系统在解决新任务时不应依赖外部帮助。如果在 Claude Code 中使用，Opus 5 的得分可能会更高。Ad","Anthropic尚未解释其优势，但有针对性的数据标注和强化学习是合理的因素。与早期模型不同，Opus 5是在ARC-AGI-3之后开发的，其格式已经公开。这可能让Anthropic能够针对基准测试的技能和谜题格式，尽管这并不表明公司在训练中使用了完全相同的任务。标注者可能已经标注了推理轨迹、有用的操作、失败尝试和类似谜题的恢复步骤。然后，强化学习可以奖励探索、计划、规则发现和自我纠正。","在Witness上的测试：https://x.com/quietnning/status/2080786711861407883，即Guanghan Ning为互动谜题游戏制作的私人基准，显示出更有限的提升。Opus 5得分43.4，统计上与Kimi K3和Fable 5持平，而相比ARC-AGI-3的提升，它相对于Opus 4.8的改进要小得多。它在采取任何行动之前就识别出了传统谜题的隐藏规则，但在机制不太熟悉的游戏中落后于Opus 4.8。Ning表示，这种模式符合针对特定类型数据的训练，尽管Witness无法识别Anthropic使用了哪些数据。Ad DEC_D_Incontent-2","ARC-AGI-3项目的研究员之一Greg Kamradt：https://x.com/GregKamradt/status/2081031602596200614 表示，这些结果并不排除更广泛的推理提升。一款基于熟悉机制的游戏无法测试对新奇事物的适应能力，而单一的低表现结果并不能抵消模型在其他任务上的整体进步，尤其是在没有具体分数数据的情况下。Witness的设计也基于ARC-AGI-3风格的谜题，因此更好的表现可能反映了真实转移，而不是记忆化。Ad","Ning随后澄清：https://x.com/quietnning/status/2081073990614061084，Opus 5确实能够在Witness上泛化，只是远不如在ARC-AGI-3上的表现。他将这一过程比作编码基准的发展过程。作为互动推理的主要目标，ARC-AGI-3可能首先吸引最多的训练投入。覆盖更多边缘情况可以帮助模型泛化到更广泛的抽象推理任务。Ning表示，编码也遵循了类似路径，从饱和基准如HumanEval发展到频繁更新的竞赛和现今的编码代理。","保持对AI的关注。清晰、有用、无废话。","关注The Decoder，获取AI新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：实践类内容的价值在于是否能被复现、是否有明确边界，以及它能否转化为稳定的开发或工作流方法。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察示例是否可复现、工具版本变化、社区反馈和实际成本。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-28T06:29:10.488Z","sourceHash":"c994f6890d7eafb9","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["技巧观点","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"Anthropic Claude Opus 5 在 ARC-AGI-3 基准上以 30.2% 得分大幅领先 GPT-5.6 Sol，展现全新推理行为","summary":"Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。","category":"技巧观点","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 在 ARC-AGI-3 基准上以 30.2% 得分大幅领先 GPT-5.6 Sol，展现全新推理行为 - Aioga AI资讯","description":"Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol （Max） 创下的 7.8% 纪录的近四倍。","url":"https://www.aioga.com/news/cms1mf9bq00qxro05gwqhi1m4/"},"en":{"title":"Anthropic Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, significantly ahead of GPT-5.6 Sol, demonstrating new reasoning behavior","summary":"Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8% set by OpenAI GPT-5.6 Sol (Max).","category":"Insights","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, significantly ahead of GPT-5.6 Sol, demonstrating new reasoning behavior - Aioga AI News","description":"Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8% set by OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/en/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:43:02.278Z"},"ja":{"title":"Anthropic Claude Opus 5 は ARC-AGI-3 ベンチマークで 30.2% のスコアを獲得し、GPT-5.6 Sol を大幅にリードし、新たな推論行動を示した","summary":"Anthropic の Claude Opus 5 は ARC-AGI-3 ベンチマークで 30.2% のスコアを達成し、これまでの OpenAI GPT-5.6 Sol（Max）が記録した 7.8% をほぼ4倍上回った。","category":"ヒントと視点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 は ARC-AGI-3 ベンチマークで 30.2% のスコアを獲得し、GPT-5.6 Sol を大幅にリードし、新たな推論行動を示した - Aioga AIニュース","description":"Anthropic の Claude Opus 5 は ARC-AGI-3 ベンチマークで 30.2% のスコアを達成し、これまでの OpenAI GPT-5.6 Sol（Max）が記録した 7.8% をほぼ4倍上回った。","url":"https://www.aioga.com/ja/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:43:21.851Z"},"ko":{"title":"Anthropic Claude Opus 5는 ARC-AGI-3 벤치마크에서 30.2%를 기록하며 GPT-5.6 Sol을 크게 앞서며 새로운 추론 행동을 보여주었다","summary":"Anthropic의 Claude Opus 5는 ARC-AGI-3 벤치마크에서 30.2%의 점수를 기록했으며, 이전에 OpenAI GPT-5.6 Sol (Max)이 세운 7.8% 기록의 거의 네 배에 달했습니다.","category":"인사이트","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5는 ARC-AGI-3 벤치마크에서 30.2%를 기록하며 GPT-5.6 Sol을 크게 앞서며 새로운 추론 행동을 보여주었다 - Aioga AI 뉴스","description":"Anthropic의 Claude Opus 5는 ARC-AGI-3 벤치마크에서 30.2%의 점수를 기록했으며, 이전에 OpenAI GPT-5.6 Sol (Max)이 세운 7.8% 기록의 거의 네 배에 달했습니다.","url":"https://www.aioga.com/ko/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:44:06.866Z"},"es":{"title":"Anthropic Claude Opus 5 obtuvo un 30,2 % en la referencia ARC-AGI-3, superando ampliamente a GPT-5.6 Sol, mostrando un nuevo comportamiento de razonamiento","summary":"Claude Opus 5 de Anthropic obtuvo un puntaje de 30.2% en el estándar ARC-AGI-3, casi cuatro veces el récord anterior de 7.8% establecido por OpenAI GPT-5.6 Sol (Max).","category":"Ideas","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 obtuvo un 30,2 % en la referencia ARC-AGI-3, superando ampliamente a GPT-5.6 Sol, mostrando un nuevo comportamiento de razonamiento - Aioga Noticias de IA","description":"Claude Opus 5 de Anthropic obtuvo un puntaje de 30.2% en el estándar ARC-AGI-3, casi cuatro veces el récord anterior de 7.8% establecido por OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/es/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:44:03.859Z"},"fr":{"title":"Anthropic Claude Opus 5 a obtenu un score de 30,2 % sur le benchmark ARC-AGI-3, devançant largement GPT-5.6 Sol et démontrant un nouveau comportement de raisonnement","summary":"Claude Opus 5 d'Anthropic a obtenu un score de 30,2 % sur le benchmark ARC-AGI-3, près de quatre fois le record précédent de 7,8 % établi par OpenAI GPT-5.6 Sol (Max).","category":"Analyses","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 a obtenu un score de 30,2 % sur le benchmark ARC-AGI-3, devançant largement GPT-5.6 Sol et démontrant un nouveau comportement de raisonnement - Aioga Actualités IA","description":"Claude Opus 5 d'Anthropic a obtenu un score de 30,2 % sur le benchmark ARC-AGI-3, près de quatre fois le record précédent de 7,8 % établi par OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/fr/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:44:48.591Z"},"de":{"title":"Anthropic Claude Opus 5 erzielte auf dem ARC-AGI-3-Benchmark 30,2 % und liegt damit deutlich vor GPT-5.6 Sol, was ein neuartiges Schlussfolgerungsverhalten zeigt","summary":"Claude Opus 5 von Anthropic erzielte im ARC-AGI-3 Benchmark 30,2 %, fast das Vierfache des bisherigen Rekords von 7,8 %, der zuvor von OpenAI GPT-5.6 Sol (Max) aufgestellt wurde.","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 erzielte auf dem ARC-AGI-3-Benchmark 30,2 % und liegt damit deutlich vor GPT-5.6 Sol, was ein neuartiges Schlussfolgerungsverhalten zeigt - Aioga KI-News","description":"Claude Opus 5 von Anthropic erzielte im ARC-AGI-3 Benchmark 30,2 %, fast das Vierfache des bisherigen Rekords von 7,8 %, der zuvor von OpenAI GPT-5.6 Sol (Max) aufgestellt wurde.","url":"https://www.aioga.com/de/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:44:48.632Z"},"pt-BR":{"title":"Anthropic Claude Opus 5 obteve 30,2% na referência ARC-AGI-3, superando significativamente o GPT-5.6 Sol, demonstrando um novo comportamento de raciocínio","summary":"Claude Opus 5 da Anthropic alcançou uma pontuação de 30,2% no benchmark ARC-AGI-3, quase quatro vezes o recorde anterior de 7,8% estabelecido pelo OpenAI GPT-5.6 Sol (Max).","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 obteve 30,2% na referência ARC-AGI-3, superando significativamente o GPT-5.6 Sol, demonstrando um novo comportamento de raciocínio - Aioga Notícias de IA","description":"Claude Opus 5 da Anthropic alcançou uma pontuação de 30,2% no benchmark ARC-AGI-3, quase quatro vezes o recorde anterior de 7,8% estabelecido pelo OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/pt-BR/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:45:30.944Z"},"ru":{"title":"Anthropic Claude Opus 5 на бенчмарке ARC-AGI-3 набрал 30,2%, значительно опередив GPT-5.6 Sol, демонстрируя новые рассуждательные способности","summary":"Claude Opus 5 от Anthropic набрал 30,2% по эталонному тесту ARC-AGI-3, что почти в четыре раза превышает предыдущий рекорд 7,8%, установленный OpenAI GPT-5.6 Sol (Max).","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 на бенчмарке ARC-AGI-3 набрал 30,2%, значительно опередив GPT-5.6 Sol, демонстрируя новые рассуждательные способности - Aioga Новости ИИ","description":"Claude Opus 5 от Anthropic набрал 30,2% по эталонному тесту ARC-AGI-3, что почти в четыре раза превышает предыдущий рекорд 7,8%, установленный OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/ru/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:45:30.735Z"},"ar":{"title":"حقق Anthropic Claude Opus 5 على معيار ARC-AGI-3 نسبة 30.2٪ متفوقًا بشكل كبير على GPT-5.6 Sol، مظهرًا سلوك استدلالي جديدًا","summary":"حقق Claude Opus 5 من Anthropic نسبة 30.2٪ في معيار ARC-AGI-3، وهو ما يقارب أربعة أضعاف الرقم القياسي السابق البالغ 7.8٪ الذي سجّلته OpenAI GPT-5.6 Sol (Max).","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"حقق Anthropic Claude Opus 5 على معيار ARC-AGI-3 نسبة 30.2٪ متفوقًا بشكل كبير على GPT-5.6 Sol، مظهرًا سلوك استدلالي جديدًا - Aioga أخبار الذكاء الاصطناعي","description":"حقق Claude Opus 5 من Anthropic نسبة 30.2٪ في معيار ARC-AGI-3، وهو ما يقارب أربعة أضعاف الرقم القياسي السابق البالغ 7.8٪ الذي سجّلته OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/ar/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:46:16.542Z"},"hi":{"title":"Anthropic Claude Opus 5 ने ARC-AGI-3 मानक पर 30.2% अंक के साथ GPT-5.6 Sol को पीछे छोड़ते हुए नई तर्कशीलता प्रदर्शित की","summary":"Anthropic का Claude Opus 5 ने ARC-AGI-3 बेंचमार्क पर 30.2% अंक हासिल किए, जो पहले OpenAI GPT-5.6 Sol (Max) द्वारा बनाए गए 7.8% के रिकॉर्ड के लगभग चार गुना हैं।","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 ने ARC-AGI-3 मानक पर 30.2% अंक के साथ GPT-5.6 Sol को पीछे छोड़ते हुए नई तर्कशीलता प्रदर्शित की - Aioga AI समाचार","description":"Anthropic का Claude Opus 5 ने ARC-AGI-3 बेंचमार्क पर 30.2% अंक हासिल किए, जो पहले OpenAI GPT-5.6 Sol (Max) द्वारा बनाए गए 7.8% के रिकॉर्ड के लगभग चार गुना हैं।","url":"https://www.aioga.com/hi/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:46:20.294Z"},"it":{"title":"Anthropic Claude Opus 5 ha ottenuto un punteggio del 30,2% sul benchmark ARC-AGI-3, superando di gran lunga GPT-5.6 Sol, mostrando un nuovo comportamento di ragionamento","summary":"Claude Opus 5 di Anthropic ha ottenuto un punteggio del 30,2% nel benchmark ARC-AGI-3, quasi quattro volte il record del 7,8% stabilito in precedenza da OpenAI GPT-5.6 Sol (Max).","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 ha ottenuto un punteggio del 30,2% sul benchmark ARC-AGI-3, superando di gran lunga GPT-5.6 Sol, mostrando un nuovo comportamento di ragionamento - Aioga Notizie IA","description":"Claude Opus 5 di Anthropic ha ottenuto un punteggio del 30,2% nel benchmark ARC-AGI-3, quasi quattro volte il record del 7,8% stabilito in precedenza da OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/it/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:47:02.860Z"},"nl":{"title":"Anthropic Claude Opus 5 behaalde een score van 30,2% op de ARC-AGI-3 benchmark en liep hiermee ver voor op GPT-5.6 Sol, waarmee het een geheel nieuw redeneergedrag toont","summary":"Claude Opus 5 van Anthropic behaalde een score van 30,2% op de ARC-AGI-3 benchmark, bijna vier keer de eerder door OpenAI GPT-5.6 Sol (Max) gevestigde record van 7,8%.","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 behaalde een score van 30,2% op de ARC-AGI-3 benchmark en liep hiermee ver voor op GPT-5.6 Sol, waarmee het een geheel nieuw redeneergedrag toont - Aioga AI-nieuws","description":"Claude Opus 5 van Anthropic behaalde een score van 30,2% op de ARC-AGI-3 benchmark, bijna vier keer de eerder door OpenAI GPT-5.6 Sol (Max) gevestigde record van 7,8%.","url":"https://www.aioga.com/nl/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:47:02.108Z"},"tr":{"title":"Anthropic Claude Opus 5, ARC-AGI-3 kıyaslamasında %30,2 puanla GPT-5.6 Sol'u büyük bir farkla geride bırakarak yeni bir akıl yürütme davranışı sergiledi","summary":"Anthropic'in Claude Opus 5'i ARC-AGI-3 kriterinde %30,2 puan aldı ve bu, daha önce OpenAI GPT-5.6 Sol (Max) tarafından elde edilmiş %7,8 rekorunun neredeyse dört katı.","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5, ARC-AGI-3 kıyaslamasında %30,2 puanla GPT-5.6 Sol'u büyük bir farkla geride bırakarak yeni bir akıl yürütme davranışı sergiledi - Aioga AI Haberleri","description":"Anthropic'in Claude Opus 5'i ARC-AGI-3 kriterinde %30,2 puan aldı ve bu, daha önce OpenAI GPT-5.6 Sol (Max) tarafından elde edilmiş %7,8 rekorunun neredeyse dört katı.","url":"https://www.aioga.com/tr/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:47:47.529Z"},"vi":{"title":"Anthropic Claude Opus 5 trên chuẩn ARC-AGI-3 đạt điểm 30,2%, dẫn đầu xa GPT-5.6 Sol, thể hiện hành vi lý luận hoàn toàn mới","summary":"Claude Opus 5 của Anthropic đạt 30,2% trên chuẩn ARC-AGI-3, gần gấp bốn lần kỷ lục 7,8% trước đó do OpenAI GPT-5.6 Sol (Max) thiết lập.","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 trên chuẩn ARC-AGI-3 đạt điểm 30,2%, dẫn đầu xa GPT-5.6 Sol, thể hiện hành vi lý luận hoàn toàn mới - Tin tức AI Aioga","description":"Claude Opus 5 của Anthropic đạt 30,2% trên chuẩn ARC-AGI-3, gần gấp bốn lần kỷ lục 7,8% trước đó do OpenAI GPT-5.6 Sol (Max) thiết lập.","url":"https://www.aioga.com/vi/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:47:49.038Z"},"id":{"title":"Anthropic Claude Opus 5 mencatat skor 30,2% di benchmark ARC-AGI-3, jauh melampaui GPT-5.6 Sol, menunjukkan perilaku penalaran baru","summary":"Claude Opus 5 dari Anthropic mencetak skor 30,2% pada tolok ukur ARC-AGI-3, hampir empat kali lipat rekor 7,8% sebelumnya yang dibuat oleh OpenAI GPT-5.6 Sol (Max).","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 mencatat skor 30,2% di benchmark ARC-AGI-3, jauh melampaui GPT-5.6 Sol, menunjukkan perilaku penalaran baru - Berita AI Aioga","description":"Claude Opus 5 dari Anthropic mencetak skor 30,2% pada tolok ukur ARC-AGI-3, hampir empat kali lipat rekor 7,8% sebelumnya yang dibuat oleh OpenAI GPT-5.6 Sol (Max).","url":"https://www.aioga.com/id/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:48:33.110Z"},"th":{"title":"Anthropic Claude Opus 5 ทำคะแนน 30.2% ในเกณฑ์มาตรฐาน ARC-AGI-3 นำ GPT-5.6 Sol ไปอย่างมาก แสดงพฤติกรรมการให้เหตุผลรูปแบบใหม่","summary":"Claude Opus 5 ของ Anthropic ทำคะแนนได้ 30.2% ในมาตรฐาน ARC-AGI-3 ซึ่งเกือบสี่เท่าของสถิติ 7.8% ที่ OpenAI GPT-5.6 Sol (Max) เคยสร้างไว้ก่อนหน้านี้","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 ทำคะแนน 30.2% ในเกณฑ์มาตรฐาน ARC-AGI-3 นำ GPT-5.6 Sol ไปอย่างมาก แสดงพฤติกรรมการให้เหตุผลรูปแบบใหม่ - ข่าว AI Aioga","description":"Claude Opus 5 ของ Anthropic ทำคะแนนได้ 30.2% ในมาตรฐาน ARC-AGI-3 ซึ่งเกือบสี่เท่าของสถิติ 7.8% ที่ OpenAI GPT-5.6 Sol (Max) เคยสร้างไว้ก่อนหน้านี้","url":"https://www.aioga.com/th/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:48:40.451Z"},"pl":{"title":"Anthropic Claude Opus 5 uzyskał wynik 30,2% w benchmarku ARC-AGI-3, znacznie wyprzedzając GPT-5.6 Sol, pokazując całkowicie nowe zachowania wnioskowania","summary":"Claude Opus 5 firmy Anthropic uzyskał wynik 30,2% w benchmarku ARC-AGI-3, prawie czterokrotnie przewyższając poprzedni rekord OpenAI GPT-5.6 Sol (Max) wynoszący 7,8%.","category":"技巧观点","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Anthropic Claude Opus 5 uzyskał wynik 30,2% w benchmarku ARC-AGI-3, znacznie wyprzedzając GPT-5.6 Sol, pokazując całkowicie nowe zachowania wnioskowania - Aioga Wiadomości AI","description":"Claude Opus 5 firmy Anthropic uzyskał wynik 30,2% w benchmarku ARC-AGI-3, prawie czterokrotnie przewyższając poprzedni rekord OpenAI GPT-5.6 Sol (Max) wynoszący 7,8%.","url":"https://www.aioga.com/pl/news/cms1mf9bq00qxro05gwqhi1m4/","contentTranslated":true,"sourceHash":"a2c92ea278654cbb","translatedAt":"2026-07-26T15:49:26.819Z"}}}}