{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"OpenAI GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA 并超越人类动作效率基线","description":"OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上，Standard harness 得分 62.7%（成本 $26K），Provider Adapter harness 得分 99.9%（成本 $19K），均为 SOTA。","url":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","mainEntityOfPage":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","datePublished":"2026-09-04T00:07:48.000Z","dateModified":"2026-09-04T00:07:48.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://arcprize.org/blog/astra","https://aihot.news/items/cmtm7yl5s01d5robnsn25onbz"],"canonicalUrl":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","directAnswer":{"@type":"Answer","text":"来源材料称，OpenAI GPT-6 Astra 在 ARC-AGI-3 Semi-Private 的 Standard harness 得分为62.7%，成本为2.6万美元；在 Provider Adapter harness 得分为99.9%，成本为1.9万美元，均被列为当前最高成绩。","url":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","dateCreated":"2026-09-04T00:07:48.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"arcprize.org source article","url":"https://arcprize.org/blog/astra","datePublished":"2026-09-04T00:07:48.000Z","provider":{"@type":"Organization","name":"arcprize.org","url":"https://arcprize.org/blog/astra"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtm7yl5s01d5robnsn25onbz","datePublished":"2026-09-04T00:07:48.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtm7yl5s01d5robnsn25onbz"}}],"aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","originalPublisher":{"name":"arcprize.org","url":"https://arcprize.org/blog/astra"},"geoDeepAnswer":null,"article":{"id":"cmtm7yl5s01d5robnsn25onbz","slug":"cmtm7yl5s01d5robnsn25onbz","url":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","title":"OpenAI GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA 并超越人类动作效率基线","title_en":"","summary":"OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上，Standard harness 得分 62.7%（成本 $26K），Provider Adapter harness 得分 99.9%（成本 $19K），均为 SOTA。","source":"Hacker News 热门（buzzing.cc 中文翻译）","sourceUrl":"https://arcprize.org/blog/astra","aiHotUrl":"https://aihot.news/items/cmtm7yl5s01d5robnsn25onbz","publishedAt":"2026-09-04T00:07:48.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["ARC-AGI-3 is a benchmark for studying agentic intelligence through novel, abstract, turn-based environments. Agents must explore, infer goals, and build internal models of environments to effectively plan actions without explicit instructions. You can play ARC-AGI-3 yourself：/tasks/ls20.","Your browser does not support embedded video.","These environments only contain core knowledge priors：https://arcprize.org/arc-agi#:~:text=towards%20general%20intelligence.-,Core%20Knowledge%20Priors,-A%20principle%20underlying and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments：/blog/arc-agi-3-human-dataset.","The goal of the ARC-AGI series is to measure the “residual gap” between current artificial intelligence and AGI. We define AGI as a system’s ability to acquire any skill a human can, as efficiently as a human can.","ARC-AGI-3 is the third generation of the ARC-AGI benchmark series：/arc-agi. It tests agentic capabilities beyond ARC-AGI-1：/arc-agi/1 and ARC-AGI-2：/arc-agi/2. Each generation expands on the one before it - as frontier AI capabilities advance, our benchmarks must advance with them.","ARC-AGI-3 tests four components of agentic intelligence:","With our Standard harness Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment. , OpenAI’s Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private：/results/openai-gpt-6-astra for $26K. With the Provider Adapter harness The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. , Astra (high) scores 99.9% for $19K. Both are state-of-the-art scores. See the full leaderboard：/leaderboard.","At max reasoning effort, Astra solves games more efficiently, requiring fewer actions and therefore lowering total cost relative to the other reasoning-effort levels.","For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.","Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted. 1：#fn-1","Beyond the scores, Astra’s replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds.","When playing ARC-AGI-3, Astra chooses which strategy notes it would like to carry forward. It tracked objects, coordinates, rules, and unfinished plans, while also using a custom domain-specific language notation it generated for the environments.","We’ve seen similar behavior in other models：https://x.com/arcprize/status/2080716567760007317, but Astra’s notes stood out for their precision and information density. It distilled the scene into a compact code-like symbolic model: where objects were, how they interacted, and exactly which actions needed to happen in what order. This is an on-the-fly algebraic shorthand rather than a fully fledged programming language. For example:","Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment. Participants were not selected for puzzle-solving experience or ability. 2：#fn-2","For each level, we defined the “human baseline” using the median action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs more actions is less action-efficient, while one that needs fewer actions is more action-efficient.","In the Provider Adapter harness The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. , Astra (max) used fewer actions than the human baseline on 96.0% of levels and used 51.7% fewer actions per level on average . This is a material milestone. This means by ARC-AGI-3’s measure of action efficiency, Astra matched and surpassed human parity.","As an aside, before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI “understands” the mechanics, it generally executes within the range of human efficiency.","Astra’s Action Efficiency Compared to Humans","Each dot represents one level that Astra (max) completed. Points below the solid line indicate fewer actions than the human baseline.","The plot above compares the number of actions Astra used to complete each level with our human baseline. This reinforces why ARC-AGI-3 measures action efficiency, not just task completion. A completion-only score would tell us that Astra completed an environment, but not how efficiently it learned to solve them.","Most benchmarks only measure cost efficiency, which measures the computational resources used, but action efficiency measures how much experience with an environment was required.","Astra’s results show that it needed fewer interactions than the human baseline to execute a solution.","We also evaluated Astra in the PRO-LONG harness：https://github.com/alexisfox7/PRO-LONG (paper：https://arxiv.org/pdf/2607.20064), an early ARC-AGI-3 red-teaming partner. In this advanced setup, Astra had access to a sandbox where it could execute custom code 3：#fn-3 .","We observed Astra create a custom set of tools for each game: board parsers, game-state models, search algorithms, planners, and persistent notes. For more involved runs, Astra even produced small, game-specific software libraries.","For example, in tu93 ：/tasks/tu93, a maze-like game with guards and moving patrols, Astra started with navigation and built maze_solver.py . It added combat rules in combat_solver.py , modeled moving patrols in patrol_solver.py , and used sync_state.py to check its predictions against observations.","Examining Astra’s performance in PRO-LONG is useful because we see what it can do with external tools. However, this represents different evaluation conditions from our controlled human testing. Our testing participants did not have a code interpreter, scratch pad, etc., so PRO-LONG’s results should be understood as the combined performance of the model and its tools.","Our Standard harness for ARC-AGI-3 asks how models compare under the same minimal, provider-neutral interface. It provides all the information required to solve each game, but leaves the model responsible for deciding what to preserve in its visible notes. We believe a future AGI should be able to solve ARC-AGI-3 under these conditions. The shared interface also gives us a consistent, apples-to-apples comparison across providers.","Alternatively, there is a separate question: how well does a model perform when it can use the context-management features its provider designed for it? For Astra, this means preserving the opaque reasoning state (which we don’t see) between requests and using compaction to manage longer conversations.","With the Provider Adapter harness, Astra's best observed score on ARC-AGI-3 Semi-Private increased from 62.7% to 99.9%. Looking across Public and Semi-Private and all reasoning levels, Provider Adapter runs were approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens across the 167 game-reasoning pairs both harnesses solved.","Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository：https://github.com/arcprize/arc-agi-3-benchmarking and testing policy：/policy document both approaches.","ARC-AGI-3 continues to be a useful playground for researchers and agents to explore unfamiliar environments, discover rules, and learn through interaction. Astra’s results are also a major milestone worth celebrating. From our perspective, Astra represents a noticeable step-function change in frontier model capabilities.","When we launched ARC-AGI-3, we made it clear：https://arxiv.org/pdf/2603.24621 that saturating the benchmark would not represent “proof of achieving AGI.” Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.","The ARC-AGI benchmark series is designed to evolve in tandem with frontier AI. This creates a feedback loop between emerging research questions and advances in AI capabilities. ARC-AGI-3 was our first interactive benchmark, which asked AI to efficiently synthesize causal world models and achieve goals without specific instructions. Astra clears this bar. At the same time, ARC-AGI-3 has a tightly bounded scope and format, and its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.","We are actively exploring the questions that should shape the next generation of benchmarks, including how to evaluate recursive self-improvement and open-ended innovation. Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.","Thank you to François Chollet, Mike Knoop, Matt Mazur, Ethan Bond, and Derek Smith for early review of this post.","Get started and receive official contest updates and news."],"articleImages":[{"sourceUrl":"https://arcprize.org/media/images/blog-greg-kamradt.jpg","alt":"Greg Kamradt","afterParagraph":0,"url":"/media/articles/cmtm7yl5s01d5robnsn25onbz/3d2edc91b00aeace.jpg"},{"sourceUrl":"https://arcprize.org/media/images/blog/astra-arc-agi-3-leaderboard.png","alt":"ARC-AGI-3 leaderboard showing GPT-6 Astra Standard and Provider Adapter results","afterParagraph":5,"url":"/media/articles/cmtm7yl5s01d5robnsn25onbz/e5b3e5742dbb233e.png"},{"sourceUrl":"https://arcprize.org/media/images/blog/astra-symbolic-model.gif","alt":"Astra playing s5i5 while recording compact symbolic notes","afterParagraph":12,"url":"/media/articles/cmtm7yl5s01d5robnsn25onbz/7e0eb7e287fa38f8.gif"},{"sourceUrl":"https://arcprize.org/media/images/blog/astra-action-efficiency.png","alt":"Scatter plot comparing Astra actions with the human baseline for each completed ARC-AGI-3 level","afterParagraph":17,"url":"/media/articles/cmtm7yl5s01d5robnsn25onbz/99c9de3802c6b5c1.png"},{"sourceUrl":"https://arcprize.org/media/images/blog/astra-pro-long-tools.gif","alt":"Astra using a custom maze solver while playing tu93 in the PRO-LONG harness","afterParagraph":25,"url":"/media/articles/cmtm7yl5s01d5robnsn25onbz/2572109f74b7be9f.gif"}],"mediaStatus":"ok","articleBodyZh":["ARC-AGI-3 是通过新颖、抽象、回合制环境研究智能智能的基准。智能体必须探索、推断目标，并构建环境内部模型，以有效规划行动而无需明确指令。你可以自己玩 ARC-AGI-3：/tasks/ls20。","你的浏览器不支持嵌入视频。","这些环境仅包含核心知识先验：https：//arcprize.org/arc-agi#：~：text=towards%20general%20intelligence.-，Core%20Knowledge%20Priors，-A%20principle%20underlying，并通过与人类参与者的受控测试进行难度校准。人类可以解决100%的环境：/blog/arc-agi-3-human-dataset。","ARC-AGI系列的目标是衡量当前人工智能与AGI之间的“残余差距”。我们将AGI定义为系统能够像人类一样高效地习得人类所能掌握的任何技能的能力。","ARC-AGI-3 是 ARC-AGI 基准系列的第三代。它测试了超越 ARC-AGI-1：/arc-agi/1 和 ARC-AGI-2：/arc-agi/2 的代理能力。每一代都在前一代基础上不断扩展——随着前沿人工智能能力的进步，我们的基准也必须跟随进步。","ARC-AGI-3 测试代理智能的四个组成部分：","通过我们的标准束，标准束允许模型在整个环境中携带其选择保留的笔记。OpenAI的Astra（最高）在ARC-AGI-3半私密：/results/openai-gpt-6-astra中得分62.7%，价格为2.6万美元。使用Provider Adapter harness的提供者适配器束保留请求间的不透明推理状态，并使用压缩处理以进行更长时间的对话，允许模型重用之前的工作。，Astra（高）得分99.9%，价格为1.9万美元。两者均为最先进的成绩。查看完整排行榜：/leaderboard。","在推理努力达到最大水平时，Astra 解决游戏更高效，所需动作更少，因此相较于其他推理努力水平，整体成本更低。","作为成本对比，在我们的受控测试中，人类参与者每90分钟游戏获得115美元，外加每场完成游戏5美元。参与者每场尝试大约九局游戏，约为12.78美元（未计奖金）。","大部分费用用于支付参与者参加测试的时间和意愿，而非大脑消耗的能量（这是与人工智能更接近的替代指标）。如果只看大脑能量，并以电为价，估计值降至每场约0.6美分，或每场游戏0.067美分。1：#fn-1","除了评分，Astra的回放展示了它如何将陌生的游戏机制转化为有用的可行模型。有三项发现尤为突出：它开发的紧凑代数符号、与人类相比的行动效率，以及它构建的自定义工具。","在玩 ARC-AGI-3 时，Astra 选择想要继续传递的策略笔记。它跟踪了对象、坐标、规则和未完成的计划，同时使用了为环境生成的自定义领域专用语言符号。","我们在其他模型中也见过类似行为：https：//x.com/arcprize/status/2080716567760007317，但Astra的笔记因其精确度和信息密度而突出。它将场景提炼成一个紧凑的代码符号模型：物体的位置、交互方式以及具体需要以何种顺序执行哪些动作。这是一种即时代数速记，而非完整的编程语言。例如：","在启动ARC-AGI-3之前，我们测试了大约500名公众成员，以建立行动效率的基线，简单来说，就是人们解决每个环境的速度。参与者的选拔并非基于解谜经验或能力。2：#fn-2","对于每个关卡，我们根据完成该关卡的玩家的中位行动数定义了“人类基线”。这为比较人类与AI表现提供了参考。需要更多动作的AI动作效率较低，而需要动作较少的AI则更高效。","在提供者适配器框架中提供者适配器框架保留请求间的不透明推理状态，并使用压缩处理以进行更长时间的对话，使模型能够重复使用之前的工作。Astra（最大）在96.0%的层级中使用了比人类基线更少的动作，平均每层的动作减少了51.7%。这是一个重要的里程碑。这意味着根据ARC-AGI-3的动作效率衡量，Astra在达到人类水平上达到并超过了。","顺便提一下，在我们推出 ARC-AGI-3 之前，我们假设行动效率将继续作为人类与 AI 之间的分水岭。我们预计，即使 AI 解决了一个环境，它可能也需要比人类更多的探索（行动）。对于蛮力方法来说，这仍然是正确的，但前沿 AI 展现出更类似二进制的模式。一旦前沿 AI “理解”了机制，它通常可以在接近人类效率的范围内执行。","阿斯特拉的行动效率与人类的比较","每个点代表 Astra（最大化）完成的一个关卡。实线下方的点表示所需行动少于人类基线。","上图比较了 Astra 完成每个关卡所使用的行动次数与我们的人类基线。这进一步说明了为什么 ARC-AGI-3 测量的是行动效率，而不仅仅是任务完成情况。仅完成任务的评分会告诉我们 Astra 完成了一个环境，但无法反映它解决问题的学习效率。","大多数基准只衡量成本效率，即使用的计算资源，而行动效率衡量的是完成任务所需的环境经验量。","Astra 的结果显示，它执行解决方案所需的交互次数少于人类基线。","我们还在 PRO-LONG 环境中评估了 Astra：https://github.com/alexisfox7/PRO-LONG（论文：https://arxiv.org/pdf/2607.20064），这是 ARC-AGI-3 的早期红队合作伙伴。在这个高级设置中，Astra 可以访问一个沙盒执行自定义代码 3：#fn-3。","我们观察到 Astra 为每个游戏创建了一套自定义工具：棋盘解析器、游戏状态模型、搜索算法、规划器及持续笔记。在更复杂的运行中，Astra 甚至生成了小型、针对特定游戏的软件库。","例如，在 tu93：/tasks/tu93，一个带守卫和移动巡逻的迷宫类游戏中，Astra 从导航开始并构建了 maze_solver.py。它在 combat_solver.py 中添加了战斗规则，在 patrol_solver.py 中模拟移动的巡逻，并使用 sync_state.py 将其预测与观察结果进行对比。","在PRO-LONG中检查Astra的表现是有用的，因为我们可以看到它在使用外部工具时的能力。然而，这代表了与我们的受控人类测试不同的评估条件。我们的测试参与者没有使用代码解释器、草稿本等工具，因此应将PRO-LONG的结果理解为模型和其工具的综合表现。","我们针对ARC-AGI-3的标准测试方法询问模型在相同的最小化、提供者中立接口下的表现如何。它提供了解决每个游戏所需的所有信息，但将决定在可见笔记中保留什么的责任留给模型。我们认为，未来的AGI应该能够在这些条件下解决ARC-AGI-3。共享接口还使我们能够在不同提供者之间进行一致、可比较的评估。","另一种情况是：当模型可以使用其提供者为其设计的上下文管理功能时，它的表现如何？对于Astra来说，这意味着在请求之间保留不透明的推理状态（我们看不到的）并使用压缩功能管理较长的对话。","在Provider Adapter测试中，Astra在ARC-AGI-3半私有环境中的最佳观察得分从62.7%提升至99.9%。在公有和半私有以及所有推理等级中，Provider Adapter的运行在总记录耗时上大约快3.66倍，并且在两种测试方法解决的167对游戏推理组合中，总使用令牌数减少了49%。","未来，我们将在ARC-AGI排行榜上报告标准测试和Provider Adapter测试的结果，并清楚标注每种评估条件。我们的开源测试仓库：https://github.com/arcprize/arc-agi-3-benchmarking 和测试政策文件：/policy 都记录了这两种方法。","ARC-AGI-3仍然是研究人员和智能体探索陌生环境、发现规则和通过互动学习的有用平台。Astra的结果也是一个值得庆祝的重要里程碑。从我们的角度来看，Astra代表了前沿模型能力的明显跃升。","当我们发布 ARC-AGI-3 时，我们明确表示：https://arxiv.org/pdf/2603.24621，饱和基准并不代表“实现 AGI 的证据”。因此，虽然我们认为 Astra 代表了在泛化方面的有意义进展，但我们并不声称它是 AGI。","ARC-AGI 基准系列设计为与前沿 AI 同步发展。这在新兴研究问题与 AI 能力进步之间建立了反馈循环。ARC-AGI-3 是我们的第一个交互式基准，它要求 AI 高效综合因果世界模型并在没有具体指令的情况下实现目标。Astra 达到了这一标准。同时，ARC-AGI-3 的范围和格式都严格限定，其环境具有确定性、封闭式的机制和目标。它并不代表现实世界的复杂性和开放性。","我们正在积极探索应当塑造下一代基准的问题，包括如何评估递归自我提升和开放式创新。Astra 的进展有助于澄清哪些 AI 能力尚无法达到，以及哪些问题仍未解决。","感谢 François Chollet、Mike Knoop、Matt Mazur、Ethan Bond 和 Derek Smith 对本文的早期审阅。","开始使用并接收官方竞赛更新和新闻。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"来源材料称，OpenAI GPT-6 Astra 在 ARC-AGI-3 Semi-Private 的 Standard harness 得分为62.7%，成本为2.6万美元；在 Provider Adapter harness 得分为99.9%，成本为1.9万美元，均被列为当前最高成绩。","background":"ARC-AGI-3用于研究智能体在新颖、抽象、回合制环境中的探索、目标推断和行动规划能力。来源称，人类能够解决全部环境；该系列旨在衡量人工智能与通用人工智能之间的剩余差距。","viewpoint":"Aioga 判断：Astra 的结果显示，不同 harness 的状态保留与推理机制可能显著影响成绩和成本，单看最高分不足以完整反映系统表现，评估时需要同时关注测试配置、行动数量与资源消耗。","implications":"可能影响：后续模型比较应明确区分 Standard harness 与 Provider Adapter harness，并同步披露成本和行动效率；99.9%的成绩不代表模型已在所有环境中达到人类水平，也不足以单独证明已实现通用人工智能。","nextStep":"后续观察：需要关注 ARC-AGI-3 完整排行榜、不同模型在统一 harness 下的表现，以及来源是否进一步公开行动数量、成本计算方式和人类参与者测试的可比细节。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-04T01:13:10.817Z","sourceHash":"01ca73406fe3fad8","review":{"approved":true,"groundedness":96,"clarity":93,"duplicationRisk":12,"blockingIssues":[],"notes":["观点部分已明确标注为“Aioga 判断”，且使用“可能”等审慎措辞，没有将推断冒充来源事实。","比较两种 harness 时还应注意推理强度不同：Standard harness 使用 Astra (max)，Provider Adapter harness 使用 Astra (high)，因此分数和成本差异不能仅归因于 harness 机制。","“99.9%的成绩不代表模型已在所有环境中达到人类水平”是合理的限制性说明，因为该成绩针对 ARC-AGI-3 Semi-Private，且来源未证明其可外推至所有环境或通用智能。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Hacker News 热门（buzzing.cc 中文翻译）"],"translations":{"zh-CN":{"title":"OpenAI GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA 并超越人类动作效率基线","summary":"OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上，Standard harness 得分 62.7%（成本 $26K），Provider Adapter harness 得分 99.9%（成本 $19K），均为 SOTA。","category":"行业动态","source":"arcprize.org","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA 并超越人类动作效率基线 - Aioga AI资讯","description":"OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上，Standard harness 得分 62.7%（成本 $26K），Provider Adapter harness 得分 99.9%（成本 $19K），均为 SOTA。","url":"https://www.aioga.com/news/cmtm7yl5s01d5robnsn25onbz/","articleBody":["ARC-AGI-3 是通过新颖、抽象、回合制环境研究智能智能的基准。智能体必须探索、推断目标，并构建环境内部模型，以有效规划行动而无需明确指令。你可以自己玩 ARC-AGI-3：/tasks/ls20。","你的浏览器不支持嵌入视频。","这些环境仅包含核心知识先验：https：//arcprize.org/arc-agi#：~：text=towards%20general%20intelligence.-，Core%20Knowledge%20Priors，-A%20principle%20underlying，并通过与人类参与者的受控测试进行难度校准。人类可以解决100%的环境：/blog/arc-agi-3-human-dataset。","ARC-AGI系列的目标是衡量当前人工智能与AGI之间的“残余差距”。我们将AGI定义为系统能够像人类一样高效地习得人类所能掌握的任何技能的能力。","ARC-AGI-3 是 ARC-AGI 基准系列的第三代。它测试了超越 ARC-AGI-1：/arc-agi/1 和 ARC-AGI-2：/arc-agi/2 的代理能力。每一代都在前一代基础上不断扩展——随着前沿人工智能能力的进步，我们的基准也必须跟随进步。","ARC-AGI-3 测试代理智能的四个组成部分：","通过我们的标准束，标准束允许模型在整个环境中携带其选择保留的笔记。OpenAI的Astra（最高）在ARC-AGI-3半私密：/results/openai-gpt-6-astra中得分62.7%，价格为2.6万美元。使用Provider Adapter harness的提供者适配器束保留请求间的不透明推理状态，并使用压缩处理以进行更长时间的对话，允许模型重用之前的工作。，Astra（高）得分99.9%，价格为1.9万美元。两者均为最先进的成绩。查看完整排行榜：/leaderboard。","在推理努力达到最大水平时，Astra 解决游戏更高效，所需动作更少，因此相较于其他推理努力水平，整体成本更低。","作为成本对比，在我们的受控测试中，人类参与者每90分钟游戏获得115美元，外加每场完成游戏5美元。参与者每场尝试大约九局游戏，约为12.78美元（未计奖金）。","大部分费用用于支付参与者参加测试的时间和意愿，而非大脑消耗的能量（这是与人工智能更接近的替代指标）。如果只看大脑能量，并以电为价，估计值降至每场约0.6美分，或每场游戏0.067美分。1：#fn-1","除了评分，Astra的回放展示了它如何将陌生的游戏机制转化为有用的可行模型。有三项发现尤为突出：它开发的紧凑代数符号、与人类相比的行动效率，以及它构建的自定义工具。","在玩 ARC-AGI-3 时，Astra 选择想要继续传递的策略笔记。它跟踪了对象、坐标、规则和未完成的计划，同时使用了为环境生成的自定义领域专用语言符号。","我们在其他模型中也见过类似行为：https：//x.com/arcprize/status/2080716567760007317，但Astra的笔记因其精确度和信息密度而突出。它将场景提炼成一个紧凑的代码符号模型：物体的位置、交互方式以及具体需要以何种顺序执行哪些动作。这是一种即时代数速记，而非完整的编程语言。例如：","在启动ARC-AGI-3之前，我们测试了大约500名公众成员，以建立行动效率的基线，简单来说，就是人们解决每个环境的速度。参与者的选拔并非基于解谜经验或能力。2：#fn-2","对于每个关卡，我们根据完成该关卡的玩家的中位行动数定义了“人类基线”。这为比较人类与AI表现提供了参考。需要更多动作的AI动作效率较低，而需要动作较少的AI则更高效。","在提供者适配器框架中提供者适配器框架保留请求间的不透明推理状态，并使用压缩处理以进行更长时间的对话，使模型能够重复使用之前的工作。Astra（最大）在96.0%的层级中使用了比人类基线更少的动作，平均每层的动作减少了51.7%。这是一个重要的里程碑。这意味着根据ARC-AGI-3的动作效率衡量，Astra在达到人类水平上达到并超过了。","顺便提一下，在我们推出 ARC-AGI-3 之前，我们假设行动效率将继续作为人类与 AI 之间的分水岭。我们预计，即使 AI 解决了一个环境，它可能也需要比人类更多的探索（行动）。对于蛮力方法来说，这仍然是正确的，但前沿 AI 展现出更类似二进制的模式。一旦前沿 AI “理解”了机制，它通常可以在接近人类效率的范围内执行。","阿斯特拉的行动效率与人类的比较","每个点代表 Astra（最大化）完成的一个关卡。实线下方的点表示所需行动少于人类基线。","上图比较了 Astra 完成每个关卡所使用的行动次数与我们的人类基线。这进一步说明了为什么 ARC-AGI-3 测量的是行动效率，而不仅仅是任务完成情况。仅完成任务的评分会告诉我们 Astra 完成了一个环境，但无法反映它解决问题的学习效率。","大多数基准只衡量成本效率，即使用的计算资源，而行动效率衡量的是完成任务所需的环境经验量。","Astra 的结果显示，它执行解决方案所需的交互次数少于人类基线。","我们还在 PRO-LONG 环境中评估了 Astra：https://github.com/alexisfox7/PRO-LONG（论文：https://arxiv.org/pdf/2607.20064），这是 ARC-AGI-3 的早期红队合作伙伴。在这个高级设置中，Astra 可以访问一个沙盒执行自定义代码 3：#fn-3。","我们观察到 Astra 为每个游戏创建了一套自定义工具：棋盘解析器、游戏状态模型、搜索算法、规划器及持续笔记。在更复杂的运行中，Astra 甚至生成了小型、针对特定游戏的软件库。","例如，在 tu93：/tasks/tu93，一个带守卫和移动巡逻的迷宫类游戏中，Astra 从导航开始并构建了 maze_solver.py。它在 combat_solver.py 中添加了战斗规则，在 patrol_solver.py 中模拟移动的巡逻，并使用 sync_state.py 将其预测与观察结果进行对比。","在PRO-LONG中检查Astra的表现是有用的，因为我们可以看到它在使用外部工具时的能力。然而，这代表了与我们的受控人类测试不同的评估条件。我们的测试参与者没有使用代码解释器、草稿本等工具，因此应将PRO-LONG的结果理解为模型和其工具的综合表现。","我们针对ARC-AGI-3的标准测试方法询问模型在相同的最小化、提供者中立接口下的表现如何。它提供了解决每个游戏所需的所有信息，但将决定在可见笔记中保留什么的责任留给模型。我们认为，未来的AGI应该能够在这些条件下解决ARC-AGI-3。共享接口还使我们能够在不同提供者之间进行一致、可比较的评估。","另一种情况是：当模型可以使用其提供者为其设计的上下文管理功能时，它的表现如何？对于Astra来说，这意味着在请求之间保留不透明的推理状态（我们看不到的）并使用压缩功能管理较长的对话。","在Provider Adapter测试中，Astra在ARC-AGI-3半私有环境中的最佳观察得分从62.7%提升至99.9%。在公有和半私有以及所有推理等级中，Provider Adapter的运行在总记录耗时上大约快3.66倍，并且在两种测试方法解决的167对游戏推理组合中，总使用令牌数减少了49%。","未来，我们将在ARC-AGI排行榜上报告标准测试和Provider Adapter测试的结果，并清楚标注每种评估条件。我们的开源测试仓库：https://github.com/arcprize/arc-agi-3-benchmarking 和测试政策文件：/policy 都记录了这两种方法。","ARC-AGI-3仍然是研究人员和智能体探索陌生环境、发现规则和通过互动学习的有用平台。Astra的结果也是一个值得庆祝的重要里程碑。从我们的角度来看，Astra代表了前沿模型能力的明显跃升。","当我们发布 ARC-AGI-3 时，我们明确表示：https://arxiv.org/pdf/2603.24621，饱和基准并不代表“实现 AGI 的证据”。因此，虽然我们认为 Astra 代表了在泛化方面的有意义进展，但我们并不声称它是 AGI。","ARC-AGI 基准系列设计为与前沿 AI 同步发展。这在新兴研究问题与 AI 能力进步之间建立了反馈循环。ARC-AGI-3 是我们的第一个交互式基准，它要求 AI 高效综合因果世界模型并在没有具体指令的情况下实现目标。Astra 达到了这一标准。同时，ARC-AGI-3 的范围和格式都严格限定，其环境具有确定性、封闭式的机制和目标。它并不代表现实世界的复杂性和开放性。","我们正在积极探索应当塑造下一代基准的问题，包括如何评估递归自我提升和开放式创新。Astra 的进展有助于澄清哪些 AI 能力尚无法达到，以及哪些问题仍未解决。","感谢 François Chollet、Mike Knoop、Matt Mazur、Ethan Bond 和 Derek Smith 对本文的早期审阅。","开始使用并接收官方竞赛更新和新闻。"]},"en":{"title":"OpenAI GPT-6 Astra achieves SOTA on ARC-AGI-3 and surpasses the human action efficiency baseline","summary":"OpenAI's GPT-6 Astra scored 62.7% with the Standard harness (cost $26K) and 99.9% with the Provider Adapter harness (cost $19K) on the ARC-AGI-3 Semi-Private, both being state-of-the-art.","category":"Industry","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra achieves SOTA on ARC-AGI-3 and surpasses the human action efficiency baseline - Aioga AI News","description":"OpenAI's GPT-6 Astra scored 62.7% with the Standard harness (cost $26K) and 99.9% with the Provider Adapter harness (cost $19K) on the ARC-AGI-3 Semi-Private, both being state-of-t...","url":"https://www.aioga.com/en/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:04:32.780Z"},"ja":{"title":"OpenAI GPT-6 Astra は ARC-AGI-3 で SOTA を達成し、人間の行動効率の基準を上回った","summary":"OpenAI の GPT-6 Astra は ARC-AGI-3 Semi-Private 上で、Standard harness のスコアは 62.7%（コスト $26K）、Provider Adapter harness のスコアは 99.9%（コスト $19K）で、いずれも SOTA です。","category":"業界動向","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra は ARC-AGI-3 で SOTA を達成し、人間の行動効率の基準を上回った - Aioga AIニュース","description":"OpenAI の GPT-6 Astra は ARC-AGI-3 Semi-Private 上で、Standard harness のスコアは 62.7%（コスト $26K）、Provider Adapter harness のスコアは 99.9%（コスト $19K）で、いずれも SOTA です。","url":"https://www.aioga.com/ja/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:04:37.970Z"},"ko":{"title":"OpenAI GPT-6 Astra는 ARC-AGI-3에서 SOTA를 달성하고 인간 행동 효율성 기준을 능가했습니다","summary":"OpenAI의 GPT-6 Astra는 ARC-AGI-3 Semi-Private에서 Standard harness 점수 62.7%(비용 $26K), Provider Adapter harness 점수 99.9%(비용 $19K)로, 모두 SOTA입니다.","category":"업계 동향","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra는 ARC-AGI-3에서 SOTA를 달성하고 인간 행동 효율성 기준을 능가했습니다 - Aioga AI 뉴스","description":"OpenAI의 GPT-6 Astra는 ARC-AGI-3 Semi-Private에서 Standard harness 점수 62.7%(비용 $26K), Provider Adapter harness 점수 99.9%(비용 $19K)로, 모두 SOTA입니다.","url":"https://www.aioga.com/ko/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:05:19.802Z"},"es":{"title":"OpenAI GPT-6 Astra logra SOTA en ARC-AGI-3 y supera la línea base de eficiencia de acción humana","summary":"El GPT-6 Astra de OpenAI en ARC-AGI-3 Semi-Private, el arnés Standard obtuvo una puntuación de 62,7% (costo $26K), el arnés Provider Adapter obtuvo una puntuación de 99,9% (costo $19K), ambos son SOTA.","category":"Industria","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra logra SOTA en ARC-AGI-3 y supera la línea base de eficiencia de acción humana - Aioga Noticias de IA","description":"El GPT-6 Astra de OpenAI en ARC-AGI-3 Semi-Private, el arnés Standard obtuvo una puntuación de 62,7% (costo $26K), el arnés Provider Adapter obtuvo una puntuación de 99,9% (costo $...","url":"https://www.aioga.com/es/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:05:16.076Z"},"fr":{"title":"OpenAI GPT-6 Astra a atteint le SOTA sur ARC-AGI-3 et a dépassé la référence d'efficacité des actions humaines","summary":"Le GPT-6 Astra d'OpenAI a obtenu un score de 62,7 % sur le Standard harness et 99,9 % sur le Provider Adapter harness sur ARC-AGI-3 Semi-Private (coût de 26 000 $ pour le Standard harness et 19 000 $ pour le Provider Adapter harness), les deux étant à l'état de l'art.","category":"Industrie","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra a atteint le SOTA sur ARC-AGI-3 et a dépassé la référence d'efficacité des actions humaines - Aioga Actualités IA","description":"Le GPT-6 Astra d'OpenAI a obtenu un score de 62,7 % sur le Standard harness et 99,9 % sur le Provider Adapter harness sur ARC-AGI-3 Semi-Private (coût de 26 000 $ pour le Standard...","url":"https://www.aioga.com/fr/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:06:02.253Z"},"de":{"title":"OpenAI GPT-6 Astra erreicht auf ARC-AGI-3 den Stand der Technik und übertrifft die Effizienz menschlicher Aktionen","summary":"Der GPT-6 Astra von OpenAI erzielte auf ARC-AGI-3 Semi-Private mit dem Standard-Harness 62,7 % (Kosten 26.000 $) und mit dem Provider-Adapter-Harness 99,9 % (Kosten 19.000 $), beide sind SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra erreicht auf ARC-AGI-3 den Stand der Technik und übertrifft die Effizienz menschlicher Aktionen - Aioga KI-News","description":"Der GPT-6 Astra von OpenAI erzielte auf ARC-AGI-3 Semi-Private mit dem Standard-Harness 62,7 % (Kosten 26.000 $) und mit dem Provider-Adapter-Harness 99,9 % (Kosten 19.000 $), beid...","url":"https://www.aioga.com/de/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:05:59.968Z"},"pt-BR":{"title":"OpenAI GPT-6 Astra alcançou SOTA no ARC-AGI-3 e superou a linha de base de eficiência de ação humana","summary":"O GPT-6 Astra da OpenAI obteve 62,7% de pontuação no ARC-AGI-3 Semi-Private com o Standard harness (custo $26K) e 99,9% com o Provider Adapter harness (custo $19K), ambos sendo SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra alcançou SOTA no ARC-AGI-3 e superou a linha de base de eficiência de ação humana - Aioga Notícias de IA","description":"O GPT-6 Astra da OpenAI obteve 62,7% de pontuação no ARC-AGI-3 Semi-Private com o Standard harness (custo $26K) e 99,9% com o Provider Adapter harness (custo $19K), ambos sendo SOT...","url":"https://www.aioga.com/pt-BR/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:06:43.797Z"},"ru":{"title":"OpenAI GPT-6 Astra достигла SOTA на ARC-AGI-3 и превзошла базовый уровень эффективности действий человека","summary":"GPT-6 Astra от OpenAI на ARC-AGI-3 Semi-Private, стандартный тест показал 62,7% (стоимость $26K), тест с адаптером провайдера показал 99,9% (стоимость $19K), обе цифры являются SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra достигла SOTA на ARC-AGI-3 и превзошла базовый уровень эффективности действий человека - Aioga Новости ИИ","description":"GPT-6 Astra от OpenAI на ARC-AGI-3 Semi-Private, стандартный тест показал 62,7% (стоимость $26K), тест с адаптером провайдера показал 99,9% (стоимость $19K), обе цифры являются SOT...","url":"https://www.aioga.com/ru/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:06:44.972Z"},"ar":{"title":"حققت OpenAI GPT-6 Astra أفضل أداء على ARC-AGI-3 وتجاوزت كفاءة الحركة البشرية الأساسية","summary":"حقق GPT-6 Astra من OpenAI على ARC-AGI-3 Semi-Private، في الاختبار القياسي Standard harness نسبة 62.7٪ (بتكلفة 26 ألف دولار)، وفي اختبار Provider Adapter harness نسبة 99.9٪ (بتكلفة 19 ألف دولار)، وكلاهما يعتبر الأفضل في مجاله SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"حققت OpenAI GPT-6 Astra أفضل أداء على ARC-AGI-3 وتجاوزت كفاءة الحركة البشرية الأساسية - Aioga أخبار الذكاء الاصطناعي","description":"حقق GPT-6 Astra من OpenAI على ARC-AGI-3 Semi-Private، في الاختبار القياسي Standard harness نسبة 62.7٪ (بتكلفة 26 ألف دولار)، وفي اختبار Provider Adapter harness نسبة 99.9٪ (بتكلفة...","url":"https://www.aioga.com/ar/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:07:25.345Z"},"hi":{"title":"OpenAI GPT-6 Astra ने ARC-AGI-3 पर SOTA हासिल किया और मानव क्रिया दक्षता मानक को पार किया","summary":"OpenAI का GPT-6 Astra ARC-AGI-3 सेमी-प्राइवेट पर, स्टैंडर्ड हार्नेस ने 62.7% स्कोर किया (लागत $26K), प्रोवाइडर एडाप्टर हार्नेस ने 99.9% स्कोर किया (लागत $19K), दोनों ही SOTA हैं।","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra ने ARC-AGI-3 पर SOTA हासिल किया और मानव क्रिया दक्षता मानक को पार किया - Aioga AI समाचार","description":"OpenAI का GPT-6 Astra ARC-AGI-3 सेमी-प्राइवेट पर, स्टैंडर्ड हार्नेस ने 62.7% स्कोर किया (लागत $26K), प्रोवाइडर एडाप्टर हार्नेस ने 99.9% स्कोर किया (लागत $19K), दोनों ही SOTA हैं।","url":"https://www.aioga.com/hi/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:07:28.655Z"},"it":{"title":"OpenAI GPT-6 Astra ha raggiunto lo SOTA su ARC-AGI-3 e ha superato il livello di efficienza delle azioni umane","summary":"GPT-6 Astra di OpenAI su ARC-AGI-3 Semi-Private, il harness Standard ha ottenuto il 62,7% (costo $26K), il harness Provider Adapter ha ottenuto il 99,9% (costo $19K), entrambi SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra ha raggiunto lo SOTA su ARC-AGI-3 e ha superato il livello di efficienza delle azioni umane - Aioga Notizie IA","description":"GPT-6 Astra di OpenAI su ARC-AGI-3 Semi-Private, il harness Standard ha ottenuto il 62,7% (costo $26K), il harness Provider Adapter ha ottenuto il 99,9% (costo $19K), entrambi SOTA...","url":"https://www.aioga.com/it/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:08:12.002Z"},"nl":{"title":"OpenAI GPT-6 Astra behaalt SOTA op ARC-AGI-3 en overtreft de menselijke actierefficiëntie-basislijn","summary":"De GPT-6 Astra van OpenAI behaalde op ARC-AGI-3 Semi-Private een score van 62,7% met de Standard harness (kosten $26K) en 99,9% met de Provider Adapter harness (kosten $19K), beide SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra behaalt SOTA op ARC-AGI-3 en overtreft de menselijke actierefficiëntie-basislijn - Aioga AI-nieuws","description":"De GPT-6 Astra van OpenAI behaalde op ARC-AGI-3 Semi-Private een score van 62,7% met de Standard harness (kosten $26K) en 99,9% met de Provider Adapter harness (kosten $19K), beide...","url":"https://www.aioga.com/nl/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:08:13.552Z"},"tr":{"title":"OpenAI GPT-6 Astra, ARC-AGI-3 üzerinde SOTA elde etti ve insan hareketi verimliliği temelini aştı","summary":"OpenAI'nin GPT-6 Astra'sı ARC-AGI-3 Yarı-Özel üzerinde, Standard harness ile %62,7 puan aldı (maliyet $26K), Provider Adapter harness ile %99,9 puan aldı (maliyet $19K), her ikisi de SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra, ARC-AGI-3 üzerinde SOTA elde etti ve insan hareketi verimliliği temelini aştı - Aioga AI Haberleri","description":"OpenAI'nin GPT-6 Astra'sı ARC-AGI-3 Yarı-Özel üzerinde, Standard harness ile %62,7 puan aldı (maliyet $26K), Provider Adapter harness ile %99,9 puan aldı (maliyet $19K), her ikisi...","url":"https://www.aioga.com/tr/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:08:54.743Z"},"vi":{"title":"OpenAI GPT-6 Astra đã đạt SOTA trên ARC-AGI-3 và vượt qua chuẩn hiệu quả hành động của con người","summary":"GPT-6 Astra của OpenAI trên ARC-AGI-3 Semi-Private, Standard harness đạt 62,7% (chi phí $26K), Provider Adapter harness đạt 99,9% (chi phí $19K), đều là SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra đã đạt SOTA trên ARC-AGI-3 và vượt qua chuẩn hiệu quả hành động của con người - Tin tức AI Aioga","description":"GPT-6 Astra của OpenAI trên ARC-AGI-3 Semi-Private, Standard harness đạt 62,7% (chi phí $26K), Provider Adapter harness đạt 99,9% (chi phí $19K), đều là SOTA.","url":"https://www.aioga.com/vi/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:08:58.076Z"},"id":{"title":"OpenAI GPT-6 Astra mencapai SOTA di ARC-AGI-3 dan melampaui tolok ukur efisiensi aksi manusia","summary":"GPT-6 Astra dari OpenAI di ARC-AGI-3 Semi-Private, Standard harness mendapatkan skor 62,7% (biaya $26K), Provider Adapter harness mendapatkan skor 99,9% (biaya $19K), keduanya merupakan SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra mencapai SOTA di ARC-AGI-3 dan melampaui tolok ukur efisiensi aksi manusia - Berita AI Aioga","description":"GPT-6 Astra dari OpenAI di ARC-AGI-3 Semi-Private, Standard harness mendapatkan skor 62,7% (biaya $26K), Provider Adapter harness mendapatkan skor 99,9% (biaya $19K), keduanya meru...","url":"https://www.aioga.com/id/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:09:38.823Z"},"th":{"title":"OpenAI GPT-6 Astra ได้สร้าง SOTA บน ARC-AGI-3 และเหนือกว่ามาตรฐานประสิทธิภาพการเคลื่อนไหวของมนุษย์","summary":"GPT-6 Astra ของ OpenAI ใน ARC-AGI-3 Semi-Private, Standard harness ได้คะแนน 62.7% (ค่าใช้จ่าย $26K), Provider Adapter harness ได้คะแนน 99.9% (ค่าใช้จ่าย $19K), ทั้งคู่เป็น SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra ได้สร้าง SOTA บน ARC-AGI-3 และเหนือกว่ามาตรฐานประสิทธิภาพการเคลื่อนไหวของมนุษย์ - ข่าว AI Aioga","description":"GPT-6 Astra ของ OpenAI ใน ARC-AGI-3 Semi-Private, Standard harness ได้คะแนน 62.7% (ค่าใช้จ่าย $26K), Provider Adapter harness ได้คะแนน 99.9% (ค่าใช้จ่าย $19K), ทั้งคู่เป็น SOTA.","url":"https://www.aioga.com/th/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:09:42.740Z"},"pl":{"title":"OpenAI GPT-6 Astra osiągnął SOTA na ARC-AGI-3 i przekroczył ludzką bazę efektywności działania","summary":"GPT-6 Astra firmy OpenAI w ARC-AGI-3 Semi-Private uzyskał w standardowym zestawie testowym wynik 62,7% (koszt $26K), a w zestawie testowym Provider Adapter wynik 99,9% (koszt $19K), oba wyniki są SOTA.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"OpenAI GPT-6 Astra osiągnął SOTA na ARC-AGI-3 i przekroczył ludzką bazę efektywności działania - Aioga Wiadomości AI","description":"GPT-6 Astra firmy OpenAI w ARC-AGI-3 Semi-Private uzyskał w standardowym zestawie testowym wynik 62,7% (koszt $26K), a w zestawie testowym Provider Adapter wynik 99,9% (koszt $19K)...","url":"https://www.aioga.com/pl/news/cmtm7yl5s01d5robnsn25onbz/","contentTranslated":true,"sourceHash":"a31b55267c4db2ef","translatedAt":"2026-09-04T01:10:26.044Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}