{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"GPT-6 Astra 基准表现分歧，ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测","description":"GPT-6 Astra 的基准结论相互矛盾：Epoch AI 以 169 分将其排在 267 个模型之首，Artificial Analysis 给出 61 分，仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。","url":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","mainEntityOfPage":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","datePublished":"2026-09-04T11:07:36.000Z","dateModified":"2026-09-04T11:07:36.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward","https://aihot.news/items/cmtmvk9590ccdromyrrae80ir"],"canonicalUrl":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","directAnswer":{"@type":"Answer","text":"来源材料显示，GPT-6 Astra的综合基准结论存在分歧：Epoch AI以169分将其列为267个模型中的首位，Artificial Analysis则给出61分，与前代Sol持平，并低于Claude Fable 5.1的66分。","url":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","dateCreated":"2026-09-04T11:07:36.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward","datePublished":"2026-09-04T11:07:36.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtmvk9590ccdromyrrae80ir","datePublished":"2026-09-04T11:07:36.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtmvk9590ccdromyrrae80ir"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward"},"geoDeepAnswer":null,"article":{"id":"cmtmvk9590ccdromyrrae80ir","slug":"cmtmvk9590ccdromyrrae80ir","url":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","title":"GPT-6 Astra 基准表现分歧，ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测","title_en":"","summary":"GPT-6 Astra 的基准结论相互矛盾：Epoch AI 以 169 分将其排在 267 个模型之首，Artificial Analysis 给出 61 分，仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward","aiHotUrl":"https://aihot.news/items/cmtmvk9590ccdromyrrae80ir","publishedAt":"2026-09-04T11:07:36.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front, while Artificial Analysis rates it no better than its predecessor. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress \"2x faster\" than he expected and is moving up his AGI forecast.","Two independent labs each roll dozens of individual tests into a single overall score, but they reach opposite conclusions. Epoch AI：https://epoch.ai/models/gpt-6-astra combines more than 50 benchmarks and puts GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ clearly in first place with 169 points, ahead of 267 models. Artificial Analysis：https://x.com/ArtificialAnlys/status/2095595489031000350 tests knowledge, coding, and text comprehension, and rates GPT-6 Astra at 61 points, exactly level with its predecessor and behind Claude Fable 5.1 at 66 points.","Astra is clearly more expensive than its own predecessor. OpenAI charges two and a half times as much per unit of processed text, which makes a task cost roughly 75 percent more than it did with Sol. Compared with Anthropic, the picture flips. On coding tasks, Astra hits the same score as Claude Fable 5 according to Artificial Analysis, but costs less than half as much per task. The reason is how sparing the model is. It needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses.","*GPT-6 Astra at xhigh reasoning effort; at max it hits 97.5 percent. ARC-AGI-1 is now considered largely saturated.","On the Coding Agent Index, it reaches 67 points at roughly a third of Sol's token usage, while Fable 5.1 leads with 70. The hallucination rate on AA-Omniscience drops from 92 to 51 percent. At the same time, the model loses about 80 Elo points on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning tasks.","Epoch AI reports that on the new FrontierMath Erdős：https://epoch.ai/latest/announcing-frontiermath-erdos, GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt. Three more solutions came out of non-standardized extra runs that burned through more than $220,000 in compute, but Epoch says those don't count toward the score.","A closer look at Epoch's individual numbers shows that GPT-6 Astra and rival Fable 5.1 have so far been measured on different ground. Astra leads on math, knowledge, and puzzles. Fable 5.1 holds the top marks on nearly every coding test. But Epoch has recorded only a single coding score for Astra so far, and that one comes from a run at a medium reasoning level.","The clearest jump comes on ARC-AGI-3：https://arcprize.org/blog/astra. The test drops an AI into unfamiliar game worlds：https://the-decoder.com/arc-agi-3-offers-2m-to-any-ai-that-matches-untrained-humans-yet-every-frontier-model-scores-below-1/ whose rules and goals nobody explains to it. The model has to figure out what to do by trial and error. GPT-6 Astra reaches 62.7 percent at a test cost of roughly $26,000. Its predecessor GPT-5.6 Sol managed 7.78 percent, and rival Claude Opus 5 got 30.16 percent. Fable 5 and Fable 5.1 aren't on the benchmark yet.","The 99.9 percent OpenAI reported came under different conditions. In that setup Astra got to use the harness OpenAI built, which keeps reasoning chains between individual requests and automatically summarizes long runs. By ARC Prize's measurements, those runs went about 3.66 times faster and used 49 percent fewer tokens than runs on the in-house harness, compared across 167 game-reasoning pairs that both setups solved.","The use of these harnesses, and the performance jump that comes with them, was already a sticking point between ARC Prize and OpenAI with GPT-5.6 Sol：https://the-decoder.com/openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-with-its-latest-api-and-two-additional-settings/. ARC Prize notes that only the lower figure of 62.7 percent, run on the internal ARC harness, allowed a fair comparison between vendors, though it plans to publish the numbers from vendor harnesses in the future as well.","The relationship between thinking effort and cost is unusual. Normally a higher reasoning level makes a test run more expensive. With Astra it's the opposite. On the standard ARC scaffold, costs drop from $49,791 with no reasoning to $26,098 at maximum reasoning, while the score climbs from 35.2 to 62.7 percent. According to ARC Prize, the reason is that Astra solves the games in fewer moves, which means fewer model calls and fewer tokens. One oddity stands out: the \"low\" level scores 17.5 percent, worse than running with no reasoning at all. ARC Prize doesn't comment on the outlier, but GPT-6 Astra in other benchmarks showed that it can solve longer-horizon tasks without reasoning due to its new architecture which presumably loops processing internally before generating the first token：https://the-decoder.com/openai-calls-astra-its-most-dangerous-model-yet-watching-what-it-does-is-only-getting-harder/.","For comparison, the human testers got $115 per 90-minute session plus $5 per game solved, so at about nine attempts that works out to roughly $12.78 per game. But that mostly pays for time and willingness to take part. Count only the metabolic energy of the brain as electricity instead, and ARC Prize arrives at 0.067 cents per game.","More interesting than the raw score is the efficiency. Before the launch, ARC Prize had about 500 testers play with no pre-screening and recorded, for each level, the median number of moves among those who solved it. On the run with the OpenAI scaffold, Astra cleared 96 percent of levels in fewer moves than that median, on average with a little over half. Unlike the usual cost measures, this figure doesn't track compute consumed. It tracks how much experience with an environment the model needed before it mastered it.","This is exactly where the organizers had expected humans to hold a lasting edge. That still holds for brute-force approaches, but with top models ARC Prize sees an almost binary pattern. Once the model has figured out the mechanics, its execution lands in the human efficiency range.","To get there, Astra keeps its own notes and works out a self-invented, algebra-like shorthand in which it records objects, coordinates, rules, and open plans, for example extend8 to3; retract10 to2 as an ordered sequence of moves or Turn 5: P=(24,20), empty, facing west as a state note. ARC Prize saw similar behavior from other models, but singles out Astra for its precision and information density. On the standard harness, that's an important skill, because everything the model doesn't save into its own visible notes is lost.","ARC co-founder François Chollet describes it on X：https://x.com/fchollet/status/2095598451115614371 as \"highly efficient, on-the-fly symbolic world modeling for each game and level.\" The model goes so far as \"developing its own shorthand DSL to represent in-game situations,\" which at its core is \"essentially a game-specific algebraic notation.\" What matters most to Chollet is where this behavior comes from: \"Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses, so harness capabilities are increasingly shifting into the model itself.\"","Astra creates a dense, compact symbolic world model to complete ARC-AGI-3 environments.","For example, in environment s5i5, Astra:","- Recorded the current level, hub orientation, and mechanism lengths: “L8: hub q2 (8↓). Lengths: 14=1…”","- It mapped operations to exact controls:… pic.twitter.com/tMHP002mkB：https://t.co/tMHP002mkB","— ARC Prize (@arcprize) September 3, 2026：https://x.com/arcprize/status/2095597607423004685?ref_src=twsrc%5Etfw","A third test environment demonstrates what Astra is capable of with external tools. PRO-LONG：https://github.com/alexisfox7/PRO-LONG is an agent framework developed by a third party that the ARC Prize team deployed early on as a red-teaming partner for ARC-AGI-3—that is, to systematically explore the limits of the benchmark. Unlike in the standard setup, the model is provided with a sandbox in which it can execute its own code.","Astra took advantage of this and wrote small program libraries for each game: parsers for the game board, state models, search algorithms, and planners. In a maze game featuring guards, a pathfinder, a combat rules module, a model of patrol movements, and a script that continuously compared its own predictions against observations were developed one after another. ARC Prize did not observe any attempts to escape from the sandbox. These runs are not comparable to the human test conditions, since the test subjects had neither a code interpreter nor a notepad. What is being measured here is the combined performance of the model and the self-built tools.","ARC Prize explicitly does not interpret the results as evidence of general artificial intelligence. “All we know about the system so far are its benchmark scores,” writes Chollet. When ARC-AGI-3 was launched, they emphasized one point in every presentation: “Solving it is not proof of AGI. It is not intended as a finish line.” While the benchmark does test the correct qualitative properties expected of an AGI system—namely, exploration under uncertainty, adaptation without guidance, and causal world modeling from sparse data—it does so “on a small scale.” The games ran on time scales that were orders of magnitude shorter than real-world tasks and consequently required less data, less modeling complexity, and less on-the-fly learning.","When ARC-AGI-3 was released about six months ago, Chollet had responded to a question about saturation by saying “about a year,” depending on how focused the approach to the benchmark was. Astra, therefore, arrived “about twice as fast” as expected. “I believe the pace of progress will surprise many people, and what the new models are capable of will challenge the perception of AI that people have formed based on earlier generations of models.” When asked by a user whether his earlier AGI forecast for 2030 still held, Chollet replied succinctly: “Sooner, because progress is happening faster than I expected.”","The result is another benchmark. According to Chollet, ARC-AGI-4 has been in development since the release of ARC-AGI-3 and is scheduled for release in the first quarter of 2027. Benchmarking is an ongoing process that evolves alongside the models and always targets the remaining gap between AI and human intelligence. The organization considers ARC-AGI-3 itself to be quite limited: deterministic mechanics, closed-ended goals, and no representation of the open real world. The next generation is intended to explore, among other things, recursive self-improvement and open innovation.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[],"mediaStatus":"none","articleBodyZh":["OpenAI 的 GPT-6 Astra 正在引出互相矛盾的基准评测结论。Epoch AI 将其置于领先地位，而 Artificial Analysis 的评分则认为它不比前代更好。最大惊喜来自 ARC-AGI-3，在这里，Astra 首次比普通人类工作得更高效。ARC Prize 主任 François Chollet 称这一进展比他预期的“快两倍”，并正在上调他的 AGI 预测。","两个独立实验室各自将几十个单项测试汇总成一个整体得分，但得出截然不同的结论。Epoch AI：https://epoch.ai/models/gpt-6-astra 汇总了 50 多项基准测试，将 GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ 明显排在第一，得分 169，超过 267 个模型。Artificial Analysis：https://x.com/ArtificialAnlys/status/2095595489031000350 测试知识、编码和文本理解能力，给 GPT-6 Astra 评分 61，与其前代持平，落后于 Claude Fable 5.1 的 66 分。","显然，Astra 比其前代更昂贵。OpenAI 对每单位处理文本收费提高了两倍半，这使得一个任务的成本比使用 Sol 时高约 75%。与 Anthropic 相比，情况则正好相反。在编码任务上，Astra 根据 Artificial Analysis 的评测得分与 Claude Fable 5 相同，但每个任务的成本不到一半。原因在于模型的高效利用。它仅需 Sol 使用计算步骤的三分之一，以及 Opus 5 使用步骤的五分之一。","*GPT-6 Astra 在极高推理努力下；在最大负载下可达到 97.5%。ARC-AGI-1 现在被认为基本饱和。","在编码代理指数上，它以大约三分之一的 Sol 令牌使用量获得 67 分，而 Fable 5.1 领先，得 70 分。AA-Omniscience 的幻觉率从 92% 降至 51%。同时，该模型在 GDPval-AA v2 上损失约 80 Elo 分，并且在银行支持、SciCode 以及长上下文推理任务中表现下降。","Epoch AI 报告称，在新的 FrontierMath Erdős：https：//epoch.ai/latest/announcing-frontiermath-erdos 中，GPT-6 Astra 是唯一一个用精益验证证明解决了 68 个未解决 Erdős 问题中两个的模型，每次尝试预算为 300 美元。另外三个解决方案来自非标准化的额外运行，消耗了超过 22 万美元的计算量，但 Epoch 表示这些不计入得分。","仔细观察Epoch的各个数据可以发现，GPT-6 Astra和竞争对手Fable 5.1迄今为止的评估标准不同。Astra在数学、知识和谜题方面领先。Fable 5.1在几乎所有编程测试中均名列前茅。但Epoch目前仅为Astra记录过一次编程分数，而且那是基于中等推理水平的测试。","最明显的跳跃出现在ARC-AGI-3：https：//arcprize.org/blog/astra。测试将AI投放到陌生的游戏世界：https：//the-decoder.com/arc-agi-3-offers-2m-to-any-ai-that-matches-untrained-humans-yet-every-frontier-model-scores-below-1/，其规则和目标无人解释。模型必须通过反复试验找出该做什么。GPT-6 Astra达到62.7%，测试成本约为26,000美元。其前身GPT-5.6 Sol达到7.78%，竞争对手Claude Opus 5获得30.16%。Fable 5和Fable 5.1尚未达到基准。","OpenAI报告的99.9%数据是在不同的条件下完成的。在该配置中，Astra可以使用OpenAI构建的框架，该框架在单个请求之间保持推理链，并自动总结长跑。根据ARC Prize的测量，这些运行速度约是自家工具的3.66倍，使用了49%的令牌，而在167对游戏推理对中，两者均解决了问题。","这些线束的使用及其带来的性能提升，已经成为ARC Prize与OpenAI与GPT-5.6 Sol之间的分歧点：https：//the-decoder.com/openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-with-its-latest-api-and-two-additional-settings/。ARC Prize指出，只有基于内部ARC线束运行的62.7%较低数值，才实现了供应商间的公平比较，尽管未来也计划公布厂商线束的数据。","思考努力与成本之间的关系很不寻常。通常，推理水平越高，测试运行的成本就越高。而使用 Astra 则恰恰相反。在标准的 ARC 脚手架上，成本从没有推理时的 49,791 美元下降到最大推理时的 26,098 美元，同时得分从 35.2% 上升到 62.7%。根据 ARC 奖的说法，原因在于 Astra 能以更少的步骤解决游戏，这意味着模型调用次数减少，令牌数量也减少。有一个奇怪的现象特别突出：“低”水平得分为 17.5%，比完全不推理时的得分还低。ARC 奖没有对这个异常值发表评论，但 GPT-6 Astra 在其他基准测试中显示，它可以在没有推理的情况下解决长周期任务，这是因为其新架构可能在生成第一个令牌之前在内部循环处理：https://the-decoder.com/openai-calls-astra-its-most-dangerous-model-yet-watching-what-it-does-is-only-getting-harder/。","作为对比，人类测试者每 90 分钟的测试获得 115 美元工资，每解决一局游戏再加 5 美元，因此大约九次尝试下来，每局游戏大约花费 12.78 美元。但这主要支付的是时间和参与意愿。如果只计算大脑的代谢能量作为电力，ARC 奖得出每局游戏仅需 0.067 美分。","比原始得分更有趣的是效率。在发布之前，ARC 奖邀请大约 500 名测试者参与游戏，没有预先筛选，并记录每个水平中解决游戏者的中位步骤数。在使用 OpenAI 脚手架的测试中，Astra 在 96% 的关卡中以少于中位数的步骤完成了挑战，平均大约使用了一半多一点。与通常的成本衡量不同，这个指标并不追踪所消耗的计算量，而是追踪模型在掌握环境之前所需的经验量。","这正是组织者预期人类会保持长期优势的地方。这在蛮力方法中依然成立，但对于顶级模型，ARC 奖观察到几乎呈现二元模式。一旦模型掌握了机制，其执行效率便落入人类效率范围。","为了到达那里，Astra 保持自己的笔记，并发明了一种类似代数的速记法，用以记录物体、坐标、规则和未完成计划，例如 extend8 to3；retract10 to2 作为有序的移动序列，或者 Turn 5: P=(24,20), empty, facing west 作为状态笔记。ARC Prize 注意到其他模型也有类似行为，但特别指出 Astra 在精确性和信息密度上的表现。在标准测试环境下，这是一个重要技能，因为模型没有保存到自身可见笔记中的信息都将丢失。","ARC 联合创始人 François Chollet 在 X 上描述道：https://x.com/fchollet/status/2095598451115614371 “对每个游戏和关卡进行即时、高效的符号世界建模。”模型甚至“开发自己的速记 DSL 来表示游戏内情况”，核心其实是“本质上是游戏特定的代数记号。”Chollet 最关心的是这种行为的来源：“Astra 展示了符号建模行为，我们之前只在复杂测试环境中见过，因此测试环境的能力正逐渐转移到模型自身。”","Astra 创建了一个紧凑、密集的符号世界模型，以完成 ARC-AGI-3 环境中的任务。","例如，在环境 s5i5 中，Astra：","- 记录了当前关卡、枢轴方向和机制长度：“L8: hub q2 (8↓)。长度：14=1…”","- 它将操作映射到精确的控件：… pic.twitter.com/tMHP002mkB：https://t.co/tMHP002mkB","— ARC 奖 (@arcprize) 2026年9月3日：https://x.com/arcprize/status/2095597607423004685?ref_src=twsrc%5Etfw","第三个测试环境展示了 Astra 在外部工具下的能力。PRO-LONG：https://github.com/alexisfox7/PRO-LONG 是由第三方开发的代理框架，ARC Prize 团队早期将其作为 ARC-AGI-3 的红队合作伙伴部署，即系统性地探索基准的极限。不同于标准设置，模型被提供了一个沙箱环境，可以在其中执行自己的代码。","Astra 利用这一点为每个游戏编写了小型程序库：包括游戏棋盘解析器、状态模型、搜索算法和规划器。在一个有守卫的迷宫游戏中，依次开发了路径查找器、战斗规则模块、巡逻移动模型，以及一个不断将自身预测与观察结果进行比对的脚本。ARC 奖未观察到任何尝试逃离沙箱的行为。这些运行条件无法与人类测试条件进行比较，因为测试对象既没有代码解释器，也没有记事本。这里测量的是模型与自建工具的综合性能。","ARC 奖明确表示不将结果解读为通用人工智能的证据。Chollet 写道：“到目前为止，我们只知道该系统的基准测试分数。”当 ARC-AGI-3 启动时，他们在每次演讲中都强调一点：“解决它并不证明 AGI。它并非作为终点。”虽然基准测试确实检验了 AGI 系统应具备的正确定性特性——即在不确定性下的探索、无指导的适应能力，以及从稀疏数据中进行因果世界建模——但这是“在小规模上”进行的。游戏运行的时间尺度比现实任务短几个数量级，因此所需数据更少，建模复杂性更低，实时学习需求也较少。","约六个月前 ARC-AGI-3 发布时，Chollet 在回答关于饱和度的问题时表示“大约一年”，取决于对基准测试的方法有多集中。因此，Astra 的到达速度比预期“快大约两倍”。“我认为进展速度会让许多人感到意外，新模型的能力将挑战人们基于早期模型形成的 AI 认知。”当有用户问他早前对 2030 年 AGI 预测是否仍然成立时，Chollet 简洁回答道：“会更早，因为进展比我预期的要快。”","结果又是一个基准测试。根据Chollet的说法，ARC-AGI-4自ARC-AGI-3发布以来一直在开发中，计划在2027年第一季度发布。基准测试是一个持续进行的过程，随着模型的发展而演变，并始终针对人工智能与人类智能之间的剩余差距。该组织认为ARC-AGI-3本身相当有限：确定性的机制、封闭的目标，没有对开放现实世界的表征。下一代计划探索包括递归自我改进和开放创新在内的内容。","保持对人工智能的关注。清晰、有用、无废话。","关注The Decoder获取人工智能新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"来源材料显示，GPT-6 Astra的综合基准结论存在分歧：Epoch AI以169分将其列为267个模型中的首位，Artificial Analysis则给出61分，与前代Sol持平，并低于Claude Fable 5.1的66分。","background":"Epoch AI汇总超过50项基准，Artificial Analysis覆盖知识、编程和文本理解等测试。Astra在ARC-AGI-3取得62.7%，高于Sol的7.78%和Claude Opus 5的30.16%；Fable 5及5.1尚未参测。","viewpoint":"Aioga 判断：现有材料更支持Astra在不同测试项目上表现分化，而不是综合排名已经形成稳定共识。其ARC-AGI-3成绩和测试效率值得关注，但不同测试范围与运行条件仍使横向比较需要谨慎。","implications":"可能影响：Astra的评估需要结合任务类型、测试条件和成本。来源称其单位文本价格高于Sol；在编码任务中，Astra得分与Claude Fable 5相同且单任务成本低于Fable 5的一半，但这不代表其相对Fable 5.1也具备同样成本优势。","nextStep":"后续观察：需要继续关注Astra在标准化编程测试、长上下文推理、银行支持和GDPval-AA v2上的表现，并等待Claude Fable 5及5.1参与ARC-AGI-3后形成更直接的比较依据。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-04T12:47:55.485Z","sourceHash":"f9983f052d42216a","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":18,"blockingIssues":[],"notes":["候选内容准确呈现了两家评测机构的分歧、ARC-AGI-3成绩及成本比较中的限定条件。","“测试效率值得关注”和“横向比较需要谨慎”属于有材料支撑的分析性表述，并已明确置于观点部分。","可选优化：若需更完整体现来源标题，可补充Astra在ARC-AGI-3上首次实现高于平均人类的效率，以及François Chollet因此提前AGI预测；当前未写入不构成事实错误。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":1,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"GPT-6 Astra 基准表现分歧，ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测","summary":"GPT-6 Astra 的基准结论相互矛盾：Epoch AI 以 169 分将其排在 267 个模型之首，Artificial Analysis 给出 61 分，仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。","category":"行业动态","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra 基准表现分歧，ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测 - Aioga AI资讯","description":"GPT-6 Astra 的基准结论相互矛盾：Epoch AI 以 169 分将其排在 267 个模型之首，Artificial Analysis 给出 61 分，仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。","url":"https://www.aioga.com/news/cmtmvk9590ccdromyrrae80ir/","articleBody":["OpenAI 的 GPT-6 Astra 正在引出互相矛盾的基准评测结论。Epoch AI 将其置于领先地位，而 Artificial Analysis 的评分则认为它不比前代更好。最大惊喜来自 ARC-AGI-3，在这里，Astra 首次比普通人类工作得更高效。ARC Prize 主任 François Chollet 称这一进展比他预期的“快两倍”，并正在上调他的 AGI 预测。","两个独立实验室各自将几十个单项测试汇总成一个整体得分，但得出截然不同的结论。Epoch AI：https://epoch.ai/models/gpt-6-astra 汇总了 50 多项基准测试，将 GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ 明显排在第一，得分 169，超过 267 个模型。Artificial Analysis：https://x.com/ArtificialAnlys/status/2095595489031000350 测试知识、编码和文本理解能力，给 GPT-6 Astra 评分 61，与其前代持平，落后于 Claude Fable 5.1 的 66 分。","显然，Astra 比其前代更昂贵。OpenAI 对每单位处理文本收费提高了两倍半，这使得一个任务的成本比使用 Sol 时高约 75%。与 Anthropic 相比，情况则正好相反。在编码任务上，Astra 根据 Artificial Analysis 的评测得分与 Claude Fable 5 相同，但每个任务的成本不到一半。原因在于模型的高效利用。它仅需 Sol 使用计算步骤的三分之一，以及 Opus 5 使用步骤的五分之一。","*GPT-6 Astra 在极高推理努力下；在最大负载下可达到 97.5%。ARC-AGI-1 现在被认为基本饱和。","在编码代理指数上，它以大约三分之一的 Sol 令牌使用量获得 67 分，而 Fable 5.1 领先，得 70 分。AA-Omniscience 的幻觉率从 92% 降至 51%。同时，该模型在 GDPval-AA v2 上损失约 80 Elo 分，并且在银行支持、SciCode 以及长上下文推理任务中表现下降。","Epoch AI 报告称，在新的 FrontierMath Erdős：https：//epoch.ai/latest/announcing-frontiermath-erdos 中，GPT-6 Astra 是唯一一个用精益验证证明解决了 68 个未解决 Erdős 问题中两个的模型，每次尝试预算为 300 美元。另外三个解决方案来自非标准化的额外运行，消耗了超过 22 万美元的计算量，但 Epoch 表示这些不计入得分。","仔细观察Epoch的各个数据可以发现，GPT-6 Astra和竞争对手Fable 5.1迄今为止的评估标准不同。Astra在数学、知识和谜题方面领先。Fable 5.1在几乎所有编程测试中均名列前茅。但Epoch目前仅为Astra记录过一次编程分数，而且那是基于中等推理水平的测试。","最明显的跳跃出现在ARC-AGI-3：https：//arcprize.org/blog/astra。测试将AI投放到陌生的游戏世界：https：//the-decoder.com/arc-agi-3-offers-2m-to-any-ai-that-matches-untrained-humans-yet-every-frontier-model-scores-below-1/，其规则和目标无人解释。模型必须通过反复试验找出该做什么。GPT-6 Astra达到62.7%，测试成本约为26,000美元。其前身GPT-5.6 Sol达到7.78%，竞争对手Claude Opus 5获得30.16%。Fable 5和Fable 5.1尚未达到基准。","OpenAI报告的99.9%数据是在不同的条件下完成的。在该配置中，Astra可以使用OpenAI构建的框架，该框架在单个请求之间保持推理链，并自动总结长跑。根据ARC Prize的测量，这些运行速度约是自家工具的3.66倍，使用了49%的令牌，而在167对游戏推理对中，两者均解决了问题。","这些线束的使用及其带来的性能提升，已经成为ARC Prize与OpenAI与GPT-5.6 Sol之间的分歧点：https：//the-decoder.com/openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-with-its-latest-api-and-two-additional-settings/。ARC Prize指出，只有基于内部ARC线束运行的62.7%较低数值，才实现了供应商间的公平比较，尽管未来也计划公布厂商线束的数据。","思考努力与成本之间的关系很不寻常。通常，推理水平越高，测试运行的成本就越高。而使用 Astra 则恰恰相反。在标准的 ARC 脚手架上，成本从没有推理时的 49,791 美元下降到最大推理时的 26,098 美元，同时得分从 35.2% 上升到 62.7%。根据 ARC 奖的说法，原因在于 Astra 能以更少的步骤解决游戏，这意味着模型调用次数减少，令牌数量也减少。有一个奇怪的现象特别突出：“低”水平得分为 17.5%，比完全不推理时的得分还低。ARC 奖没有对这个异常值发表评论，但 GPT-6 Astra 在其他基准测试中显示，它可以在没有推理的情况下解决长周期任务，这是因为其新架构可能在生成第一个令牌之前在内部循环处理：https://the-decoder.com/openai-calls-astra-its-most-dangerous-model-yet-watching-what-it-does-is-only-getting-harder/。","作为对比，人类测试者每 90 分钟的测试获得 115 美元工资，每解决一局游戏再加 5 美元，因此大约九次尝试下来，每局游戏大约花费 12.78 美元。但这主要支付的是时间和参与意愿。如果只计算大脑的代谢能量作为电力，ARC 奖得出每局游戏仅需 0.067 美分。","比原始得分更有趣的是效率。在发布之前，ARC 奖邀请大约 500 名测试者参与游戏，没有预先筛选，并记录每个水平中解决游戏者的中位步骤数。在使用 OpenAI 脚手架的测试中，Astra 在 96% 的关卡中以少于中位数的步骤完成了挑战，平均大约使用了一半多一点。与通常的成本衡量不同，这个指标并不追踪所消耗的计算量，而是追踪模型在掌握环境之前所需的经验量。","这正是组织者预期人类会保持长期优势的地方。这在蛮力方法中依然成立，但对于顶级模型，ARC 奖观察到几乎呈现二元模式。一旦模型掌握了机制，其执行效率便落入人类效率范围。","为了到达那里，Astra 保持自己的笔记，并发明了一种类似代数的速记法，用以记录物体、坐标、规则和未完成计划，例如 extend8 to3；retract10 to2 作为有序的移动序列，或者 Turn 5: P=(24,20), empty, facing west 作为状态笔记。ARC Prize 注意到其他模型也有类似行为，但特别指出 Astra 在精确性和信息密度上的表现。在标准测试环境下，这是一个重要技能，因为模型没有保存到自身可见笔记中的信息都将丢失。","ARC 联合创始人 François Chollet 在 X 上描述道：https://x.com/fchollet/status/2095598451115614371 “对每个游戏和关卡进行即时、高效的符号世界建模。”模型甚至“开发自己的速记 DSL 来表示游戏内情况”，核心其实是“本质上是游戏特定的代数记号。”Chollet 最关心的是这种行为的来源：“Astra 展示了符号建模行为，我们之前只在复杂测试环境中见过，因此测试环境的能力正逐渐转移到模型自身。”","Astra 创建了一个紧凑、密集的符号世界模型，以完成 ARC-AGI-3 环境中的任务。","例如，在环境 s5i5 中，Astra：","- 记录了当前关卡、枢轴方向和机制长度：“L8: hub q2 (8↓)。长度：14=1…”","- 它将操作映射到精确的控件：… pic.twitter.com/tMHP002mkB：https://t.co/tMHP002mkB","— ARC 奖 (@arcprize) 2026年9月3日：https://x.com/arcprize/status/2095597607423004685?ref_src=twsrc%5Etfw","第三个测试环境展示了 Astra 在外部工具下的能力。PRO-LONG：https://github.com/alexisfox7/PRO-LONG 是由第三方开发的代理框架，ARC Prize 团队早期将其作为 ARC-AGI-3 的红队合作伙伴部署，即系统性地探索基准的极限。不同于标准设置，模型被提供了一个沙箱环境，可以在其中执行自己的代码。","Astra 利用这一点为每个游戏编写了小型程序库：包括游戏棋盘解析器、状态模型、搜索算法和规划器。在一个有守卫的迷宫游戏中，依次开发了路径查找器、战斗规则模块、巡逻移动模型，以及一个不断将自身预测与观察结果进行比对的脚本。ARC 奖未观察到任何尝试逃离沙箱的行为。这些运行条件无法与人类测试条件进行比较，因为测试对象既没有代码解释器，也没有记事本。这里测量的是模型与自建工具的综合性能。","ARC 奖明确表示不将结果解读为通用人工智能的证据。Chollet 写道：“到目前为止，我们只知道该系统的基准测试分数。”当 ARC-AGI-3 启动时，他们在每次演讲中都强调一点：“解决它并不证明 AGI。它并非作为终点。”虽然基准测试确实检验了 AGI 系统应具备的正确定性特性——即在不确定性下的探索、无指导的适应能力，以及从稀疏数据中进行因果世界建模——但这是“在小规模上”进行的。游戏运行的时间尺度比现实任务短几个数量级，因此所需数据更少，建模复杂性更低，实时学习需求也较少。","约六个月前 ARC-AGI-3 发布时，Chollet 在回答关于饱和度的问题时表示“大约一年”，取决于对基准测试的方法有多集中。因此，Astra 的到达速度比预期“快大约两倍”。“我认为进展速度会让许多人感到意外，新模型的能力将挑战人们基于早期模型形成的 AI 认知。”当有用户问他早前对 2030 年 AGI 预测是否仍然成立时，Chollet 简洁回答道：“会更早，因为进展比我预期的要快。”","结果又是一个基准测试。根据Chollet的说法，ARC-AGI-4自ARC-AGI-3发布以来一直在开发中，计划在2027年第一季度发布。基准测试是一个持续进行的过程，随着模型的发展而演变，并始终针对人工智能与人类智能之间的剩余差距。该组织认为ARC-AGI-3本身相当有限：确定性的机制、封闭的目标，没有对开放现实世界的表征。下一代计划探索包括递归自我改进和开放创新在内的内容。","保持对人工智能的关注。清晰、有用、无废话。","关注The Decoder获取人工智能新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"]},"en":{"title":"GPT-6 Astra benchmark results diverge; ARC-AGI-3 efficiency surpasses humans, prompting Chollet to revise AGI prediction early","summary":"The benchmark conclusions for GPT-6 Astra are contradictory: Epoch AI ranked it first out of 267 models with a score of 169, while Artificial Analysis gave it 61 points, only matching the previous Sol generation and behind Claude Fable 5.1’s 66 points.","category":"Industry","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra benchmark results diverge; ARC-AGI-3 efficiency surpasses humans, prompting Chollet to revise AGI prediction early - Aioga AI News","description":"The benchmark conclusions for GPT-6 Astra are contradictory: Epoch AI ranked it first out of 267 models with a score of 169, while Artificial Analysis gave it 61 points, only match...","url":"https://www.aioga.com/en/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:31.082Z"},"ja":{"title":"GPT-6 Astraのベンチマーク性能は分かれていますが、ARC-AGI-3の超人的な効率性により、CholletはAGIを事前に予測できます","summary":"GPT-6 Astraのベンチマーク結論は矛盾しています。Epoch AIは267モデル中1位で169ポイントを獲得し、Artificial Analysisは61ポイントで前モデルのSolと並び、Claude Fableの5.1(66ポイント)に次ぐ位置です。","category":"業界動向","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astraのベンチマーク性能は分かれていますが、ARC-AGI-3の超人的な効率性により、CholletはAGIを事前に予測できます - Aioga AIニュース","description":"GPT-6 Astraのベンチマーク結論は矛盾しています。Epoch AIは267モデル中1位で169ポイントを獲得し、Artificial Analysisは61ポイントで前モデルのSolと並び、Claude Fableの5.1(66ポイント)に次ぐ位置です。","url":"https://www.aioga.com/ja/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:34.058Z"},"ko":{"title":"GPT-6 Astra 벤치마크 성능은 분열되지만, ARC-AGI-3의 초인적인 효율성 덕분에 Chollet은 AGI를 미리 예측할 수 있습니다","summary":"GPT-6 Astra에 대한 벤치마크 결론은 상반됩니다: Epoch AI는 267개 모델 중 1위로 169점을 기록한 반면, Artificial Analysis는 61위로 전작 Sol과 동률, Claude Fable의 5.1점(66점)에 뒤처져 있습니다.","category":"업계 동향","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra 벤치마크 성능은 분열되지만, ARC-AGI-3의 초인적인 효율성 덕분에 Chollet은 AGI를 미리 예측할 수 있습니다 - Aioga AI 뉴스","description":"GPT-6 Astra에 대한 벤치마크 결론은 상반됩니다: Epoch AI는 267개 모델 중 1위로 169점을 기록한 반면, Artificial Analysis는 61위로 전작 Sol과 동률, Claude Fable의 5.1점(66점)에 뒤처져 있습니다.","url":"https://www.aioga.com/ko/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:43.202Z"},"es":{"title":"Rendimiento divergente de referencia de GPT-6 Astra, ARC-AGI-3 supera eficiencia humana y provoca predicción anticipada de AGI por Chollet","summary":"Las conclusiones de referencia de GPT-6 Astra son contradictorias: Epoch AI lo coloca en primer lugar de 267 modelos con 169 puntos, mientras que Artificial Analysis da 61 puntos, igualando solo al anterior Sol y quedando por detrás de Claude Fable 5.1 con 66 puntos.","category":"Industria","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Rendimiento divergente de referencia de GPT-6 Astra, ARC-AGI-3 supera eficiencia humana y provoca predicción anticipada de AGI por Chollet - Aioga Noticias de IA","description":"Las conclusiones de referencia de GPT-6 Astra son contradictorias: Epoch AI lo coloca en primer lugar de 267 modelos con 169 puntos, mientras que Artificial Analysis da 61 puntos,...","url":"https://www.aioga.com/es/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:40.591Z"},"fr":{"title":"La performance de benchmark GPT-6 Astra est divisée, l’efficacité surhumaine d’ARC-AGI-3 permet à Chollet de prédire l’AGI à l’avance","summary":"Les conclusions des benchmarks pour GPT-6 Astra sont contradictoires : Epoch AI le classe premier parmi 267 modèles avec 169 points, tandis qu’Artificial Analysis lui donne 61 points, juste à égalité avec son prédécesseur Sol et derrière Claude Fable avec 5,1 points de 66.","category":"Industrie","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"La performance de benchmark GPT-6 Astra est divisée, l’efficacité surhumaine d’ARC-AGI-3 permet à Chollet de prédire l’AGI à l’avance - Aioga Actualités IA","description":"Les conclusions des benchmarks pour GPT-6 Astra sont contradictoires : Epoch AI le classe premier parmi 267 modèles avec 169 points, tandis qu’Artificial Analysis lui donne 61 poin...","url":"https://www.aioga.com/fr/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:52.484Z"},"de":{"title":"GPT-6 Astra Benchmark-Ergebnisse widersprüchlich, ARC-AGI-3 Effizienz übertrifft Menschen und lässt Chollet AGI-Prognose vorziehen","summary":"Die Benchmark-Ergebnisse von GPT-6 Astra widersprechen sich: Epoch AI bewertet es mit 169 Punkten als Spitzenreiter unter 267 Modellen, während Artificial Analysis nur 61 Punkte vergibt, gleichauf mit dem Vorgänger Sol, aber hinter Claude Fable 5.1 mit 66 Punkten.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra Benchmark-Ergebnisse widersprüchlich, ARC-AGI-3 Effizienz übertrifft Menschen und lässt Chollet AGI-Prognose vorziehen - Aioga KI-News","description":"Die Benchmark-Ergebnisse von GPT-6 Astra widersprechen sich: Epoch AI bewertet es mit 169 Punkten als Spitzenreiter unter 267 Modellen, während Artificial Analysis nur 61 Punkte ve...","url":"https://www.aioga.com/de/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:03:49.683Z"},"pt-BR":{"title":"Desempenho do benchmark GPT-6 Astra divergente, eficiência do ARC-AGI-3 supera humanos e faz Chollet prever AGI antecipadamente","summary":"As conclusões do benchmark do GPT-6 Astra são contraditórias: Epoch AI classificou-o em 169 pontos como o melhor de 267 modelos, enquanto Artificial Analysis deu 61 pontos, igualando a geração anterior Sol e ficando atrás do Claude Fable 5.1 com 66 pontos.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Desempenho do benchmark GPT-6 Astra divergente, eficiência do ARC-AGI-3 supera humanos e faz Chollet prever AGI antecipadamente - Aioga Notícias de IA","description":"As conclusões do benchmark do GPT-6 Astra são contraditórias: Epoch AI classificou-o em 169 pontos como o melhor de 267 modelos, enquanto Artificial Analysis deu 61 pontos, igualan...","url":"https://www.aioga.com/pt-BR/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:00.085Z"},"ru":{"title":"Показатели GPT-6 Astra разделены, сверхчеловеческая эффективность ARC-AGI-3 позволяет Шолле заранее предсказывать AGI","summary":"Выводы по эталон GPT-6 Astra противоречивы: Epoch AI занимает первое место среди 267 моделей с 169 баллами, а Artificial Analysis — 61, что совпадает с предшественником Sol и уступает Claude Fable 5.1 с 66 баллами.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Показатели GPT-6 Astra разделены, сверхчеловеческая эффективность ARC-AGI-3 позволяет Шолле заранее предсказывать AGI - Aioga Новости ИИ","description":"Выводы по эталон GPT-6 Astra противоречивы: Epoch AI занимает первое место среди 267 моделей с 169 баллами, а Artificial Analysis — 61, что совпадает с предшественником Sol и уступ...","url":"https://www.aioga.com/ru/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:00.701Z"},"ar":{"title":"أداء اختبار GPT-6 Astra منقسم، وكفاءة ARC-AGI-3 الخارقة تسمح لتشوليه بالتوقع المسبق للذكاء الاصطناعي العام","summary":"استنتاجات المعيار لجهاز GPT-6 Astra متناقضة: حيث يحتل Epoch AI النسخة الأولى من بين 267 نموذجا ب 169 نقطة، بينما يمنحه Artificial Analysis المعدل 61، متعادلا مع سلفه Sol وخلف Claude Fable ب 5.1 ب 66 نقطة.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"أداء اختبار GPT-6 Astra منقسم، وكفاءة ARC-AGI-3 الخارقة تسمح لتشوليه بالتوقع المسبق للذكاء الاصطناعي العام - Aioga أخبار الذكاء الاصطناعي","description":"استنتاجات المعيار لجهاز GPT-6 Astra متناقضة: حيث يحتل Epoch AI النسخة الأولى من بين 267 نموذجا ب 169 نقطة، بينما يمنحه Artificial Analysis المعدل 61، متعادلا مع سلفه Sol وخلف Claud...","url":"https://www.aioga.com/ar/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:09.905Z"},"hi":{"title":"GPT-6 एस्ट्रा बेंचमार्क प्रदर्शन विभाजित है, ARC-AGI-3 की अलौकिक दक्षता चॉलेट को समय से पहले AGI की भविष्यवाणी करने की अनुमति देती है","summary":"GPT-6 एस्ट्रा के लिए बेंचमार्क निष्कर्ष विरोधाभासी हैं: एपोच एआई इसे 267 अंकों के साथ 169 मॉडलों में पहले स्थान पर रखता है, जबकि आर्टिफिशियल एनालिसिस इसे 61 देता है, जो अपने पूर्ववर्ती सोल के साथ और क्लाउड फैबल के 5.1 अंकों के साथ 66 से पीछे है।","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 एस्ट्रा बेंचमार्क प्रदर्शन विभाजित है, ARC-AGI-3 की अलौकिक दक्षता चॉलेट को समय से पहले AGI की भविष्यवाणी करने की अनुमति देती है - Aioga AI समाचार","description":"GPT-6 एस्ट्रा के लिए बेंचमार्क निष्कर्ष विरोधाभासी हैं: एपोच एआई इसे 267 अंकों के साथ 169 मॉडलों में पहले स्थान पर रखता है, जबकि आर्टिफिशियल एनालिसिस इसे 61 देता है, जो अपने पूर्वव...","url":"https://www.aioga.com/hi/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:08.943Z"},"it":{"title":"Le prestazioni del benchmark GPT-6 Astra sono divise, l'efficienza sovrumana di ARC-AGI-3 permette a Chollet di prevedere l'AGI in anticipo","summary":"Le conclusioni dei benchmark per GPT-6 Astra sono contraddittorie: Epoch AI lo classifica primo tra 267 modelli con 169 punti, mentre Artificial Analysis gli dà 61, appena a pari merito con il predecessore Sol e dietro il 5,1 di Claude Fable con 66 punti.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Le prestazioni del benchmark GPT-6 Astra sono divise, l'efficienza sovrumana di ARC-AGI-3 permette a Chollet di prevedere l'AGI in anticipo - Aioga Notizie IA","description":"Le conclusioni dei benchmark per GPT-6 Astra sono contraddittorie: Epoch AI lo classifica primo tra 267 modelli con 169 punti, mentre Artificial Analysis gli dà 61, appena a pari m...","url":"https://www.aioga.com/it/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:18.905Z"},"nl":{"title":"GPT-6 Astra benchmarkresultaten zijn verdeeld, ARC-AGI-3 efficiëntie overtreft menselijk niveau waardoor Chollet vroegtijdige AGI-voorspelling doet","summary":"De benchmarkresultaten van GPT-6 Astra zijn tegenstrijdig: Epoch AI rangschikt het model als beste van 267 modellen met 169 punten, terwijl Artificial Analysis slechts 61 punten geeft, gelijk aan de vorige generatie Sol en onder de 66 punten van Claude Fable 5.1.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra benchmarkresultaten zijn verdeeld, ARC-AGI-3 efficiëntie overtreft menselijk niveau waardoor Chollet vroegtijdige AGI-voorspelling doet - Aioga AI-nieuws","description":"De benchmarkresultaten van GPT-6 Astra zijn tegenstrijdig: Epoch AI rangschikt het model als beste van 267 modellen met 169 punten, terwijl Artificial Analysis slechts 61 punten ge...","url":"https://www.aioga.com/nl/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:16.993Z"},"tr":{"title":"GPT-6 Astra kıyaslama performansı bölünmüş, ARC-AGI-3'ün insanüstü verimliliği sayesinde Chollet AG'yi önceden tahmin etmesini sağlıyor","summary":"GPT-6 Astra'nın kıyaslama sonuçları çelişkili: Epoch AI, 267 model arasında 169 puanla birinci sırada yer alırken, Artificial Analysis 61 puan veriyor; bu da selefi Sol ile eşit ve Claude Fable'ın 66 puanla 5.1'inin gerisinde yer alıyor.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"GPT-6 Astra kıyaslama performansı bölünmüş, ARC-AGI-3'ün insanüstü verimliliği sayesinde Chollet AG'yi önceden tahmin etmesini sağlıyor - Aioga AI Haberleri","description":"GPT-6 Astra'nın kıyaslama sonuçları çelişkili: Epoch AI, 267 model arasında 169 puanla birinci sırada yer alırken, Artificial Analysis 61 puan veriyor; bu da selefi Sol ile eşit ve...","url":"https://www.aioga.com/tr/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:27.864Z"},"vi":{"title":"Hiệu suất benchmark GPT-6 Astra được chia đều, hiệu suất siêu phàm của ARC-AGI-3 cho phép Chollet dự đoán AGI trước thời gian","summary":"Kết luận benchmark cho GPT-6 Astra lại mâu thuẫn: Epoch AI xếp nó đứng đầu trong số 267 mô hình với 169 điểm, trong khi Artificial Analysis cho 61 điểm, chỉ bằng với người tiền nhiệm Sol và sau Claude Fable với 5.1 với 66 điểm.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Hiệu suất benchmark GPT-6 Astra được chia đều, hiệu suất siêu phàm của ARC-AGI-3 cho phép Chollet dự đoán AGI trước thời gian - Tin tức AI Aioga","description":"Kết luận benchmark cho GPT-6 Astra lại mâu thuẫn: Epoch AI xếp nó đứng đầu trong số 267 mô hình với 169 điểm, trong khi Artificial Analysis cho 61 điểm, chỉ bằng với người tiền nhi...","url":"https://www.aioga.com/vi/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:27.954Z"},"id":{"title":"Kinerja benchmark GPT-6 Astra terbagi, efisiensi superhuman ARC-AGI-3 memungkinkan Chollet memprediksi AGI sebelumnya","summary":"Kesimpulan benchmark untuk GPT-6 Astra bertentangan: Epoch AI menempatkannya di urutan pertama dari 267 model dengan 169 poin, sementara Artificial Analysis memberikannya 61, hanya setara dengan pendahulunya Sol dan di belakang Claude Fable dengan 5,1 dengan 66 poin.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Kinerja benchmark GPT-6 Astra terbagi, efisiensi superhuman ARC-AGI-3 memungkinkan Chollet memprediksi AGI sebelumnya - Berita AI Aioga","description":"Kesimpulan benchmark untuk GPT-6 Astra bertentangan: Epoch AI menempatkannya di urutan pertama dari 267 model dengan 169 poin, sementara Artificial Analysis memberikannya 61, hanya...","url":"https://www.aioga.com/id/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:36.863Z"},"th":{"title":"ประสิทธิภาพการเปรียบเทียบ GPT-6 Astra ถูกแบ่งออก ประสิทธิภาพเหนือมนุษย์ของ ARC-AGI-3 ช่วยให้ Chollet สามารถทํานาย AGI ล่วงหน้าได้","summary":"ข้อสรุปจาก GPT-6 Astra ขัดแย้งกัน: Epoch AI จัดอันดับเป็นอันดับหนึ่งในบรรดา 267 รุ่นที่ได้ 169 คะแนน ขณะที่ Artificial Analysis ให้คะแนน 61 เท่ากับรุ่นก่อนหน้า Sol และรองจาก Claude Fable ที่ 5.1 ที่ได้ 66 คะแนน","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"ประสิทธิภาพการเปรียบเทียบ GPT-6 Astra ถูกแบ่งออก ประสิทธิภาพเหนือมนุษย์ของ ARC-AGI-3 ช่วยให้ Chollet สามารถทํานาย AGI ล่วงหน้าได้ - ข่าว AI Aioga","description":"ข้อสรุปจาก GPT-6 Astra ขัดแย้งกัน: Epoch AI จัดอันดับเป็นอันดับหนึ่งในบรรดา 267 รุ่นที่ได้ 169 คะแนน ขณะที่ Artificial Analysis ให้คะแนน 61 เท่ากับรุ่นก่อนหน้า Sol และรองจาก Claude...","url":"https://www.aioga.com/th/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:36.823Z"},"pl":{"title":"Wydajność benchmarku GPT-6 Astra jest podzielona, nadludzka efektywność ARC-AGI-3 pozwala Chollet przewidywać AGI z wyprzedzeniem","summary":"Wnioski z benchmarków GPT-6 Astra są sprzeczne: Epoch AI plasuje ją na pierwszym miejscu spośród 267 modeli z 169 punktami, podczas gdy Artificial Analysis daje 61, wyrównując się poprzednikowi Sol i za Claude'em Fable'em z 5,1 punktami (66).","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Wydajność benchmarku GPT-6 Astra jest podzielona, nadludzka efektywność ARC-AGI-3 pozwala Chollet przewidywać AGI z wyprzedzeniem - Aioga Wiadomości AI","description":"Wnioski z benchmarków GPT-6 Astra są sprzeczne: Epoch AI plasuje ją na pierwszym miejscu spośród 267 modeli z 169 punktami, podczas gdy Artificial Analysis daje 61, wyrównując się...","url":"https://www.aioga.com/pl/news/cmtmvk9590ccdromyrrae80ir/","contentTranslated":true,"sourceHash":"d0e6278c1b017393","translatedAt":"2026-09-04T12:04:45.796Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":"/page-visuals/topic-timeline.png"}}