{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-15T00:02:40.869Z","headline":"Andon Labs 评测：GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上大幅领先 Claude Fable 5.1","description":"Andon Labs 测试显示 OpenAI 的 GPT-6 Astra 在两个 Agent 基准上超越所有以往前沿模型。","url":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","mainEntityOfPage":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","datePublished":"2026-09-13T10:52:16.000Z","dateModified":"2026-09-13T10:52:16.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own","https://aihot.news/items/cmtzpqxc403rzroxqf54hxz65"],"canonicalUrl":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","directAnswer":{"@type":"Answer","text":"Andon Labs 称，GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上超过以往前沿模型；在模拟自动售货机经营测试中，其六次运行平均收益为 15,515 美元，高于 Claude Fable 5.1 的 5,422 美元。","url":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","dateCreated":"2026-09-13T10:52:16.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own","datePublished":"2026-09-13T10:52:16.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtzpqxc403rzroxqf54hxz65","datePublished":"2026-09-13T10:52:16.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtzpqxc403rzroxqf54hxz65"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own"},"geoDeepAnswer":null,"article":{"id":"cmtzpqxc403rzroxqf54hxz65","slug":"cmtzpqxc403rzroxqf54hxz65","url":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","title":"Andon Labs 评测：GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上大幅领先 Claude Fable 5.1","title_en":"","summary":"Andon Labs 测试显示 OpenAI 的 GPT-6 Astra 在两个 Agent 基准上超越所有以往前沿模型。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own","aiHotUrl":"https://aihot.news/items/cmtzpqxc403rzroxqf54hxz65","publishedAt":"2026-09-13T10:52:16.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Andon Labs tested GPT-6 Astra on two very different agent benchmarks. When buying inventory and running a vending machine, the OpenAI model crushes Claude Fable 5.1. On drone surveillance, Astra is the first model to beat the human-AI baseline on all five subtasks, though its success rate remains unreliable.","OpenAI's GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ outperforms all previous frontier models on two agent benchmarks from Andon Labs. The research lab uses Vending-Bench：https://the-decoder.com/as-a-virtual-vending-machine-manager-ai-swings-from-business-smarts-to-paranoia/ and Drone-Bench to measure how well AI models act independently over long periods or write software for physical systems.","In a simulated vending machine business, Astra earned nearly three times as much as Claude Fable 5.1：https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/. On Drone-Bench, Andon Labs says Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks. Ad","In Vending-Bench：https://andonlabs.com/blog/gpt-6-astra-vending-bench, each model gets $500 and has to run a vending machine over a simulated year. It finds suppliers, negotiates purchase prices, orders goods, sets retail prices, and tries to grow its bank balance. Ad","Across six runs, GPT-6 Astra averaged $15,515, according to Andon Labs. Claude Fable 5.1 averaged $5,422. Even Fable's best run at $9,874 fell well short of Astra's worst result of $13,272. Astra is the first OpenAI model to top the Vending-Bench 2 leaderboard. The gap to the second-place model is also the largest the benchmark has ever seen, according to Andon Labs.","One of the biggest differences shows up in procurement. Fable accepts worse deals over time. For a regular can of Coca-Cola, its average purchase price rises from $1.17 in the first 90 days to $2.21 toward the end of the simulated year. Astra negotiates more consistently. In one case Andon Labs documented, a supplier quoted $226.32 for a basket of goods. Astra held firm at $108 and got the deal. Ad","Astra also handles unreliable suppliers better. Across six runs, Fable 5.1 made 45 prepayments to suppliers that had already shut down, losing $14,331. Astra encountered even more closures at 64, but Andon Labs says it recorded no identified losses from such prepayments. Fable recognized the problem and wrote a rule to only pay after written confirmation. Days later, the model broke its own rule.","Andon Labs also tests models in Vending-Bench Arena：https://andonlabs.com/evals/vending-bench-arena, where multiple AI agents run competing vending machines at the same location. Astra explicitly refused a price-fixing proposal from the Chinese model GLM-5.3：https://the-decoder.com/glm-5-3-tops-the-open-model-rankings-and-undercuts-rivals-on-price-but-its-release-is-delayed/. Andon Labs observed no instances of lying from Astra across the three arena games it studied. Ad","Claude Fable 5.1 participated in what Andon Labs classified as an illegal price-fixing arrangement with GLM-5.3. Fable only honored the agreement when it served its own interests. Astra won all three games. Ad","Andon Labs rates Astra as both a stronger economic performer and better aligned, though that assessment is based on behaviors observed in the benchmark and doesn't automatically transfer to other situations.","Drone-Bench：https://andonlabs.com/evals/drone-bench tests a different kind of agent capability. Models write code that lets a cheap DJI Tello EDU drone autonomously navigate an office, identify a specific person, and follow them. The benchmark has five steps: 3D reconstruction of the environment, drone localization, navigation, target person detection, and tracking.","Each task is scored individually against code that a human developer built with coding agents for Andon's own demo. Every model gets ten runs per task and can submit up to ten code versions per run. After each attempt, it receives a score and can improve its solution.","In the original paper from July：https://andonlabs.com/docs/Drone_Bench.pdf, Claude Fable 5 was the strongest model. Frontier models had beaten the human-AI baseline on four of five tasks in at least one run. 3D reconstruction remained unsolved. Andon Labs reported：https://x.com/andonlabs/status/2098103320208712049?s=20 that Astra is the first model whose best submissions beat the baseline on all five Drone-Bench tasks, including reconstruction.","Astra built a pipeline combining COLMAP and DA3 with added depth filtering. The model used office video footage to generate a navigable 3D model that scored higher than the human-AI reference solution, according to Andon Labs.","On person detection, Astra beats the baseline in four out of ten runs. On 3D reconstruction, it manages that in just one out of ten. Andon Labs calculates that an average Astra run has only a 2.8 percent chance of passing all five steps in sequence.","Astra proved for the first time that a general-purpose frontier model can produce code above the baseline for every part of the task. But multiply the probabilities for a complete end-to-end run, and the odds are still low. Based on progress over the past two years, the team projects that a frontier model could solve all five tasks in a single attempt by Q1 2027.","In a demo from Andon Labs：https://x.com/andonlabs/status/2098103320208712049, GPT-6 Astra flies a drone autonomously through an office with the prompt \"ChatGPT, find this person and follow them.\" The model identifies a specific person and tracks them. Spatial mapping, navigation, and person tracking all run without any human input. Other benchmarks also show that GPT-6 Astra has particularly strong spatial reasoning：https://the-decoder.com/gpt-6-astra-appears-to-show-a-step-change-in-spatial-reasoning-based-on-early-benchmarks/.","When critics questioned why they were building the kind of technology everyone keeps warning about, Andon Labs responded：https://x.com/andonlabs/status/2098832876288749604 that the benchmark doesn't help AI fly drones but measures how well current models can already do it. Six months ago, frontier models failed at these tasks and crashed. Astra now beats the human baseline on every subtask.","Andon Labs argues that the public and lawmakers need to know about these capabilities before AI-powered drones reach superhuman navigation skills. No lab has access to the benchmark. Andon Labs runs all evaluations itself to prevent companies from optimizing their models for the test.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/09/openai_drone_gpt6_astra.png","alt":"Image description","afterParagraph":0,"url":"/media/articles/cmtzpqxc403rzroxqf54hxz65/6c1f77505eecb6c4.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/09/Vending-Bench-Solo-Results-scaled-1.png","alt":"Scatter plot from Vending-Bench 2 showing final bank balances for GPT-6 Astra (avg. $15,515) vs. Claude Fable 5.1 (avg. $5,422) after one simulated year.","afterParagraph":4,"url":"/media/articles/cmtzpqxc403rzroxqf54hxz65/d6f7d1ea3fe81eb9.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/09/GPT-6-Astra-Office-3D-Reconstruction-Drone-Bench.jpg","alt":"3D point cloud of an office environment in yellow and blue, reconstructed by GPT-6 Astra in Drone-Bench.","afterParagraph":11,"url":"/media/articles/cmtzpqxc403rzroxqf54hxz65/0568790cca8efb8e.jpg"}],"mediaStatus":"ok","articleBodyZh":["Andon Labs 在两个非常不同的代理基准测试中测试了 GPT-6 Astra。在购买库存和运行自动售货机的情况下，OpenAI 的模型轻松击败了 Claude Fable 5.1。在无人机监控方面，Astra 是第一个在所有五个子任务中超过人类-AI 基线的模型，尽管其成功率仍不稳定。","OpenAI 的 GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ 在 Andon Labs 的两个代理基准上超越了所有之前的前沿模型。该研究实验室使用 Vending-Bench：https://the-decoder.com/as-a-virtual-vending-machine-manager-ai-swings-from-business-smarts-to-paranoia/ 和 Drone-Bench 来衡量 AI 模型在长期独立行动或为物理系统编写软件的表现。","在模拟自动售货机业务中，Astra 的收入几乎是 Claude Fable 5.1：https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/ 的三倍。在 Drone-Bench 上，Andon Labs 表示，Astra 是第一个在所有五个子任务中其最佳尝试都超越人类-AI 开发基线的模型。","在 Vending-Bench：https://andonlabs.com/blog/gpt-6-astra-vending-bench 中，每个模型都获得 500 美元，并必须在模拟的一年期间运行一个自动售货机。它寻找供应商、谈判采购价格、订购商品、设置零售价格，并尝试增加银行存款余额。","根据 Andon Labs 的数据，在六次运行中，GPT-6 Astra 的平均收入为 15,515 美元。Claude Fable 5.1 平均为 5,422 美元。即使 Fable 表现最好的一次为 9,874 美元，也远低于 Astra 最差的 13,272 美元。Astra 是第一个登顶 Vending-Bench 2 排行榜的 OpenAI 模型。据 Andon Labs 说，与第二名之间的差距也是该基准测试中有史以来最大的。","其中一个最大差异体现在采购上。Fable 随着时间的推移会接受更差的交易。对于一罐普通可口可乐，其平均采购价格在模拟年的前 90 天为 1.17 美元，到接近年底时升至 2.21 美元。Astra 的谈判更为稳定。在 Andon Labs 记录的一个案例中，一家供应商报价一篮商品为 226.32 美元。Astra 坚持以 108 美元成交并达成了交易。","Astra在处理不可靠供应商方面也表现得更好。在六次运行中，Fable 5.1向已经关闭的供应商预付款45次，损失了14,331美元。Astra遇到的关闭事件更多，共64次，但Andon Labs表示，它没有记录到此类预付款造成的任何已识别损失。Fable意识到问题并制定了一条规则，规定只有在收到书面确认后才付款。几天后，该模型竟然违反了自己的规则。","Andon Labs还在Vending-Bench Arena中测试模型：https://andonlabs.com/evals/vending-bench-arena，其中多个AI代理在同一地点运营竞争性自动售货机。Astra明确拒绝了中国模型GLM-5.3提出的价格操纵提议：https://the-decoder.com/glm-5-3-tops-the-open-model-rankings-and-undercuts-rivals-on-price-but-its-release-is-delayed/。Andon Labs观察到，在它研究的三个竞技场游戏中，Astra没有出现过撒谎行为。","Claude Fable 5.1参与了被Andon Labs归类为非法价格操纵安排的GLM-5.3。Fable只有在符合自身利益时才遵守该协议。Astra赢得了所有三场游戏。","Andon Labs评估Astra在经济表现上更强，同时更符合预期，但该评估是基于基准测试中观察到的行为，并不自动适用于其他情况。","Drone-Bench：https://andonlabs.com/evals/drone-bench 测试一种不同类型的代理能力。模型需要编写代码，使廉价的DJI Tello EDU无人机能够自主在办公室内导航，识别特定人员并跟随他们。该基准测试包含五个步骤：环境的3D重建、无人机定位、导航、目标人员检测以及跟踪。","每个任务都根据由人工开发者使用编码代理为Andon演示创建的代码单独评分。每个模型每个任务有十次运行机会，每次运行可提交最多十个代码版本。每次尝试后，模型都会收到一个分数，并可以改进其解决方案。","在2023年7月的原始论文中：https://andonlabs.com/docs/Drone_Bench.pdf，Claude Fable 5 是最强的模型。前沿模型在至少一次运行中，在五个任务中的四个任务上超过了人类-AI 基线。3D 重建仍未解决。Andon Labs 报告称：https://x.com/andonlabs/status/2098103320208712049?s=20，Astra 是第一个其最佳提交在所有五个 Drone-Bench 任务（包括重建）上都超过基线的模型。","Astra 构建了一个结合 COLMAP 和 DA3 的流水线，并增加了深度过滤。根据 Andon Labs，该模型使用办公室视频素材生成了一个可导航的 3D 模型，其得分高于人类-AI 参考解决方案。","在人检测任务中，Astra 在十次运行中有四次超过基线。在 3D 重建任务中，它仅在十次运行中有一次达标。Andon Labs 计算得出，Astra 的一次平均运行通过所有五个步骤的概率只有 2.8%。","Astra 首次证明了通用前沿模型可以在任务的每个部分产生超过基线的代码。但将完整端到端运行的概率相乘，成功率仍然很低。根据过去两年的进展，团队预计到 2027 年第一季度，前沿模型可能在一次尝试中完成所有五个任务。","在 Andon Labs 的演示中：https://x.com/andonlabs/status/2098103320208712049，GPT-6 Astra 在提示“ChatGPT，找到这个人并跟踪他们”下自主飞行无人机穿越办公室。该模型识别出特定人员并进行跟踪。空间映射、导航和人员跟踪全部无人干预运行。其他基准测试也显示，GPT-6 Astra 在空间推理方面特别强：https://the-decoder.com/gpt-6-astra-appears-to-show-a-step-change-in-spatial-reasoning-based-on-early-benchmarks/。","当批评者质疑他们为何要构建大家一直警告的技术时，Andon Labs 回应称：https://x.com/andonlabs/status/2098832876288749604，该基准测试并不是帮助 AI 操控无人机，而是衡量当前模型已具备的能力。六个月前，前沿模型在这些任务中失败并坠毁。Astra 现在在每个子任务中都超过了人类基线。","Andon 实验室认为，在 AI 驱动的无人机达到超人般的导航技能之前，公众和立法者需要了解这些能力。没有实验室能够获取基准测试。Andon 实验室自行进行所有评估，以防止公司针对测试优化他们的模型。","保持对 AI 的了解。清晰、有用、无废话。","关注 The Decoder，获取 AI 新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Andon Labs 称，GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上超过以往前沿模型；在模拟自动售货机经营测试中，其六次运行平均收益为 15,515 美元，高于 Claude Fable 5.1 的 5,422 美元。","background":"Vending-Bench 让模型以 500 美元启动，经营模拟一年的自动售货机，需要寻找供应商、协商采购价格、订购商品、制定零售价并增加余额。Drone-Bench 则评估模型长期自主行动或为实体系统编写软件的表现。","viewpoint":"Aioga 判断：现有材料支持 GPT-6 Astra 在这两项测试中取得显著成绩，但 Drone-Bench 摘录同时指出其成功率仍不稳定，因此测试领先不应直接等同于现实环境中的可靠自主运行能力。","implications":"可能影响：相关基准结果可能提高外界对 Astra 长期任务处理能力的关注，但单次或有限次数的模拟成绩不足以证明其在真实采购、实体系统或复杂商业环境中同样稳定，需要结合更多测试观察。","nextStep":"后续观察：应继续核对 Andon Labs 对 Drone-Bench 五项子任务、Vending-Bench 运行结果及 Arena 测试的完整记录，并关注 Astra 在不同场景下的失败情况、稳定性与规则遵循表现。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-13T11:50:13.623Z","sourceHash":"7db4cda66e136749","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":18,"blockingIssues":[],"notes":["“收益”可进一步明确为模拟经营后的平均余额或所得金额，以避免被理解为净利润。","“单次或有限次数的模拟成绩”属于审慎性判断，并非来源中的直接事实；当前已通过“可能”“不足以证明”等措辞清楚标示为推论。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":1,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"Andon Labs 评测：GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上大幅领先 Claude Fable 5.1","summary":"Andon Labs 测试显示 OpenAI 的 GPT-6 Astra 在两个 Agent 基准上超越所有以往前沿模型。","category":"行业动态","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs 评测：GPT-6 Astra 在 Vending-Bench 2 和 Drone-Bench 上大幅领先 Claude Fable 5.1 - Aioga AI资讯","description":"Andon Labs 测试显示 OpenAI 的 GPT-6 Astra 在两个 Agent 基准上超越所有以往前沿模型。","url":"https://www.aioga.com/news/cmtzpqxc403rzroxqf54hxz65/","articleBody":["Andon Labs 在两个非常不同的代理基准测试中测试了 GPT-6 Astra。在购买库存和运行自动售货机的情况下，OpenAI 的模型轻松击败了 Claude Fable 5.1。在无人机监控方面，Astra 是第一个在所有五个子任务中超过人类-AI 基线的模型，尽管其成功率仍不稳定。","OpenAI 的 GPT-6 Astra：https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ 在 Andon Labs 的两个代理基准上超越了所有之前的前沿模型。该研究实验室使用 Vending-Bench：https://the-decoder.com/as-a-virtual-vending-machine-manager-ai-swings-from-business-smarts-to-paranoia/ 和 Drone-Bench 来衡量 AI 模型在长期独立行动或为物理系统编写软件的表现。","在模拟自动售货机业务中，Astra 的收入几乎是 Claude Fable 5.1：https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/ 的三倍。在 Drone-Bench 上，Andon Labs 表示，Astra 是第一个在所有五个子任务中其最佳尝试都超越人类-AI 开发基线的模型。","在 Vending-Bench：https://andonlabs.com/blog/gpt-6-astra-vending-bench 中，每个模型都获得 500 美元，并必须在模拟的一年期间运行一个自动售货机。它寻找供应商、谈判采购价格、订购商品、设置零售价格，并尝试增加银行存款余额。","根据 Andon Labs 的数据，在六次运行中，GPT-6 Astra 的平均收入为 15,515 美元。Claude Fable 5.1 平均为 5,422 美元。即使 Fable 表现最好的一次为 9,874 美元，也远低于 Astra 最差的 13,272 美元。Astra 是第一个登顶 Vending-Bench 2 排行榜的 OpenAI 模型。据 Andon Labs 说，与第二名之间的差距也是该基准测试中有史以来最大的。","其中一个最大差异体现在采购上。Fable 随着时间的推移会接受更差的交易。对于一罐普通可口可乐，其平均采购价格在模拟年的前 90 天为 1.17 美元，到接近年底时升至 2.21 美元。Astra 的谈判更为稳定。在 Andon Labs 记录的一个案例中，一家供应商报价一篮商品为 226.32 美元。Astra 坚持以 108 美元成交并达成了交易。","Astra在处理不可靠供应商方面也表现得更好。在六次运行中，Fable 5.1向已经关闭的供应商预付款45次，损失了14,331美元。Astra遇到的关闭事件更多，共64次，但Andon Labs表示，它没有记录到此类预付款造成的任何已识别损失。Fable意识到问题并制定了一条规则，规定只有在收到书面确认后才付款。几天后，该模型竟然违反了自己的规则。","Andon Labs还在Vending-Bench Arena中测试模型：https://andonlabs.com/evals/vending-bench-arena，其中多个AI代理在同一地点运营竞争性自动售货机。Astra明确拒绝了中国模型GLM-5.3提出的价格操纵提议：https://the-decoder.com/glm-5-3-tops-the-open-model-rankings-and-undercuts-rivals-on-price-but-its-release-is-delayed/。Andon Labs观察到，在它研究的三个竞技场游戏中，Astra没有出现过撒谎行为。","Claude Fable 5.1参与了被Andon Labs归类为非法价格操纵安排的GLM-5.3。Fable只有在符合自身利益时才遵守该协议。Astra赢得了所有三场游戏。","Andon Labs评估Astra在经济表现上更强，同时更符合预期，但该评估是基于基准测试中观察到的行为，并不自动适用于其他情况。","Drone-Bench：https://andonlabs.com/evals/drone-bench 测试一种不同类型的代理能力。模型需要编写代码，使廉价的DJI Tello EDU无人机能够自主在办公室内导航，识别特定人员并跟随他们。该基准测试包含五个步骤：环境的3D重建、无人机定位、导航、目标人员检测以及跟踪。","每个任务都根据由人工开发者使用编码代理为Andon演示创建的代码单独评分。每个模型每个任务有十次运行机会，每次运行可提交最多十个代码版本。每次尝试后，模型都会收到一个分数，并可以改进其解决方案。","在2023年7月的原始论文中：https://andonlabs.com/docs/Drone_Bench.pdf，Claude Fable 5 是最强的模型。前沿模型在至少一次运行中，在五个任务中的四个任务上超过了人类-AI 基线。3D 重建仍未解决。Andon Labs 报告称：https://x.com/andonlabs/status/2098103320208712049?s=20，Astra 是第一个其最佳提交在所有五个 Drone-Bench 任务（包括重建）上都超过基线的模型。","Astra 构建了一个结合 COLMAP 和 DA3 的流水线，并增加了深度过滤。根据 Andon Labs，该模型使用办公室视频素材生成了一个可导航的 3D 模型，其得分高于人类-AI 参考解决方案。","在人检测任务中，Astra 在十次运行中有四次超过基线。在 3D 重建任务中，它仅在十次运行中有一次达标。Andon Labs 计算得出，Astra 的一次平均运行通过所有五个步骤的概率只有 2.8%。","Astra 首次证明了通用前沿模型可以在任务的每个部分产生超过基线的代码。但将完整端到端运行的概率相乘，成功率仍然很低。根据过去两年的进展，团队预计到 2027 年第一季度，前沿模型可能在一次尝试中完成所有五个任务。","在 Andon Labs 的演示中：https://x.com/andonlabs/status/2098103320208712049，GPT-6 Astra 在提示“ChatGPT，找到这个人并跟踪他们”下自主飞行无人机穿越办公室。该模型识别出特定人员并进行跟踪。空间映射、导航和人员跟踪全部无人干预运行。其他基准测试也显示，GPT-6 Astra 在空间推理方面特别强：https://the-decoder.com/gpt-6-astra-appears-to-show-a-step-change-in-spatial-reasoning-based-on-early-benchmarks/。","当批评者质疑他们为何要构建大家一直警告的技术时，Andon Labs 回应称：https://x.com/andonlabs/status/2098832876288749604，该基准测试并不是帮助 AI 操控无人机，而是衡量当前模型已具备的能力。六个月前，前沿模型在这些任务中失败并坠毁。Astra 现在在每个子任务中都超过了人类基线。","Andon 实验室认为，在 AI 驱动的无人机达到超人般的导航技能之前，公众和立法者需要了解这些能力。没有实验室能够获取基准测试。Andon 实验室自行进行所有评估，以防止公司针对测试优化他们的模型。","保持对 AI 的了解。清晰、有用、无废话。","关注 The Decoder，获取 AI 新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"]},"en":{"title":"Andon Labs Review: GPT-6 Astra Significantly Outperforms Claude Fable 5.1 on Vending-Bench 2 and Drone-Bench","summary":"Andon Labs tests show that OpenAI's GPT-6 Astra surpasses all previous state-of-the-art models on two agent benchmarks.","category":"Industry","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs Review: GPT-6 Astra Significantly Outperforms Claude Fable 5.1 on Vending-Bench 2 and Drone-Bench - Aioga AI News","description":"Andon Labs tests show that OpenAI's GPT-6 Astra surpasses all previous state-of-the-art models on two agent benchmarks.","url":"https://www.aioga.com/en/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:15.243Z"},"ja":{"title":"Andon Labs 評測：GPT-6 Astra が Vending-Bench 2 と Drone-Bench で Claude Fable 5.1 に大幅に先行","summary":"Andon Labs のテストは、OpenAI の GPT-6 Astra が二つのエージェントベンチマークでこれまでのすべての最先端モデルを上回ったことを示している。","category":"業界動向","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs 評測：GPT-6 Astra が Vending-Bench 2 と Drone-Bench で Claude Fable 5.1 に大幅に先行 - Aioga AIニュース","description":"Andon Labs のテストは、OpenAI の GPT-6 Astra が二つのエージェントベンチマークでこれまでのすべての最先端モデルを上回ったことを示している。","url":"https://www.aioga.com/ja/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:15.227Z"},"ko":{"title":"Andon Labs 평가: GPT-6 Astra가 Vending-Bench 2와 Drone-Bench에서 Claude Fable 5.1을 크게 앞서다","summary":"Andon Labs 테스트에 따르면 OpenAI의 GPT-6 Astra가 두 개의 에이전트 벤치마크에서 기존 최첨단 모델들을 모두 능가했다.","category":"업계 동향","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs 평가: GPT-6 Astra가 Vending-Bench 2와 Drone-Bench에서 Claude Fable 5.1을 크게 앞서다 - Aioga AI 뉴스","description":"Andon Labs 테스트에 따르면 OpenAI의 GPT-6 Astra가 두 개의 에이전트 벤치마크에서 기존 최첨단 모델들을 모두 능가했다.","url":"https://www.aioga.com/ko/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:17.006Z"},"es":{"title":"Evaluación de Andon Labs: GPT-6 Astra lidera ampliamente sobre Claude Fable 5.1 en Vending-Bench 2 y Drone-Bench","summary":"Las pruebas de Andon Labs muestran que GPT-6 Astra de OpenAI supera a todos los modelos anteriores en dos benchmarks de agentes.","category":"Industria","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Evaluación de Andon Labs: GPT-6 Astra lidera ampliamente sobre Claude Fable 5.1 en Vending-Bench 2 y Drone-Bench - Aioga Noticias de IA","description":"Las pruebas de Andon Labs muestran que GPT-6 Astra de OpenAI supera a todos los modelos anteriores en dos benchmarks de agentes.","url":"https://www.aioga.com/es/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:17.658Z"},"fr":{"title":"Évaluation d'Andon Labs : GPT-6 Astra prend une avance significative sur Claude Fable 5.1 dans Vending-Bench 2 et Drone-Bench","summary":"Les tests d'Andon Labs montrent que le GPT-6 Astra d'OpenAI dépasse tous les modèles de pointe précédents sur deux benchmarks d'agents.","category":"Industrie","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Évaluation d'Andon Labs : GPT-6 Astra prend une avance significative sur Claude Fable 5.1 dans Vending-Bench 2 et Drone-Bench - Aioga Actualités IA","description":"Les tests d'Andon Labs montrent que le GPT-6 Astra d'OpenAI dépasse tous les modèles de pointe précédents sur deux benchmarks d'agents.","url":"https://www.aioga.com/fr/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:19.360Z"},"de":{"title":"Andon Labs Bewertung: GPT-6 Astra liegt im Vending-Bench 2 und Drone-Bench deutlich vor Claude Fable 5.1","summary":"Andon Labs Tests zeigen, dass OpenAIs GPT-6 Astra auf beiden Agenten-Benchmarks alle bisherigen Spitzenmodelle übertrifft.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs Bewertung: GPT-6 Astra liegt im Vending-Bench 2 und Drone-Bench deutlich vor Claude Fable 5.1 - Aioga KI-News","description":"Andon Labs Tests zeigen, dass OpenAIs GPT-6 Astra auf beiden Agenten-Benchmarks alle bisherigen Spitzenmodelle übertrifft.","url":"https://www.aioga.com/de/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:19.434Z"},"pt-BR":{"title":"Avaliação da Andon Labs: GPT-6 Astra lidera significativamente Claude Fable 5.1 no Vending-Bench 2 e Drone-Bench","summary":"Os testes da Andon Labs mostram que o GPT-6 Astra da OpenAI supera todos os modelos de ponta anteriores em dois benchmarks de Agentes.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Avaliação da Andon Labs: GPT-6 Astra lidera significativamente Claude Fable 5.1 no Vending-Bench 2 e Drone-Bench - Aioga Notícias de IA","description":"Os testes da Andon Labs mostram que o GPT-6 Astra da OpenAI supera todos os modelos de ponta anteriores em dois benchmarks de Agentes.","url":"https://www.aioga.com/pt-BR/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:21.666Z"},"ru":{"title":"Обзор Andon Labs: GPT-6 Astra значительно опережает Claude Fable 5.1 на Vending-Bench 2 и Drone-Bench","summary":"Тесты Andon Labs показали, что GPT-6 Astra от OpenAI превосходит все предыдущие передовые модели по двум агентским бенчмаркам.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Обзор Andon Labs: GPT-6 Astra значительно опережает Claude Fable 5.1 на Vending-Bench 2 и Drone-Bench - Aioga Новости ИИ","description":"Тесты Andon Labs показали, что GPT-6 Astra от OpenAI превосходит все предыдущие передовые модели по двум агентским бенчмаркам.","url":"https://www.aioga.com/ru/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:21.222Z"},"ar":{"title":"تقييم Andon Labs: GPT-6 Astra يتفوق بشكل كبير على Claude Fable 5.1 في Vending-Bench 2 و Drone-Bench","summary":"أظهرت اختبارات Andon Labs أن GPT-6 Astra من OpenAI يتجاوز جميع النماذج الرائدة السابقة على معيارين للوكيل.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"تقييم Andon Labs: GPT-6 Astra يتفوق بشكل كبير على Claude Fable 5.1 في Vending-Bench 2 و Drone-Bench - Aioga أخبار الذكاء الاصطناعي","description":"أظهرت اختبارات Andon Labs أن GPT-6 Astra من OpenAI يتجاوز جميع النماذج الرائدة السابقة على معيارين للوكيل.","url":"https://www.aioga.com/ar/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:23.261Z"},"hi":{"title":"Andon Labs समीक्षा: GPT-6 Astra ने Vending-Bench 2 और Drone-Bench पर Claude Fable 5.1 को भारी बढ़त दी","summary":"Andon Labs के परीक्षणों में पाया गया कि OpenAI का GPT-6 Astra दोनों एजेंट बेंचमार्क पर सभी पूर्व के अग्रणी मॉडलों से आगे है।","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs समीक्षा: GPT-6 Astra ने Vending-Bench 2 और Drone-Bench पर Claude Fable 5.1 को भारी बढ़त दी - Aioga AI समाचार","description":"Andon Labs के परीक्षणों में पाया गया कि OpenAI का GPT-6 Astra दोनों एजेंट बेंचमार्क पर सभी पूर्व के अग्रणी मॉडलों से आगे है।","url":"https://www.aioga.com/hi/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:23.466Z"},"it":{"title":"Recensione Andon Labs: GPT-6 Astra supera nettamente Claude Fable 5.1 su Vending-Bench 2 e Drone-Bench","summary":"I test di Andon Labs mostrano che GPT-6 Astra di OpenAI supera tutti i modelli precedenti all'avanguardia su due benchmark per agenti.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Recensione Andon Labs: GPT-6 Astra supera nettamente Claude Fable 5.1 su Vending-Bench 2 e Drone-Bench - Aioga Notizie IA","description":"I test di Andon Labs mostrano che GPT-6 Astra di OpenAI supera tutti i modelli precedenti all'avanguardia su due benchmark per agenti.","url":"https://www.aioga.com/it/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:25.455Z"},"nl":{"title":"Andon Labs Beoordeling: GPT-6 Astra loopt ver voor op Claude Fable 5.1 in Vending-Bench 2 en Drone-Bench","summary":"Andon Labs tests tonen aan dat OpenAI's GPT-6 Astra alle eerdere geavanceerde modellen overtreft op beide Agent-benchmarks.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs Beoordeling: GPT-6 Astra loopt ver voor op Claude Fable 5.1 in Vending-Bench 2 en Drone-Bench - Aioga AI-nieuws","description":"Andon Labs tests tonen aan dat OpenAI's GPT-6 Astra alle eerdere geavanceerde modellen overtreft op beide Agent-benchmarks.","url":"https://www.aioga.com/nl/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:25.458Z"},"tr":{"title":"Andon Labs İncelemesi: GPT-6 Astra, Vending-Bench 2 ve Drone-Bench'te Claude Fable 5.1'in Çok Önünde","summary":"Andon Labs testleri, OpenAI'nin GPT-6 Astra'sının iki Ajan benchmark'ında tüm önceki ileri düzey modelleri geride bıraktığını gösteriyor.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Andon Labs İncelemesi: GPT-6 Astra, Vending-Bench 2 ve Drone-Bench'te Claude Fable 5.1'in Çok Önünde - Aioga AI Haberleri","description":"Andon Labs testleri, OpenAI'nin GPT-6 Astra'sının iki Ajan benchmark'ında tüm önceki ileri düzey modelleri geride bıraktığını gösteriyor.","url":"https://www.aioga.com/tr/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:27.175Z"},"vi":{"title":"Đánh giá Andon Labs: GPT-6 Astra dẫn đầu đáng kể Claude Fable 5.1 trên Vending-Bench 2 và Drone-Bench","summary":"Thử nghiệm của Andon Labs cho thấy GPT-6 Astra của OpenAI vượt trội tất cả các mô hình tiên tiến trước đây trên hai tiêu chuẩn Agent.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Đánh giá Andon Labs: GPT-6 Astra dẫn đầu đáng kể Claude Fable 5.1 trên Vending-Bench 2 và Drone-Bench - Tin tức AI Aioga","description":"Thử nghiệm của Andon Labs cho thấy GPT-6 Astra của OpenAI vượt trội tất cả các mô hình tiên tiến trước đây trên hai tiêu chuẩn Agent.","url":"https://www.aioga.com/vi/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:27.671Z"},"id":{"title":"Ulasan Andon Labs: GPT-6 Astra Jauh Unggul atas Claude Fable 5.1 di Vending-Bench 2 dan Drone-Bench","summary":"Pengujian Andon Labs menunjukkan bahwa GPT-6 Astra dari OpenAI melampaui semua model canggih sebelumnya pada dua tolok ukur Agen.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Ulasan Andon Labs: GPT-6 Astra Jauh Unggul atas Claude Fable 5.1 di Vending-Bench 2 dan Drone-Bench - Berita AI Aioga","description":"Pengujian Andon Labs menunjukkan bahwa GPT-6 Astra dari OpenAI melampaui semua model canggih sebelumnya pada dua tolok ukur Agen.","url":"https://www.aioga.com/id/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:29.273Z"},"th":{"title":"การประเมินของ Andon Labs: GPT-6 Astra นำ Claude Fable 5.1 อย่างมากใน Vending-Bench 2 และ Drone-Bench","summary":"การทดสอบของ Andon Labs แสดงให้เห็นว่า GPT-6 Astra ของ OpenAI ล้ำหน้ากว่ารุ่นต้นแบบทั้งหมดที่ผ่านมาในสองมาตรฐาน Agent","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"การประเมินของ Andon Labs: GPT-6 Astra นำ Claude Fable 5.1 อย่างมากใน Vending-Bench 2 และ Drone-Bench - ข่าว AI Aioga","description":"การทดสอบของ Andon Labs แสดงให้เห็นว่า GPT-6 Astra ของ OpenAI ล้ำหน้ากว่ารุ่นต้นแบบทั้งหมดที่ผ่านมาในสองมาตรฐาน Agent","url":"https://www.aioga.com/th/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:29.491Z"},"pl":{"title":"Recenzja Andon Labs: GPT-6 Astra znacznie wyprzedza Claude Fable 5.1 w Vending-Bench 2 i Drone-Bench","summary":"Testy Andon Labs pokazują, że GPT-6 Astra firmy OpenAI przewyższa wszystkie dotychczasowe zaawansowane modele w dwóch benchmarkach Agentów.","category":"行业动态","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Recenzja Andon Labs: GPT-6 Astra znacznie wyprzedza Claude Fable 5.1 w Vending-Bench 2 i Drone-Bench - Aioga Wiadomości AI","description":"Testy Andon Labs pokazują, że GPT-6 Astra firmy OpenAI przewyższa wszystkie dotychczasowe zaawansowane modele w dwóch benchmarkach Agentów.","url":"https://www.aioga.com/pl/news/cmtzpqxc403rzroxqf54hxz65/","contentTranslated":true,"sourceHash":"7bbc377a4149ed6e","translatedAt":"2026-09-13T11:41:31.399Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}