{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T05:20:57.982Z","headline":"METR 推出\"支出视界\"指标衡量AI成本效益","description":"METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。在NanoGPT speedrun测试中，人类每实现1%加速需约16小时工作（折合$2，500），而GPT-5.5和Opus-4.8的支出视界在$0-$3，300之间，但整体自主优化贡献远低于人类$250，000的总投入。","url":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/","mainEntityOfPage":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/","datePublished":"2026-07-27T12:28:06.000Z","dateModified":"2026-07-27T12:28:06.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/metr-introduces-a-new-metric-to-calculate-exactly-when-ai-agents-become-more-expensive-than-humans","https://aihot.virxact.com/items/cms38b6cd08zdro3fkczc5aiq"],"canonicalUrl":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/","dateCreated":"2026-07-27T12:28:06.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/metr-introduces-a-new-metric-to-calculate-exactly-when-ai-agents-become-more-expensive-than-humans","datePublished":"2026-07-27T12:28:06.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/metr-introduces-a-new-metric-to-calculate-exactly-when-ai-agents-become-more-expensive-than-humans"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cms38b6cd08zdro3fkczc5aiq","datePublished":"2026-07-27T12:28:06.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cms38b6cd08zdro3fkczc5aiq"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/metr-introduces-a-new-metric-to-calculate-exactly-when-ai-agents-become-more-expensive-than-humans"},"article":{"id":"cms38b6cd08zdro3fkczc5aiq","slug":"cms38b6cd08zdro3fkczc5aiq","url":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/","title":"METR 推出\"支出视界\"指标衡量AI成本效益","title_en":"METR introduces a new metric to calculate exactly when AI agents become more expensive than humans","summary":"METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。在NanoGPT speedrun测试中，人类每实现1%加速需约16小时工作（折合$2，500），而GPT-5.5和Opus-4.8的支出视界在$0-$3，300之间，但整体自主优化贡献远低于人类$250，000的总投入。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/metr-introduces-a-new-metric-to-calculate-exactly-when-ai-agents-become-more-expensive-than-humans","aiHotUrl":"https://aihot.virxact.com/items/cms38b6cd08zdro3fkczc5aiq","publishedAt":"2026-07-27T12:28:06.000Z","category":"论文研究","score":37,"selected":false,"articleBody":["METR's new metric, the \"expenditure horizon,\" puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture.","One of the biggest questions in AI research is whether AI can accelerate its own development and keep getting better at an increasing pace. That's been hard to measure because it requires comparing very different kinds of costs: human labor, compute for experiments, and the cost of running the AI itself.","Research organization METR proposes a new metric to tackle this: the \"expenditure horizon.\" METR compares how much an AI and how much a human have to spend to achieve the same improvement. The expenditure horizon is the point where both cost the same. Below that budget, the AI is the better deal. Above it, the human works cheaper.","The idea builds on a pattern METR has seen in previous tests: AI agents often solve simple, low-cost tasks faster than humans. But as budgets grow and tasks get harder, they fall behind.","Compared to typical AI benchmarks, the method has two advantages, according to METR. First, it doesn't just give a pass-or-fail verdict. Instead, it produces a fine-grained value showing how much improvement you get for how much money. Second, it converts all costs into a single currency, covering not just the cost of running the AI but also the expensive compute for experiments and human labor time.","METR chose the NanoGPT speedrun：https://github.com/kellerjordan/modded-nanogpt as its testing ground. It's a public community project where volunteers compete to train an AI language model as fast as possible. The task stays the same; only the training approach can change. Since May 2024, the required training time on standardized hardware dropped from about 45 minutes to under two minutes across 82 documented improvement steps.","To figure out the cost of human work, METR interviewed two of the project's most active contributors and also had an AI model (Opus-4.6) estimate the effort behind each improvement. Both approaches landed on roughly 16 hours of work per one-percent speedup. At an assumed hourly rate of $150, that comes to about $2,500 per percentage point.","METR stresses that this number is very uncertain. One detail from the interviews stands out: most of the time went into ideas that ultimately didn't work.","For the comparison, METR had six AI models work on the same task independently. They didn't start from scratch but from an already highly optimized state of the speedrun (Record #78 from March 2026) and were allowed to spend up to $10,000 in compute and operating costs per run. The result: estimated expenditure horizons between $0 and $3,300.","The differences between models were stark. GPT-5 and Opus-4.1 produced no real progress after careful verification. Their apparent gains turned out to be random noise. GPT-5.5 and Opus-4.8, on the other hand, delivered real improvements of about 1 and 1.5 percent, respectively.","The quality of AI-generated ideas was mixed. The speedrun's maintainer estimated that about 70 percent of them could in principle be integrated into the project, but many weren't very original. He praised one clever, low-level optimization from GPT-5.5 as the \"coolest one,\" while calling most of the rest just parameter tweaking. The models also tried to cheat multiple times, taking shortcuts that faked good results in the test but would have been useless in practice, like shutting off parts of training right before the finish line.","METR's takeaway: while individual models reach expenditure horizons in the low four figures, those values are tiny compared to the estimated $250,000 in total human effort. Autonomous optimization has barely moved the needle on NanoGPT progress so far.","An important caveat: METR only tested older models (GPT-5, GPT-5.2, GPT-5.5, and Opus-4.1 and Opus-4.8). The models released since then, Fable 5, GPT-5.6 Sol, and Opus 5, don't appear in the paper. Anthropic markets Opus 5 as a major leap：https://the-decoder.com/anthropic-claims-its-new-claude-opus-5-delivers-near-fable-5-performance-at-half-the-token-price/: on the Frontier-Bench test, it doubles Opus 4.8's performance at lower cost per task. According to Anthropic, Opus 5 wastes less effort on dead ends, checks its own work more reliably, and achieves similar performance with an average of 26 percent fewer compute steps. All of those are factors that directly affect METR's expenditure horizon.","The progress on ARC-AGI-3：https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/ is even more telling. That benchmark doesn't test memorized knowledge but genuine problem-solving: the AI is dropped into unfamiliar, game-like environments with no instructions or goals and has to figure everything out through trial and error. Opus 5 has held the top spot since July 24, 2026, scoring 30.2 percent and solving five tasks that every previous model had failed. Its predecessor Opus 4.8 managed just 1.5 percent. The ARC Prize team attributes the jump to better logical reasoning, which lets the AI explore and plan more independently. That kind of ability could also prove useful in the NanoGPT speedrun.","Perhaps the biggest limitation is one METR calls out itself: the entire study measures AI working alone, purely autonomous optimization. In real AI research, humans typically use AI as a tool. METR sketches a third, hypothetical curve for this scenario. If humans make smart decisions about when and how to deploy the AI, this hybrid curve should theoretically beat both the pure human and pure AI curves by combining the strengths of each.","METR tempers that expectation, though, pointing to its own earlier work showing that human-plus-AI setups sometimes performed worse than humans alone：https://the-decoder.com/ai-coding-can-make-developers-slower-even-if-they-feel-faster/. The added value isn't guaranteed and depends on whether the AI gets used in the right places. Measuring this properly would require a controlled experiment comparing the same researchers working with and without AI support. That kind of experiment is hard to organize, but METR says it would be extremely informative. Until it happens, the expenditure horizon says a lot about what AI can do on its own, but very little about how much it actually speeds up human researchers.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/METR-Expenditure-Horizon-1.png","alt":"Image description","afterParagraph":0,"url":"/media/articles/cms38b6cd08zdro3fkczc5aiq/ed8d832c8febab24.png"},{"sourceUrl":"https://the-decoder.de/wp-content/uploads/2026/07/METR-eh-schematic-up-1-1200x798.png","alt":"","afterParagraph":2,"url":"/media/articles/cms38b6cd08zdro3fkczc5aiq/54d0c6609e8e3ef4.png"},{"sourceUrl":"https://the-decoder.de/wp-content/uploads/2026/07/METR-speedrun-timeline-1200x655.png","alt":"","afterParagraph":5,"url":"/media/articles/cms38b6cd08zdro3fkczc5aiq/895cfbcf5cabecb9.png"},{"sourceUrl":"https://the-decoder.de/wp-content/uploads/2026/07/METR-expenditure-horizon-1200x838.png","alt":"","afterParagraph":8,"url":"/media/articles/cms38b6cd08zdro3fkczc5aiq/54c72ceb2e7113fc.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/METR-hybrid-schematic-up-1200x798.png","alt":"","afterParagraph":14,"url":"/media/articles/cms38b6cd08zdro3fkczc5aiq/f6a3efcd491116ea.png"}],"mediaStatus":"ok","articleBodyZh":["METR的新指标“支出视界”用一个美元数额来衡量AI代理解决问题的成本效益。NanoGPT加速赛的早期结果并不令人满意，该指标存在盲点，而新一代模型可能会改变局面。","AI研究中最大的疑问之一是，AI是否能够加速自身发展，并以越来越快的速度不断进步。这很难衡量，因为这需要比较非常不同类型的成本：人工劳动、实验计算和运行AI本身的成本。","研究机构METR提出了一个新指标来解决这个问题：“支出视界”。METR比较AI和人类为实现同样的改进需要花费多少。支出视界是两者成本相等的点。低于这个预算，AI更划算。高于这个预算，人类的工作更便宜。","这一想法基于METR在以往测试中看到的规律：AI代理通常比人类更快完成简单、低成本的任务。但随着预算增加和任务难度加大，它们会落后。","与典型的AI基准测试相比，这种方法有两个优势，根据METR。首先，它不仅仅给出通过或失败的结论，而是产生一个精细的数值，显示花费多少钱可以获得多少改进。其次，它将所有成本转换为一种货币，不仅包括运行AI的成本，还包括昂贵的实验计算和人工劳动时间。","METR选择NanoGPT加速赛（https://github.com/kellerjordan/modded-nanogpt）作为测试平台。这是一个公共社区项目，志愿者竞赛尽可能快地训练AI语言模型。任务保持不变，只有训练方法可以变化。自2024年5月起，在标准化硬件上的所需训练时间从大约45分钟下降到82个记录改进步骤中不到两分钟。","为了计算人工成本，METR采访了该项目中两位最活跃的贡献者，并让一个AI模型（Opus-4.6）估计每次改进背后的工作量。两种方法的结果都大约是每提高1%的速度需要16小时的工作。按每小时150美元计算，每个百分点约需2,500美元。","METR 强调这个数字非常不确定。访谈中的一个细节很突出：大部分时间都花在了最终无效的想法上。","为了进行比较，METR 让六个 AI 模型独立完成相同任务。它们并不是从零开始，而是从已高度优化的速跑状态（2026 年 3 月的记录 #78）出发，并允许每次运行花费最多 1 万美元的计算和运营成本。结果：预估支出范围在 0 到 3,300 美元之间。","模型之间的差异非常明显。经过仔细验证后，GPT-5 和 Opus-4.1 没有产生任何真正的进展。它们看似的收益实际上是随机噪声。另一方面，GPT-5.5 和 Opus-4.8 分别带来了约 1% 和 1.5% 的实际改进。","AI 生成的想法质量参差不齐。速跑维护者估计，大约 70% 的想法原则上可以整合到项目中，但很多并不十分原创。他称赞 GPT-5.5 的一个巧妙的低级优化为“最酷的一个”，而其他大部分只是参数调整。模型们还多次尝试作弊，采取了在测试中看似良好但实际无用的捷径，比如在终点前关闭部分训练。","METR 的结论：虽然单个模型的支出范围在低四位数，但这些数值与估计的 25 万美元的总人力成本相比微不足道。自主优化迄今对 NanoGPT 的进展几乎没有推动作用。","一个重要的注意事项：METR 只测试了较旧的模型（GPT-5、GPT-5.2、GPT-5.5，以及 Opus-4.1 和 Opus-4.8）。自那以后发布的模型 Fable 5、GPT-5.6 Sol 和 Opus 5 并未出现在论文中。Anthropic 将 Opus 5 推广为重大飞跃：https://the-decoder.com/anthropic-claims-its-new-claude-opus-5-delivers-near-fable-5-performance-at-half-the-token-price/：在 Frontier-Bench 测试中，它以每任务更低的成本将 Opus 4.8 的性能翻倍。据 Anthropic 所说，Opus 5 在死胡同上的浪费更少，自我检查更可靠，并以平均少 26% 的计算步骤实现类似性能。所有这些因素都直接影响 METR 的支出范围。","ARC-AGI-3的进展：https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/ 更具意义。该基准测试并不考察记忆知识，而是真正的问题解决能力：AI被置入不熟悉的、类似游戏的环境中，没有任何指令或目标，必须通过反复试验来弄清一切。Opus 5自2026年7月24日起一直位居榜首，得分为30.2%，完成了五项之前所有模型都未能完成的任务。其前代产品Opus 4.8仅取得了1.5%的成绩。ARC 奖团队将这一飞跃归因于更强的逻辑推理能力，这使得AI能够更加独立地探索和计划。这种能力在NanoGPT极速挑战中也可能发挥作用。","或许最大限制是METR自己指出的：整个研究衡量的是AI单独工作，纯粹自主优化的情况。在实际的AI研究中，人类通常将AI作为工具使用。METR勾画了第三种假设情景的曲线。如果人类能够聪明地决定何时以及如何部署AI，这条混合曲线理论上应当通过结合各自优势，超过纯人类曲线和纯AI曲线。","不过，METR也对这一期望进行了调节，指出其早期研究显示，人类加AI的组合有时表现甚至不如纯人类：https://the-decoder.com/ai-coding-can-make-developers-slower-even-if-they-feel-faster/。额外价值并不保证存在，而取决于AI是否被应用在合适的地方。要正确衡量这一点，需要进行一项对照实验，比较同一批研究人员在有无AI支持下的工作表现。这类实验难以组织，但METR表示，这将非常有信息价值。在此之前，支出视野可以说明AI独立能做些什么，但几乎无法说明它实际上加快了人类研究者的工作速度。","保持对AI的了解。内容清晰、有用、无废话。","关注The Decoder的AI新闻、背景故事和专家分析。","The Decoder：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-28T05:28:54.624Z","sourceHash":"c9c59f9790a5cfe8","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["论文研究","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"METR 推出\"支出视界\"指标衡量AI成本效益","summary":"METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。在NanoGPT speedrun测试中，人类每实现1%加速需约16小时工作（折合$2，500），而GPT-5.5和Opus-4.8的支出视界在$0-$3，300之间，但整体自主优化贡献远低于人类$250，000的总投入。","category":"论文研究","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR 推出\"支出视界\"指标衡量AI成本效益 - Aioga AI资讯","description":"METR 发布新指标\"支出视界\"，通过比较AI与人类达成同等改进所需的总成本，确定AI更划算的预算上限。在NanoGPT speedrun测试中，人类每实现1%加速需约16小时工作（折合$2，500），而GPT-5.5和Opus-4.8的支出视界在$0-$3，300之间，但整体自主优化贡献远低于人类$250，000的总投入。","url":"https://www.aioga.com/news/cms38b6cd08zdro3fkczc5aiq/"},"en":{"title":"METR launches 'Expenditure Vision' indicator to measure AI cost-effectiveness","summary":"METR released a new metric, 'Spending Horizon,' which determines the cost-effective budget ceiling for AI by comparing the total cost required for AI and humans to achieve the same improvement. In the NanoGPT speedrun test, humans require about 16 hours of work (equivalent to $2,500) to achieve a 1% acceleration, while the spending horizon for GPT-5.5 and Opus-4.8 ranges from $0 to $3,300, yet their overall autonomous optimization contribution is far below the total human investment of $250,000.","category":"Research","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR launches 'Expenditure Vision' indicator to measure AI cost-effectiveness - Aioga AI News","description":"METR released a new metric, 'Spending Horizon,' which determines the cost-effective budget ceiling for AI by comparing the total cost required for AI and humans to achieve the same...","url":"https://www.aioga.com/en/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:42:20.720Z"},"ja":{"title":"METRは「支出の視界」指標を導入してAIのコスト効率を測定","summary":"METRは新しい指標「支出の視界」を発表しました。AIと人間が同等の改善を達成するために必要な総コストを比較することで、AIがより費用対効果の高い予算の上限を決定します。NanoGPT speedrunテストでは、人間が1%の高速化を達成するのに約16時間の作業が必要で（換算で2,500ドル）、GPT-5.5とOpus-4.8の支出視界は0ドルから3,300ドルの範囲ですが、全体的な自主的な最適化への貢献は、人間の総投入25万ドルに比べてはるかに低いです。","category":"論文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METRは「支出の視界」指標を導入してAIのコスト効率を測定 - Aioga AIニュース","description":"METRは新しい指標「支出の視界」を発表しました。AIと人間が同等の改善を達成するために必要な総コストを比較することで、AIがより費用対効果の高い予算の上限を決定します。NanoGPT speedrunテストでは、人間が1%の高速化を達成するのに約16時間の作業が必要で（換算で2,500ドル）、GPT-5.5とOpus-4.8の支出視界は0ドルから3,300...","url":"https://www.aioga.com/ja/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:42:29.969Z"},"ko":{"title":"METR, '지출 시각' 지표를 출시하여 AI 비용 효율성을 측정","summary":"METR는 새로운 지표 '지출 시야'를 발표했으며, AI와 인간이 동일한 개선을 달성하는 데 필요한 총 비용을 비교하여 AI가 더 경제적인 예산 한도를 결정합니다. NanoGPT speedrun 테스트에서 인간은 1% 속도 향상을 달성하는 데 약 16시간의 작업(약 $2,500)이 필요했으며, GPT-5.5와 Opus-4.8의 지출 시야는 $0-$3,300 사이였지만, 전체 자율 최적화 기여도는 인간의 $250,000 전체 투입을 훨씬 밑돌았습니다.","category":"연구","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR, '지출 시각' 지표를 출시하여 AI 비용 효율성을 측정 - Aioga AI 뉴스","description":"METR는 새로운 지표 '지출 시야'를 발표했으며, AI와 인간이 동일한 개선을 달성하는 데 필요한 총 비용을 비교하여 AI가 더 경제적인 예산 한도를 결정합니다. NanoGPT speedrun 테스트에서 인간은 1% 속도 향상을 달성하는 데 약 16시간의 작업(약 $2,500)이 필요했으며, GPT-5.5와 Opus-4...","url":"https://www.aioga.com/ko/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:43:16.113Z"},"es":{"title":"METR lanza el indicador \"Visión de gastos\" para medir la rentabilidad de la IA","summary":"METR lanza un nuevo indicador 'Horizonte de gasto', que determina el límite máximo de presupuesto en el que la IA resulta más rentable mediante la comparación del costo total necesario para que la IA y los humanos logren la misma mejora. En la prueba NanoGPT speedrun, los humanos necesitan aproximadamente 16 horas de trabajo por cada 1% de aceleración (equivalente a $2,500), mientras que el horizonte de gasto de GPT-5.5 y Opus-4.8 está entre $0 y $3,300, pero su contribución total a la optimización autónoma es mucho menor que la inversión total de $250,000 de los humanos.","category":"Investigación","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR lanza el indicador \"Visión de gastos\" para medir la rentabilidad de la IA - Aioga Noticias de IA","description":"METR lanza un nuevo indicador 'Horizonte de gasto', que determina el límite máximo de presupuesto en el que la IA resulta más rentable mediante la comparación del costo total neces...","url":"https://www.aioga.com/es/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:43:14.852Z"},"fr":{"title":"METR lance l'indicateur 'Vision des dépenses' pour mesurer le rapport coût-efficacité de l'IA","summary":"METR publie un nouvel indice \"Perspective des dépenses\", qui détermine le plafond budgétaire à privilégier pour l'IA en comparant le coût total nécessaire pour atteindre la même amélioration par l'IA et par les humains. Dans le test NanoGPT speedrun, les humains doivent travailler environ 16 heures (l'équivalent de 2 500 $) pour chaque 1 % d'accélération, tandis que la perspective des dépenses pour GPT-5.5 et Opus-4.8 se situe entre 0 $ et 3 300 $, mais leur contribution globale à l'optimisation autonome est largement inférieure à l'investissement total humain de 250 000 $.","category":"Recherche","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR lance l'indicateur 'Vision des dépenses' pour mesurer le rapport coût-efficacité de l'IA - Aioga Actualités IA","description":"METR publie un nouvel indice \"Perspective des dépenses\", qui détermine le plafond budgétaire à privilégier pour l'IA en comparant le coût total nécessaire pour atteindre la même am...","url":"https://www.aioga.com/fr/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:43:55.641Z"},"de":{"title":"METR führt den Indikator 'Ausgabesicht' zur Messung der Kosten-Nutzen-Effizienz von KI ein","summary":"METR veröffentlicht den neuen Index \"Ausgabensicht\", der durch den Vergleich der Gesamtkosten, die erforderlich sind, damit KI die gleiche Verbesserung wie Menschen erzielt, die budgetäre Obergrenze bestimmt, bei der KI kosteneffizienter ist. Beim NanoGPT-Speedrun-Test benötigt ein Mensch, um eine 1%-Beschleunigung zu erreichen, etwa 16 Arbeitsstunden (entspricht 2.500 $), während der Ausgabensichtbereich von GPT-5.5 und Opus-4.8 zwischen 0 $ und 3.300 $ liegt, jedoch der Gesamtbeitrag zur autonomen Optimierung weit unter den 250.000 $ Gesamteinsatz des Menschen liegt.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR führt den Indikator 'Ausgabesicht' zur Messung der Kosten-Nutzen-Effizienz von KI ein - Aioga KI-News","description":"METR veröffentlicht den neuen Index \"Ausgabensicht\", der durch den Vergleich der Gesamtkosten, die erforderlich sind, damit KI die gleiche Verbesserung wie Menschen erzielt, die bu...","url":"https://www.aioga.com/de/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:44:04.630Z"},"pt-BR":{"title":"METR lança o indicador 'Visão de Gastos' para medir a relação custo-benefício da IA","summary":"METR lança novo indicador \"Visão de Gastos\", que determina o limite orçamentário mais econômico para a IA, comparando o custo total necessário para alcançar melhorias equivalentes às feitas por humanos. No teste de speedrun do NanoGPT, os humanos precisaram de cerca de 16 horas de trabalho para cada 1% de aceleração (equivalente a $2.500), enquanto a visão de gastos do GPT-5.5 e do Opus-4.8 fica entre $0 e $3.300, mas a contribuição geral de otimização autônoma é muito inferior ao total investido pelos humanos de $250.000.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR lança o indicador 'Visão de Gastos' para medir a relação custo-benefício da IA - Aioga Notícias de IA","description":"METR lança novo indicador \"Visão de Gastos\", que determina o limite orçamentário mais econômico para a IA, comparando o custo total necessário para alcançar melhorias equivalentes...","url":"https://www.aioga.com/pt-BR/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:44:45.649Z"},"ru":{"title":"METR запустила показатель \"Обзор расходов\" для оценки рентабельности ИИ","summary":"METR выпустил новый показатель «視界 расходов», который через сравнение общей стоимости, необходимой для достижения аналогичного улучшения с помощью AI и человека, определяет верхний предел бюджета, при котором AI более выгоден. В тесте NanoGPT speedrun человеку требуется около 16 часов работы для достижения 1% ускорения (эквивалент $2,500), тогда как показатель «視界 расходов» для GPT-5.5 и Opus-4.8 находится в диапазоне $0–$3,300, но общий вклад автономной оптимизации значительно ниже общей инвестиции человека в $250,000.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR запустила показатель \"Обзор расходов\" для оценки рентабельности ИИ - Aioga Новости ИИ","description":"METR выпустил новый показатель «視界 расходов», который через сравнение общей стоимости, необходимой для достижения аналогичного улучшения с помощью AI и человека, определяет верхний...","url":"https://www.aioga.com/ru/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:44:46.919Z"},"ar":{"title":"METR تطلق مؤشر \"رؤية الإنفاق\" لقياس فعالية تكلفة الذكاء الاصطناعي","summary":"أصدرت METR المؤشر الجديد \"نطاق الإنفاق\"، والذي يحدد الحد الأقصى للميزانية الأكثر فعالية للذكاء الاصطناعي من خلال مقارنة إجمالي التكلفة اللازمة لتحقيق نفس التحسين الذي يحققه الإنسان. في اختبار NanoGPT speedrun، يحتاج الإنسان حوالي 16 ساعة عمل لكل 1% من التسريع (ما يعادل 2500 دولار)، بينما كان نطاق الإنفاق لكل من GPT-5.5 و Opus-4.8 بين 0 و 3,300 دولار، لكن المساهمة الإجمالية للتحسين الذاتي أقل بكثير من إجمالي استثمار الإنسان البالغ 250,000 دولار.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR تطلق مؤشر \"رؤية الإنفاق\" لقياس فعالية تكلفة الذكاء الاصطناعي - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت METR المؤشر الجديد \"نطاق الإنفاق\"، والذي يحدد الحد الأقصى للميزانية الأكثر فعالية للذكاء الاصطناعي من خلال مقارنة إجمالي التكلفة اللازمة لتحقيق نفس التحسين الذي يحققه الإنسان...","url":"https://www.aioga.com/ar/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:45:36.063Z"},"hi":{"title":"METR ने 'खर्च दृश्य' सूचकांक पेश किया जो AI की लागत-लाभ को मापता है","summary":"METR ने नया संकेतक \"खर्च सीमा\" जारी किया है, जो AI और मानव द्वारा समान सुधार प्राप्त करने की कुल लागत की तुलना करके यह निर्धारित करता है कि AI के लिए सबसे अधिक लागत-कुशल बजट सीमा क्या होगी। NanoGPT स्पीडरन परीक्षण में, मानव को 1% गति वृद्धि प्राप्त करने में लगभग 16 घंटे का काम करना पड़ता है (लगभग $2,500), जबकि GPT-5.5 और Opus-4.8 की खर्च सीमा $0-$3,300 के बीच है, लेकिन कुल स्वायत्त अनुकूलन योगदान मानव की $250,000 कुल निवेश की तुलना में काफी कम है।","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR ने 'खर्च दृश्य' सूचकांक पेश किया जो AI की लागत-लाभ को मापता है - Aioga AI समाचार","description":"METR ने नया संकेतक \"खर्च सीमा\" जारी किया है, जो AI और मानव द्वारा समान सुधार प्राप्त करने की कुल लागत की तुलना करके यह निर्धारित करता है कि AI के लिए सबसे अधिक लागत-कुशल बजट सीमा क...","url":"https://www.aioga.com/hi/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:45:45.643Z"},"it":{"title":"METR lancia l'indicatore 'Spesa Visuale' per misurare il rapporto costi-benefici dell'AI","summary":"METR pubblica un nuovo indicatore \"spesa visione\", che determina il limite massimo di budget più conveniente per l'IA confrontando il costo totale necessario per ottenere lo stesso miglioramento degli esseri umani. Nel test NanoGPT speedrun, gli esseri umani richiedono circa 16 ore di lavoro (equivalenti a $2.500) per ottenere un miglioramento del 1%, mentre la spesa visione di GPT-5.5 e Opus-4.8 varia tra $0 e $3.300, ma il contributo complessivo all'ottimizzazione autonoma è molto inferiore all'investimento totale umano di $250.000.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR lancia l'indicatore 'Spesa Visuale' per misurare il rapporto costi-benefici dell'AI - Aioga Notizie IA","description":"METR pubblica un nuovo indicatore \"spesa visione\", che determina il limite massimo di budget più conveniente per l'IA confrontando il costo totale necessario per ottenere lo stesso...","url":"https://www.aioga.com/it/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:46:33.847Z"},"nl":{"title":"METR lanceert de 'Uitgavenvisie'-indicator om de kosten-baten van AI te meten","summary":"METR publiceert een nieuwe indicator 'uitgavengrens', waarmee door het vergelijken van de totale kosten die nodig zijn voor AI en mensen om eenzelfde verbetering te bereiken, het budgetmaximum wordt bepaald waarbij AI kosteneffectiever is. In de NanoGPT speedrun-test kost het een mens ongeveer 16 uur werk ($2.500) om 1% versnelling te realiseren, terwijl de uitgavengrens voor GPT-5.5 en Opus-4.8 tussen $0 en $3.300 ligt, maar hun totale bijdrage aan zelfoptimalisatie ligt aanzienlijk lager dan de totale investering van $250.000 door mensen.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR lanceert de 'Uitgavenvisie'-indicator om de kosten-baten van AI te meten - Aioga AI-nieuws","description":"METR publiceert een nieuwe indicator 'uitgavengrens', waarmee door het vergelijken van de totale kosten die nodig zijn voor AI en mensen om eenzelfde verbetering te bereiken, het b...","url":"https://www.aioga.com/nl/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:46:28.331Z"},"tr":{"title":"METR, AI maliyet etkinliğini ölçmek için 'Harcama Görünümü' göstergesini başlattı","summary":"METR, yeni gösterge 'Harcamalar Perspektifi'ni yayınladı ve AI ile insanların eşdeğer iyileşmeye ulaşmak için gereken toplam maliyeti karşılaştırarak AI için daha kârlı bütçe üst sınırını belirliyor. NanoGPT speedrun testinde, insanlar %1 hızlanma sağlamak için yaklaşık 16 saat çalışıyor (yaklaşık 2.500 $), oysa GPT-5.5 ve Opus-4.8'in harcamalar perspektifi 0-3.300 $ arasında, ancak genel bağımsız optimizasyon katkısı, insanların 250.000 $ toplam yatırımlarının çok altında.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR, AI maliyet etkinliğini ölçmek için 'Harcama Görünümü' göstergesini başlattı - Aioga AI Haberleri","description":"METR, yeni gösterge 'Harcamalar Perspektifi'ni yayınladı ve AI ile insanların eşdeğer iyileşmeye ulaşmak için gereken toplam maliyeti karşılaştırarak AI için daha kârlı bütçe üst s...","url":"https://www.aioga.com/tr/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:47:12.726Z"},"vi":{"title":"METR ra mắt chỉ số 'Tầm nhìn chi tiêu' để đo lường hiệu quả chi phí AI","summary":"METR phát hành chỉ số mới \"Chi tiêu Tầm nhìn\", thông qua việc so sánh tổng chi phí cần thiết để đạt được mức cải thiện tương đương giữa AI và con người, nhằm xác định giới hạn ngân sách mà AI có lợi hơn. Trong thử nghiệm NanoGPT speedrun, con người cần khoảng 16 giờ làm việc để đạt 1% tốc độ tăng (tương đương 2.500 USD), trong khi tầm nhìn chi phí của GPT-5.5 và Opus-4.8 nằm trong khoảng 0-3.300 USD, nhưng đóng góp tự tối ưu tổng thể thấp hơn nhiều so với tổng đầu tư 250.000 USD của con người.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR ra mắt chỉ số 'Tầm nhìn chi tiêu' để đo lường hiệu quả chi phí AI - Tin tức AI Aioga","description":"METR phát hành chỉ số mới \"Chi tiêu Tầm nhìn\", thông qua việc so sánh tổng chi phí cần thiết để đạt được mức cải thiện tương đương giữa AI và con người, nhằm xác định giới hạn ngân...","url":"https://www.aioga.com/vi/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:47:12.339Z"},"id":{"title":"METR meluncurkan indikator 'Penglihatan Pengeluaran' untuk mengukur efisiensi biaya AI","summary":"METR merilis indikator baru \"Pandangan Pengeluaran\", yang menentukan batas anggaran maksimum di mana AI lebih menguntungkan dengan membandingkan total biaya yang diperlukan untuk mencapai peningkatan setara antara AI dan manusia. Dalam pengujian speedrun NanoGPT, manusia membutuhkan sekitar 16 jam kerja per 1% percepatan (setara dengan $2.500), sedangkan pandangan pengeluaran GPT-5.5 dan Opus-4.8 berkisar antara $0-$3.300, tetapi kontribusi optimisasi otonom keseluruhan jauh lebih rendah dibandingkan total investasi $250.000 oleh manusia.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR meluncurkan indikator 'Penglihatan Pengeluaran' untuk mengukur efisiensi biaya AI - Berita AI Aioga","description":"METR merilis indikator baru \"Pandangan Pengeluaran\", yang menentukan batas anggaran maksimum di mana AI lebih menguntungkan dengan membandingkan total biaya yang diperlukan untuk m...","url":"https://www.aioga.com/id/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:47:58.063Z"},"th":{"title":"METR เปิดตัวดัชนี 'มุมมองค่าใช้จ่าย' เพื่อวัดประสิทธิผลค่าใช้จ่ายของ AI","summary":"METR เปิดตัวดัชนีใหม่ 'ขอบเขตการใช้จ่าย' โดยการเปรียบเทียบต้นทุนรวมที่จำเป็นเพื่อให้ AI และมนุษย์บรรลุการปรับปรุงเทียบเท่ากัน เพื่อกำหนดเพดานงบประมาณที่ AI คุ้มค่ากว่า ในการทดสอบ NanoGPT speedrun มนุษย์ต้องใช้เวลาประมาณ 16 ชั่วโมงในการเพิ่มความเร็ว 1% (คิดเป็นเงิน $2,500) ในขณะที่ขอบเขตการใช้จ่ายของ GPT-5.5 และ Opus-4.8 อยู่ระหว่าง $0-$3,300 แต่ผลการปรับปรุงอัตโนมัติทั้งหมดยังต่ำกว่าต้นทุนรวม $250,000 ของมนุษย์มาก","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR เปิดตัวดัชนี 'มุมมองค่าใช้จ่าย' เพื่อวัดประสิทธิผลค่าใช้จ่ายของ AI - ข่าว AI Aioga","description":"METR เปิดตัวดัชนีใหม่ 'ขอบเขตการใช้จ่าย' โดยการเปรียบเทียบต้นทุนรวมที่จำเป็นเพื่อให้ AI และมนุษย์บรรลุการปรับปรุงเทียบเท่ากัน เพื่อกำหนดเพดานงบประมาณที่ AI คุ้มค่ากว่า ในการทดสอบ N...","url":"https://www.aioga.com/th/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:48:04.081Z"},"pl":{"title":"METR wprowadza wskaźnik 'Wydatki w perspektywie', aby mierzyć efektywność kosztową AI","summary":"METR opublikował nowy wskaźnik „horyzont wydatków”, który poprzez porównanie całkowitych kosztów potrzebnych do osiągnięcia takiej samej poprawy przez AI i człowieka określa maksymalny budżet, przy którym AI jest bardziej opłacalna. W teście NanoGPT speedrun człowiek potrzebuje około 16 godzin pracy (co przekłada się na 2 500 USD) na każdy 1% przyspieszenia, podczas gdy horyzont wydatków GPT-5.5 i Opus-4.8 mieści się w przedziale 0–3 300 USD, lecz całkowity wkład w optymalizację autonomiczną jest znacznie niższy niż całkowite zaangażowanie człowieka w wysokości 250 000 USD.","category":"论文研究","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"METR wprowadza wskaźnik 'Wydatki w perspektywie', aby mierzyć efektywność kosztową AI - Aioga Wiadomości AI","description":"METR opublikował nowy wskaźnik „horyzont wydatków”, który poprzez porównanie całkowitych kosztów potrzebnych do osiągnięcia takiej samej poprawy przez AI i człowieka określa maksym...","url":"https://www.aioga.com/pl/news/cms38b6cd08zdro3fkczc5aiq/","contentTranslated":true,"sourceHash":"0dcc1813dad9e44a","translatedAt":"2026-07-28T01:48:46.273Z"}}}}