{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-11T09:21:12.743Z","headline":"MirrorCode：AI 能独立完成的最大规模软件项目基准测试","description":"Anthropic 与 METR 联合推出 MirrorCode 基准，要求 AI 在无源码、无联网条件下端到端重写完整程序。Claude Opus 4.7 用 14 小时、$251 重写了约 16，000 行 Go 代码的生物信息学工具 gotree，而人类工程师预计需 2-17 周。该基准单次运行最高花费 $2，600，AI 连续工作 19 天，目前 22/25 个目标程序已开源。","url":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","mainEntityOfPage":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","datePublished":"2026-08-04T22:31:50.984Z","dateModified":"2026-08-04T22:31:50.984Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://epoch.ai/MirrorCode","https://aihot.virxact.com/items/cmsf9lwvw1kbaro2ek5d6lqpx"],"canonicalUrl":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","directAnswer":{"@type":"Answer","text":"MirrorCode 由 Anthropic 与 METR 联合推出，用于测试 AI 在无源码、无联网条件下端到端重写完整程序的能力，并通过包含隐藏用例的端到端测试核对输出。25 个目标程序覆盖多类计算领域。","url":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","dateCreated":"2026-08-04T22:31:50.984Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"epoch.ai source article","url":"https://epoch.ai/MirrorCode","datePublished":"2026-08-04T22:31:50.984Z","provider":{"@type":"Organization","name":"epoch.ai","url":"https://epoch.ai/MirrorCode"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsf9lwvw1kbaro2ek5d6lqpx","datePublished":"2026-08-04T22:31:50.984Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsf9lwvw1kbaro2ek5d6lqpx"}}],"aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","originalPublisher":{"name":"epoch.ai","url":"https://epoch.ai/MirrorCode"},"geoDeepAnswer":null,"article":{"id":"cmsf9lwvw1kbaro2ek5d6lqpx","slug":"cmsf9lwvw1kbaro2ek5d6lqpx","url":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","title":"MirrorCode：AI 能独立完成的最大规模软件项目基准测试","title_en":"人工智能能够独立完成的最大规模的软件项目是什么？","summary":"Anthropic 与 METR 联合推出 MirrorCode 基准，要求 AI 在无源码、无联网条件下端到端重写完整程序。Claude Opus 4.7 用 14 小时、$251 重写了约 16，000 行 Go 代码的生物信息学工具 gotree，而人类工程师预计需 2-17 周。该基准单次运行最高花费 $2，600，AI 连续工作 19 天，目前 22/25 个目标程序已开源。","source":"Hacker News 热门（buzzing.cc 中文翻译）","sourceUrl":"https://epoch.ai/MirrorCode","aiHotUrl":"https://aihot.virxact.com/items/cmsf9lwvw1kbaro2ek5d6lqpx","publishedAt":"2026-08-04T22:31:50.984Z","category":"论文研究","score":67,"selected":false,"articleBody":["AI has made rapid progress on software engineering benchmarks in the past few years. However, most such benchmarks tend to focus on shorter tasks like fixing bugs or implementing individual features. MirrorCode is our benchmark, co-developed with METR, to test AI models on long-horizon coding tasks. In a MirrorCode task, AI models are tasked with reimplementing an entire program end-to-end, without access to the original source code. AI-generated solutions must match the original program’s output exactly on end-to-end tests, including held-out tests. MirrorCode’s 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.","Crucially, we provide a large enough inference budget to make a serious attempt at MirrorCode tasks. Many existing software engineering benchmarks limit inference spending to around $1–10, even when the task would take weeks for a human to complete. For example, one of the largest MirrorCode tasks cost $2,600 for a single run and involved AI working for 19 days without human intervention.","Reimplementing entire programs is extremely challenging for human software engineers. We believe a human engineer without AI would take months to solve the most complex MirrorCode tasks. However, MirrorCode tasks are also feasible; we know that there is enough information for the tasks to be fair.","We sandbox AI models, requiring them to conduct their work without access to the internet, without access to the original codebase, and with no way to cheat on the task. There are end-to-end tests that models never see while developing their code, so they cannot simply create a lookup table to mimic the original program's outputs.","AI can already solve long-horizon MirrorCode tasks, despite their difficulty. For example, Claude Opus 4.7 reimplemented gotree: a bioinformatics toolkit with ~16,000 lines of Go and 40+ commands. 1：#user-content-fn-1 We believe this same task would take a human engineer without AI assistance 2–17 weeks. Opus 4.7 solved it in 14 hours, costing $251.","One important caveat to these results is data contamination. Because MirrorCode tasks involve reimplementing open-source programs, AI models are likely to have seen the original codebases in pretraining. This might lead to inflated performance on the benchmark. However, AI successfully reimplemented several target programs that passed our memorization screen, and failed to reimplement programs where the screen showed evidence of memorization. This suggests that the results were not dominated by memorization, but we cannot rule out the possibility that memorization contributes to AI performance. Overall, we expect that the capabilities measured by MirrorCode would generalize to an unseen codebase. We discuss this further, along with more results and details on benchmark construction, in the paper：https://arxiv.org/pdf/2606.30182.","MirrorCode is not fully solved. For our regularly updated leaderboard, we report MirrorCode (ML, +Private, 2L) . This means we run the 15 target programs from the Medium and Large buckets, and drop the Small bucket. Each target program is evaluated in two implementation languages (generally Go and Ada) giving 30 tasks. We run each task three times, with a budget of 10 billion tokens and 7 days per attempt. 2：#user-content-fn-2","We release our scaffold and 22 of the 25 MirrorCode target programs (totaling 132 task instances across the six supported programming languages) as open-source：https://github.com/epoch-research/MirrorCode, with the other three targets held out as a private test set.","This work was co-developed with METR and supported by a grant from METR. The authors of MirrorCode are Tom Adamczewski, David Owen, and David Rein. Florian Brand, Giles Edkins, Allen Hart, and Daniel O’Connell contributed additional target programs. Rasmus Faber-Espensen made crucial infrastructure improvements and gave advice on engineering","The best-scoring AI gotree implementations passed 2000/2001 tests, but failed a single edge-case test for a niche command to manipulate date annotations. Consequently, they do not strictly solve the task to 100% completion, but we consider the reimplementation near-perfect, covering essentially all scoped functionality. ：#user-content-fnref-1","See definitions in the “Suggested naming conventions” of the paper. “+Private” indicates that the private test set is included: here, private_M and private_L, the private targets in the Medium and Large buckets. Under the 2L mapping, target programs generally use Go and Ada (one mainstream language and one low-resource language), with two exceptions. These scores are not directly comparable with the paper, which evaluated all 25 target programs, used all 6 agent implementation languages on the Small and Medium buckets, and gave attempts a budget of 1 billion tokens except on Large targets, with no time limit. ：#user-content-fnref-2","Epoch AI’s work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons Attribution license：https://creativecommons.org/licenses/by/4.0/.","Get the latest from Epoch AI in your inbox","Have a question? Noticed something wrong? Let us know.","If you would like a reply, please include your name and email address.","MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code."],"articleImages":[],"mediaStatus":"none","articleBodyZh":["近年来，AI在软件工程基准测试上取得了快速进展。然而，大多数此类基准测试往往专注于较短的任务，例如修复错误或实现单个功能。MirrorCode 是我们与 METR 共同开发的基准测试，用于测试 AI 模型在长周期编码任务中的表现。在 MirrorCode 任务中，AI 模型的任务是从头重实现整个程序，不得访问原始源代码。AI 生成的解决方案必须在端到端测试中（包括保留测试）完全匹配原程序的输出。MirrorCode 的 25 个目标程序涵盖不同的计算领域：Unix 工具、数据序列化与查询工具、生物信息学、解释器、静态分析、密码学及压缩。","关键是，我们提供了足够大的推理预算，以便对 MirrorCode 任务进行认真尝试。许多现有的软件工程基准测试将推理开支限制在约 1–10 美元，即使任务对人类来说需要几周时间才能完成。例如，MirrorCode 中最大的一项任务单次运行成本为 2,600 美元，AI 连续工作了 19 天，没有人工干预。","从头重实现整个程序对人类软件工程师来说极具挑战。我们相信，即使没有 AI 的帮助，人类工程师也需要数月时间才能完成最复杂的 MirrorCode 任务。然而，MirrorCode 任务也是可行的；我们知道这类任务的信息量足够，使其具有公平性。","我们对 AI 模型进行沙箱管理，要求它们在进行工作时无法访问互联网、无法访问原始代码库，且无法在任务中作弊。还有一些模型在开发代码时永远不会看到的端到端测试，因此它们不能简单地创建查找表来模拟原程序的输出。","尽管困难重重，AI 已经能够解决长周期的 MirrorCode 任务。例如，Claude Opus 4.7 成功重实现了 gotree：一个包含约 16,000 行 Go 代码和 40 多条命令的生物信息学工具包。我们认为，此任务没有 AI 辅助的人类工程师需要 2–17 周的时间，而 Opus 4.7 在 14 小时内完成，成本为 251 美元。","这些结果的一个重要警告是数据污染。因为MirrorCode任务涉及重新实现开源程序，AI模型在预训练中很可能已经见过原始代码库。这可能导致基准测试结果被高估。然而，AI成功重新实现了一些通过我们记忆筛选的目标程序，而在记忆筛选显示有记忆证据的程序上失败了。这表明结果并非完全由记忆主导，但我们不能排除记忆对AI性能有贡献的可能性。总体而言，我们预计MirrorCode测量的能力会推广到未见过的代码库。我们在论文中对此进行了进一步讨论，并提供了更多结果和基准构建细节：https://arxiv.org/pdf/2606.30182。","MirrorCode并未完全解决。在我们定期更新的排行榜中，我们报告MirrorCode (ML, +Private, 2L)。这意味着我们运行来自中大型类别的15个目标程序，并丢弃小型类别的程序。每个目标程序在两种实现语言中进行评估（通常是Go和Ada），共30个任务。我们对每个任务运行三次，每次尝试的预算为100亿个token和7天。","我们发布了我们的框架和25个MirrorCode目标程序中的22个（涵盖六种支持的编程语言，总计132个任务实例）作为开源：https://github.com/epoch-research/MirrorCode，其他三个目标保留作为私有测试集。","这项工作与METR共同开发，并由METR资助。MirrorCode的作者是Tom Adamczewski、David Owen 和 David Rein。Florian Brand、Giles Edkins、Allen Hart和Daniel O’Connell贡献了额外的目标程序。Rasmus Faber-Espensen在关键基础设施改进和工程建议方面提供了支持。","得分最高的AI gotree实现通过了2000/2001个测试，但在一个针对操作日期注释的冷门命令的边缘案例测试中失败。因此，它们并未严格达到100%的任务完成度，但我们认为重新实现几乎完美，涵盖了基本上所有的功能范围。","请参阅论文中“建议命名规范”的定义。“+Private”表示包含私有测试集：这里指 private_M 和 private_L，即中等和大型类别中的私有目标。在 2L 映射下，目标程序通常使用 Go 和 Ada（一种主流语言和一种低资源语言），有两个例外。这些分数与论文中的分数不可直接比较，论文评估了所有 25 个目标程序，在 Small 和 Medium 类别中使用了所有 6 种代理实现语言，并为尝试分配了 10 亿个 token 的预算（大型目标除外），且没有时间限制。：#user-content-fnref-2","Epoch AI 的作品可以自由使用、分发和复制，但必须在署名的情况下遵循知识共享署名许可协议：https://creativecommons.org/licenses/by/4.0/。","在收件箱中获取 Epoch AI 的最新信息","有问题吗？发现了错误？请告诉我们。","如果您希望收到回复，请提供您的姓名和电子邮件地址。","MirrorCode 是 Epoch AI 的长期编码基准测试：AI 可以从头到尾重新实现整个程序，而无需访问原始源代码。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"MirrorCode 由 Anthropic 与 METR 联合推出，用于测试 AI 在无源码、无联网条件下端到端重写完整程序的能力，并通过包含隐藏用例的端到端测试核对输出。25 个目标程序覆盖多类计算领域。","background":"现有软件工程基准多聚焦修复缺陷或实现单项功能，且推理预算通常约为 1 至 10 美元。MirrorCode 转向长周期任务，允许更高预算；其中一次大型运行花费 2600 美元，AI 在无人干预下工作 19 天。","viewpoint":"Aioga 判断，MirrorCode 的价值在于把评估对象从局部代码修改扩展到完整程序复现，并设置沙箱、隐藏测试与较长运行预算。不过开源代码可能进入预训练数据，研究方也明确表示无法排除记忆对成绩的贡献。","implications":"Claude Opus 4.7 在 14 小时内以 251 美元重写约 1.6 万行 Go、包含 40 多个命令的 gotree；研究方估计人类工程师不用 AI 需 2 至 17 周。值得关注的是，该结果仅代表特定任务与测试条件，基准尚未被完全解决。","nextStep":"后续可关注定期更新的排行榜、私有测试集表现，以及通过记忆筛查的目标程序能否持续复现当前结果。项目已开放脚手架和 25 个目标程序中的 22 个，其余 3 个保留为私有测试集。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-05T00:03:09.652Z","sourceHash":"8714ac1e67be2398","review":{"approved":true,"groundedness":96,"clarity":92,"duplicationRisk":18,"blockingIssues":[],"notes":["background 中“推理预算通常约为 1 至 10 美元”可更严谨地改为“许多现有软件工程基准将推理预算限制在约 1 至 10 美元”，以贴合来源原文的限定范围。","viewpoint 已将“Aioga 判断”明确标为观点，且关于预训练数据污染与记忆贡献的表述符合研究方披露的 caveat。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["论文研究","Hacker News 热门（buzzing.cc 中文翻译）"],"translations":{"zh-CN":{"title":"MirrorCode：AI 能独立完成的最大规模软件项目基准测试","summary":"Anthropic 与 METR 联合推出 MirrorCode 基准，要求 AI 在无源码、无联网条件下端到端重写完整程序。Claude Opus 4.7 用 14 小时、$251 重写了约 16，000 行 Go 代码的生物信息学工具 gotree，而人类工程师预计需 2-17 周。该基准单次运行最高花费 $2，600，AI 连续工作 19 天，目前 22/25 个目标程序已开源。","category":"论文研究","source":"epoch.ai","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode：AI 能独立完成的最大规模软件项目基准测试 - Aioga AI资讯","description":"Anthropic 与 METR 联合推出 MirrorCode 基准，要求 AI 在无源码、无联网条件下端到端重写完整程序。Claude Opus 4.7 用 14 小时、$251 重写了约 16，000 行 Go 代码的生物信息学工具 gotree，而人类工程师预计需 2-17 周。该基准单次运行最高花费 $2，600，AI 连续工作 19 天，目前 2...","url":"https://www.aioga.com/news/cmsf9lwvw1kbaro2ek5d6lqpx/","articleBody":["近年来，AI在软件工程基准测试上取得了快速进展。然而，大多数此类基准测试往往专注于较短的任务，例如修复错误或实现单个功能。MirrorCode 是我们与 METR 共同开发的基准测试，用于测试 AI 模型在长周期编码任务中的表现。在 MirrorCode 任务中，AI 模型的任务是从头重实现整个程序，不得访问原始源代码。AI 生成的解决方案必须在端到端测试中（包括保留测试）完全匹配原程序的输出。MirrorCode 的 25 个目标程序涵盖不同的计算领域：Unix 工具、数据序列化与查询工具、生物信息学、解释器、静态分析、密码学及压缩。","关键是，我们提供了足够大的推理预算，以便对 MirrorCode 任务进行认真尝试。许多现有的软件工程基准测试将推理开支限制在约 1–10 美元，即使任务对人类来说需要几周时间才能完成。例如，MirrorCode 中最大的一项任务单次运行成本为 2,600 美元，AI 连续工作了 19 天，没有人工干预。","从头重实现整个程序对人类软件工程师来说极具挑战。我们相信，即使没有 AI 的帮助，人类工程师也需要数月时间才能完成最复杂的 MirrorCode 任务。然而，MirrorCode 任务也是可行的；我们知道这类任务的信息量足够，使其具有公平性。","我们对 AI 模型进行沙箱管理，要求它们在进行工作时无法访问互联网、无法访问原始代码库，且无法在任务中作弊。还有一些模型在开发代码时永远不会看到的端到端测试，因此它们不能简单地创建查找表来模拟原程序的输出。","尽管困难重重，AI 已经能够解决长周期的 MirrorCode 任务。例如，Claude Opus 4.7 成功重实现了 gotree：一个包含约 16,000 行 Go 代码和 40 多条命令的生物信息学工具包。我们认为，此任务没有 AI 辅助的人类工程师需要 2–17 周的时间，而 Opus 4.7 在 14 小时内完成，成本为 251 美元。","这些结果的一个重要警告是数据污染。因为MirrorCode任务涉及重新实现开源程序，AI模型在预训练中很可能已经见过原始代码库。这可能导致基准测试结果被高估。然而，AI成功重新实现了一些通过我们记忆筛选的目标程序，而在记忆筛选显示有记忆证据的程序上失败了。这表明结果并非完全由记忆主导，但我们不能排除记忆对AI性能有贡献的可能性。总体而言，我们预计MirrorCode测量的能力会推广到未见过的代码库。我们在论文中对此进行了进一步讨论，并提供了更多结果和基准构建细节：https://arxiv.org/pdf/2606.30182。","MirrorCode并未完全解决。在我们定期更新的排行榜中，我们报告MirrorCode (ML, +Private, 2L)。这意味着我们运行来自中大型类别的15个目标程序，并丢弃小型类别的程序。每个目标程序在两种实现语言中进行评估（通常是Go和Ada），共30个任务。我们对每个任务运行三次，每次尝试的预算为100亿个token和7天。","我们发布了我们的框架和25个MirrorCode目标程序中的22个（涵盖六种支持的编程语言，总计132个任务实例）作为开源：https://github.com/epoch-research/MirrorCode，其他三个目标保留作为私有测试集。","这项工作与METR共同开发，并由METR资助。MirrorCode的作者是Tom Adamczewski、David Owen 和 David Rein。Florian Brand、Giles Edkins、Allen Hart和Daniel O’Connell贡献了额外的目标程序。Rasmus Faber-Espensen在关键基础设施改进和工程建议方面提供了支持。","得分最高的AI gotree实现通过了2000/2001个测试，但在一个针对操作日期注释的冷门命令的边缘案例测试中失败。因此，它们并未严格达到100%的任务完成度，但我们认为重新实现几乎完美，涵盖了基本上所有的功能范围。","请参阅论文中“建议命名规范”的定义。“+Private”表示包含私有测试集：这里指 private_M 和 private_L，即中等和大型类别中的私有目标。在 2L 映射下，目标程序通常使用 Go 和 Ada（一种主流语言和一种低资源语言），有两个例外。这些分数与论文中的分数不可直接比较，论文评估了所有 25 个目标程序，在 Small 和 Medium 类别中使用了所有 6 种代理实现语言，并为尝试分配了 10 亿个 token 的预算（大型目标除外），且没有时间限制。：#user-content-fnref-2","Epoch AI 的作品可以自由使用、分发和复制，但必须在署名的情况下遵循知识共享署名许可协议：https://creativecommons.org/licenses/by/4.0/。","在收件箱中获取 Epoch AI 的最新信息","有问题吗？发现了错误？请告诉我们。","如果您希望收到回复，请提供您的姓名和电子邮件地址。","MirrorCode 是 Epoch AI 的长期编码基准测试：AI 可以从头到尾重新实现整个程序，而无需访问原始源代码。"]},"en":{"title":"MirrorCode: The largest software project benchmarking AI can independently complete","summary":"Anthropic and METR jointly launched the MirrorCode benchmark, requiring AI to rewrite complete programs end-to-end without source code or internet connectivity. Claude Opus 4.7 rewrote the bioinformatics tool gotree, which had about 16,000 lines of Go code, in 14 hours and at $251, while human engineers expected 2-17 weeks. The benchmark runs up to $2,600 per run, the AI has worked continuously for 19 days, and currently 22 out of 25 target programs are open source.","category":"Research","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: The largest software project benchmarking AI can independently complete - Aioga AI News","description":"Anthropic and METR jointly launched the MirrorCode benchmark, requiring AI to rewrite complete programs end-to-end without source code or internet connectivity. Claude Opus 4.7 rew...","url":"https://www.aioga.com/en/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:33.327Z"},"ja":{"title":"MirrorCode:AIが独立して完了できる最大のソフトウェアベンチマーキングプロジェクト","summary":"AnthropicとMETRは共同でMirrorCodeベンチマークを開始し、AIがソースコードやインターネット接続なしで完全なプログラムをエンドツーエンドで書き換えることを要求しました。 Claude Opus 4.7は、約16,000行のGoコードを持つバイオインフォマティクスツールGotreeを14時間で251ドルで書き直しました。人間のエンジニアは2〜17週間の要約を予想していました。 ベンチマークは1回の実行で最大2,600ドルの費用がかかり、AIは19日間連続稼働しており、現在25のターゲットプログラムのうち22はオープンソースです。","category":"論文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode:AIが独立して完了できる最大のソフトウェアベンチマーキングプロジェクト - Aioga AIニュース","description":"AnthropicとMETRは共同でMirrorCodeベンチマークを開始し、AIがソースコードやインターネット接続なしで完全なプログラムをエンドツーエンドで書き換えることを要求しました。 Claude Opus 4.7は、約16,000行のGoコードを持つバイオインフォマティクスツールGotreeを14時間で251ドルで書き直しました。人間のエンジニアは2...","url":"https://www.aioga.com/ja/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:32.752Z"},"ko":{"title":"MirrorCode: AI가 독립적으로 수행할 수 있는 가장 큰 소프트웨어 벤치마킹 프로젝트","summary":"Anthropic과 METR은 공동으로 MirrorCode 벤치마크를 출시했으며, AI가 소스 코드나 인터넷 연결 없이 완전한 프로그램을 처음부터 끝까지 다시 작성하도록 요구했습니다. 클로드 Opus 4.7은 약 16,000줄의 Go 코드를 가진 생물정보학 도구를 14시간 만에 251달러에 재작성했으며, 인간 엔지니어들은 2주에서 17주를 예상했습니다. 벤치마크는 실행당 최대 $2,600이며, AI는 19일간 연속 작동했고, 현재 25개 대상 프로그램 중 22개가 오픈 소스입니다.","category":"연구","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: AI가 독립적으로 수행할 수 있는 가장 큰 소프트웨어 벤치마킹 프로젝트 - Aioga AI 뉴스","description":"Anthropic과 METR은 공동으로 MirrorCode 벤치마크를 출시했으며, AI가 소스 코드나 인터넷 연결 없이 완전한 프로그램을 처음부터 끝까지 다시 작성하도록 요구했습니다. 클로드 Opus 4.7은 약 16,000줄의 Go 코드를 가진 생물정보학 도구를 14시간 만에 251달러에 재작성했으며, 인간 엔지니어들은...","url":"https://www.aioga.com/ko/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:41.861Z"},"es":{"title":"MirrorCode: El mayor proyecto de benchmarking de software que la IA puede completar de forma independiente","summary":"Anthropic y METR lanzaron conjuntamente el benchmark MirrorCode, que requiere que la IA reescriba programas completos de principio a fin sin código fuente ni conexión a internet. Claude Opus 4.7 reescribió la herramienta de bioinformática gotree, que tenía unas 16.000 líneas de código Go, en 14 horas y a 251 dólares, mientras que los ingenieros humanos esperaban entre 2 y 17 semanas. El benchmark puede llegar a 2.600 dólares por partida, la IA ha trabajado de forma continua durante 19 días y actualmente 22 de los 25 programas objetivo son de código abierto.","category":"Investigación","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: El mayor proyecto de benchmarking de software que la IA puede completar de forma independiente - Aioga Noticias de IA","description":"Anthropic y METR lanzaron conjuntamente el benchmark MirrorCode, que requiere que la IA reescriba programas completos de principio a fin sin código fuente ni conexión a internet. C...","url":"https://www.aioga.com/es/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:41.996Z"},"fr":{"title":"MirrorCode : Le plus grand projet logiciel de benchmarking que l’IA peut réaliser de manière autonome","summary":"Anthropic et METR ont conjointement lancé le benchmark MirrorCode, exigeant que l’IA réécrive des programmes complets de bout en bout sans code source ni connexion internet. Claude Opus 4.7 a réécrit l’outil de bioinformatique gotree, qui comptait environ 16 000 lignes de code Go, en 14 heures et à 251 $, tandis que les ingénieurs humains s’attendaient à 2 à 17 semaines. Le benchmark peut atteindre 2 600 $ par exécution, l’IA fonctionne sans interruption depuis 19 jours, et actuellement 22 des 25 programmes ciblés sont open source.","category":"Recherche","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode : Le plus grand projet logiciel de benchmarking que l’IA peut réaliser de manière autonome - Aioga Actualités IA","description":"Anthropic et METR ont conjointement lancé le benchmark MirrorCode, exigeant que l’IA réécrive des programmes complets de bout en bout sans code source ni connexion internet. Claude...","url":"https://www.aioga.com/fr/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:50.715Z"},"de":{"title":"MirrorCode: Das größte Softwareprojekt, das KI unabhängig abschließen kann, als Benchmark","summary":"Anthropic und METR haben gemeinsam den MirrorCode-Benchmark eingeführt, der von KI verlangt, komplette Programme End-to-End-ohne Quellcode oder Internetverbindung neu zu schreiben. Claude Opus 4.7 schrieb das bioinformatische Tool gotree, das etwa 16.000 Zeilen Go-Code enthielt, in 14 Stunden und für 251 Dollar um, während menschliche Ingenieure mit 2–17 Wochen rechneten. Der Benchmark liegt bei bis zu 2.600 US-Dollar pro Durchlauf, die KI arbeitet seit 19 Tagen ununterbrochen, und derzeit sind 22 von 25 Zielprogrammen Open Source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Das größte Softwareprojekt, das KI unabhängig abschließen kann, als Benchmark - Aioga KI-News","description":"Anthropic und METR haben gemeinsam den MirrorCode-Benchmark eingeführt, der von KI verlangt, komplette Programme End-to-End-ohne Quellcode oder Internetverbindung neu zu schreiben....","url":"https://www.aioga.com/de/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:50.700Z"},"pt-BR":{"title":"MirrorCode: O maior projeto de benchmarking de software que a IA pode concluir de forma independente","summary":"A Anthropic e a METR lançaram conjuntamente o benchmark MirrorCode, exigindo que a IA reescreva programas completos de ponta a ponta, sem código-fonte ou conectividade à internet. Claude Opus 4.7 reescreveu a ferramenta de bioinformática gotree, que tinha cerca de 16.000 linhas de código Go, em 14 horas e a US$ 251, enquanto engenheiros humanos esperavam de 2 a 17 semanas. O benchmark custa até $2.600 por execução, a IA tem funcionado continuamente por 19 dias e, atualmente, 22 dos 25 programas-alvo são open source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: O maior projeto de benchmarking de software que a IA pode concluir de forma independente - Aioga Notícias de IA","description":"A Anthropic e a METR lançaram conjuntamente o benchmark MirrorCode, exigindo que a IA reescreva programas completos de ponta a ponta, sem código-fonte ou conectividade à internet....","url":"https://www.aioga.com/pt-BR/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:59.574Z"},"ru":{"title":"MirrorCode: крупнейший программный проект для бенчмаркинга ИИ может самостоятельно реализовать","summary":"Anthropic и METR совместно запустили бенчмарк MirrorCode, требующий от ИИ переписывать полные программы от самого начала до конца без исходного кода или подключения к интернету. Клод Опус 4.7 переписал биоинформатический инструмент Gotree, который содержал около 16 000 строк кода Go, за 14 часов и стоил $251, тогда как инженеры-люди рассчитывали на 2–17 недель. Бенчмарк составляет до $2,600 за запуск, ИИ работает непрерывно уже 19 дней, и в настоящее время 22 из 25 целевых программ являются открытыми по исходному коду.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: крупнейший программный проект для бенчмаркинга ИИ может самостоятельно реализовать - Aioga Новости ИИ","description":"Anthropic и METR совместно запустили бенчмарк MirrorCode, требующий от ИИ переписывать полные программы от самого начала до конца без исходного кода или подключения к интернету. Кл...","url":"https://www.aioga.com/ru/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:21:59.116Z"},"ar":{"title":"MirrorCode: أكبر مشروع برمجي يمكن للذكاء الاصطناعي إكماله بشكل مستقل","summary":"أطلقت Anthropic وMETR معا معيار MirrorCode، الذي يتطلب من الذكاء الاصطناعي إعادة كتابة البرامج الكاملة من البداية إلى النهاية دون الحاجة إلى الشفرة المصدرية أو الاتصال بالإنترنت. أعاد كلود أوبوس 4.7 كتابة أداة المعلوماتية الحيوية gotree، التي تحتوي على حوالي 16,000 سطر من كود Go، في 14 ساعة وبسعر 251 دولارا، بينما توقع المهندسون البشريون من 2 إلى 17 أسبوعا. يصل سعر الاختبار إلى 2,600 دولار لكل تشغيل، ويعمل الذكاء الاصطناعي بشكل متواصل لمدة 19 يوما، وحاليا 22 من أصل 25 برنامجا مستهدفا مفتوحة المصدر.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: أكبر مشروع برمجي يمكن للذكاء الاصطناعي إكماله بشكل مستقل - Aioga أخبار الذكاء الاصطناعي","description":"أطلقت Anthropic وMETR معا معيار MirrorCode، الذي يتطلب من الذكاء الاصطناعي إعادة كتابة البرامج الكاملة من البداية إلى النهاية دون الحاجة إلى الشفرة المصدرية أو الاتصال بالإنترنت. أ...","url":"https://www.aioga.com/ar/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:08.195Z"},"hi":{"title":"मिररकोड: सबसे बड़ा सॉफ्टवेयर प्रोजेक्ट बेंचमार्किंग एआई स्वतंत्र रूप से पूरा कर सकता है","summary":"एंथ्रोपिक और एमईटीआर ने संयुक्त रूप से मिररकोड बेंचमार्क लॉन्च किया, जिसमें एआई को स्रोत कोड या इंटरनेट कनेक्टिविटी के बिना पूरे कार्यक्रमों को फिर से लिखने की आवश्यकता होती है। क्लाउड ओपस 4.7 ने जैव सूचना विज्ञान उपकरण गोट्री को फिर से लिखा, जिसमें 16,000 घंटे में और $ 251 पर गो कोड की लगभग 251 लाइनें थीं, जबकि मानव इंजीनियरों को 2-17 सप्ताह की उम्मीद थी। बेंचमार्क $2,600 प्रति रन तक चलता है, एआई ने 19 दिनों तक लगातार काम किया है, और वर्तमान में 22 में से 25 लक्ष्य कार्यक्रम ओपन सोर्स हैं।","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"मिररकोड: सबसे बड़ा सॉफ्टवेयर प्रोजेक्ट बेंचमार्किंग एआई स्वतंत्र रूप से पूरा कर सकता है - Aioga AI समाचार","description":"एंथ्रोपिक और एमईटीआर ने संयुक्त रूप से मिररकोड बेंचमार्क लॉन्च किया, जिसमें एआई को स्रोत कोड या इंटरनेट कनेक्टिविटी के बिना पूरे कार्यक्रमों को फिर से लिखने की आवश्यकता होती है। क्...","url":"https://www.aioga.com/hi/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:07.876Z"},"it":{"title":"MirrorCode: Il più grande progetto software di benchmarking che l'IA possa completare in modo indipendente","summary":"Anthropic e METR hanno lanciato congiuntamente il benchmark MirrorCode, che richiede all'IA di riscrivere programmi completi end-to-end senza codice sorgente o connettività internet. Claude Opus 4.7 ha riscritto lo strumento bioinformatico gotree, che aveva circa 16.000 righe di codice Go, in 14 ore e a 251 dollari, mentre gli ingegneri umani si aspettavano da 2 a 17 settimane. Il benchmark arriva fino a 2.600 dollari per run, l'IA ha funzionato ininterrottamente per 19 giorni e attualmente 22 dei 25 programmi target sono open source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Il più grande progetto software di benchmarking che l'IA possa completare in modo indipendente - Aioga Notizie IA","description":"Anthropic e METR hanno lanciato congiuntamente il benchmark MirrorCode, che richiede all'IA di riscrivere programmi completi end-to-end senza codice sorgente o connettività interne...","url":"https://www.aioga.com/it/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:16.893Z"},"nl":{"title":"MirrorCode: Het grootste softwareproject dat AI zelfstandig kan benchmarken","summary":"Anthropic en METR lanceerden gezamenlijk de MirrorCode-benchmark, waarbij AI volledige programma's end-to-end moest herschrijven zonder broncode of internetverbinding. Claude Opus 4.7 herschreef de bio-informaticatool gotree, die ongeveer 16.000 regels Go-code had, in 14 uur en voor $251, terwijl menselijke ingenieurs 2-17 weken verwachtten. De benchmark loopt op tot $2.600 per run, de AI werkt al 19 dagen onafgebroken en momenteel zijn 22 van de 25 doelprogramma's open source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Het grootste softwareproject dat AI zelfstandig kan benchmarken - Aioga AI-nieuws","description":"Anthropic en METR lanceerden gezamenlijk de MirrorCode-benchmark, waarbij AI volledige programma's end-to-end moest herschrijven zonder broncode of internetverbinding. Claude Opus...","url":"https://www.aioga.com/nl/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:16.855Z"},"tr":{"title":"MirrorCode: Bağımsız olarak tamamlayabileceği en büyük yazılım projesi kıyaslama yapay zekası","summary":"Anthropic ve METR, yapay zekanın kaynak kodu veya internet bağlantısı olmadan tüm programları uçtan uca yeniden yazmasını gerektiren MirrorCode kıyaslamasını birlikte başlattı. Claude Opus 4.7, yaklaşık 16.000 satır Go kodu içeren biyoinformatik araç gotree'yi 14 saatte ve 251 dolara yeniden yazdı, insan mühendisler ise 2-17 hafta beklerken. Benchmark her bir koşu için $2.600'a kadar çıkıyor, yapay zeka 19 gün boyunca aralıksız çalışıyor ve şu anda 25 hedef programdan 22'si açık kaynaklı.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Bağımsız olarak tamamlayabileceği en büyük yazılım projesi kıyaslama yapay zekası - Aioga AI Haberleri","description":"Anthropic ve METR, yapay zekanın kaynak kodu veya internet bağlantısı olmadan tüm programları uçtan uca yeniden yazmasını gerektiren MirrorCode kıyaslamasını birlikte başlattı. Cla...","url":"https://www.aioga.com/tr/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:25.540Z"},"vi":{"title":"MirrorCode: Dự án phần mềm lớn nhất về benchmarking AI có thể tự thực hiện độc lập","summary":"Anthropic và METR đã cùng nhau khởi động bài kiểm tra MirrorCode, yêu cầu AI phải viết lại toàn bộ chương trình từ đầu đến cuối mà không cần mã nguồn hay kết nối internet. Claude Opus 4.7 đã viết lại công cụ tin sinh học Gotree, với khoảng 16.000 dòng mã Go, chỉ trong 14 giờ với giá 251 đô la, trong khi các kỹ sư con người dự kiến sẽ mất từ 2 đến 17 tuần. Benchmark lên đến $2,600 mỗi lần chạy, AI hoạt động liên tục trong 19 ngày, và hiện tại 22 trong số 25 chương trình mục tiêu là mã nguồn mở.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Dự án phần mềm lớn nhất về benchmarking AI có thể tự thực hiện độc lập - Tin tức AI Aioga","description":"Anthropic và METR đã cùng nhau khởi động bài kiểm tra MirrorCode, yêu cầu AI phải viết lại toàn bộ chương trình từ đầu đến cuối mà không cần mã nguồn hay kết nối internet. Claude O...","url":"https://www.aioga.com/vi/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:25.506Z"},"id":{"title":"MirrorCode: Proyek perangkat lunak benchmarking terbesar yang dapat dilakukan AI secara mandiri","summary":"Anthropic dan METR bersama-sama meluncurkan benchmark MirrorCode, yang mengharuskan AI menulis ulang program lengkap secara end-to-end tanpa kode sumber atau konektivitas internet. Claude Opus 4.7 menulis ulang alat bioinformatika gotree, yang memiliki sekitar 16.000 baris kode Go, dalam 14 jam dan seharga $251, sementara insinyur manusia memperkirakan 2-17 minggu. Tolok ukur ini mencapai $2.600 per run, AI telah bekerja terus-menerus selama 19 hari, dan saat ini 22 dari 25 program target adalah open source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Proyek perangkat lunak benchmarking terbesar yang dapat dilakukan AI secara mandiri - Berita AI Aioga","description":"Anthropic dan METR bersama-sama meluncurkan benchmark MirrorCode, yang mengharuskan AI menulis ulang program lengkap secara end-to-end tanpa kode sumber atau konektivitas internet....","url":"https://www.aioga.com/id/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:34.224Z"},"th":{"title":"MirrorCode: โครงการซอฟต์แวร์ที่ใหญ่ที่สุดที่สามารถทําการทดสอบประสิทธิภาพ AI ได้ด้วยตนเอง","summary":"Anthropic และ METR ร่วมกันเปิดตัวการทดสอบ MirrorCode ซึ่งกําหนดให้ AI ต้องเขียนโปรแกรมใหม่ทั้งหมดโดยไม่ต้องใช้ซอร์สโค้ดหรือการเชื่อมต่ออินเทอร์เน็ต Claude Opus 4.7 เขียนเครื่องมือชีวสารสนเทศ Gotree ซึ่งมีโค้ด Go ประมาณ 16,000 บรรทัด ในเวลา 14 ชั่วโมงและราคา $251 ขณะที่วิศวกรมนุษย์คาดว่าจะใช้เวลา 2-17 สัปดาห์ เบนช์มาร์กนี้คิดค่าบริการสูงสุดถึง $2,600 ต่อรอบ AI ทํางานต่อเนื่องมา 19 วัน และปัจจุบัน 22 จาก 25 โปรแกรมเป้าหมายเป็นโอเพ่นซอร์ส","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: โครงการซอฟต์แวร์ที่ใหญ่ที่สุดที่สามารถทําการทดสอบประสิทธิภาพ AI ได้ด้วยตนเอง - ข่าว AI Aioga","description":"Anthropic และ METR ร่วมกันเปิดตัวการทดสอบ MirrorCode ซึ่งกําหนดให้ AI ต้องเขียนโปรแกรมใหม่ทั้งหมดโดยไม่ต้องใช้ซอร์สโค้ดหรือการเชื่อมต่ออินเทอร์เน็ต Claude Opus 4.7 เขียนเครื่องมือช...","url":"https://www.aioga.com/th/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:34.363Z"},"pl":{"title":"MirrorCode: Największy projekt programistyczny benchmarkingowy, jaki AI może samodzielnie zrealizować","summary":"Anthropic i METR wspólnie wprowadziły benchmark MirrorCode, wymagający od AI przepisywania kompletnych programów end-to-end bez kodu źródłowego czy połączenia z internetem. Claude Opus 4.7 przepisał narzędzie bioinformatyczne gotree, które miało około 16 000 linii kodu Go, w 14 godzin i kosztował 251 dolarów, podczas gdy inżynierowie ludzki spodziewali się 2-17 tygodni. Benchmark kosztuje do $2,600 za przejście, AI działa nieprzerwanie od 19 dni, a obecnie 22 z 25 docelowych programów jest open source.","category":"论文研究","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","pageTitle":"MirrorCode: Największy projekt programistyczny benchmarkingowy, jaki AI może samodzielnie zrealizować - Aioga Wiadomości AI","description":"Anthropic i METR wspólnie wprowadziły benchmark MirrorCode, wymagający od AI przepisywania kompletnych programów end-to-end bez kodu źródłowego czy połączenia z internetem. Claude...","url":"https://www.aioga.com/pl/news/cmsf9lwvw1kbaro2ek5d6lqpx/","contentTranslated":true,"sourceHash":"4eb515f52007a9a5","translatedAt":"2026-08-04T23:22:43.032Z"}}}}