{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-11T09:21:12.743Z","headline":"AI智能体尚无法开展开放式AI研究","description":"普林斯顿大学团队通过\"影子评估\"测试前沿AI智能体的开放式研究能力：让智能体在六天时间内、使用数千美元API额度和算力，回答两篇未发表论文的核心研究问题，结果两篇论文均被原作者明确拒绝。分析显示，智能体缺乏研究判断力、资源意识、创造性反馈应对能力和有效回溯能力，且未遵循具体指令。研究团队认为，开放式研究对前沿AI智能体仍具挑战性，但结果尚属初步，需扩大样本量并持续评估。","url":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","mainEntityOfPage":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","datePublished":"2026-08-05T13:49:19.000Z","dateModified":"2026-08-05T13:49:19.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended","https://aihot.virxact.com/items/cmsg6qeks07otrolg6som6vy0"],"canonicalUrl":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","directAnswer":{"@type":"Answer","text":"普林斯顿大学团队让前沿AI智能体用六天时间及数千美元API额度和算力，回答两篇未发表论文的核心问题。原作者明确拒绝了两份智能体论文，团队认为开放式研究仍具挑战。","url":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","dateCreated":"2026-08-05T13:49:19.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"normaltech.ai source article","url":"https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended","datePublished":"2026-08-05T13:49:19.000Z","provider":{"@type":"Organization","name":"normaltech.ai","url":"https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsg6qeks07otrolg6som6vy0","datePublished":"2026-08-05T13:49:19.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsg6qeks07otrolg6som6vy0"}}],"aggregationSource":"AI as Normal Technology（RSS）","originalPublisher":{"name":"normaltech.ai","url":"https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended"},"geoDeepAnswer":null,"article":{"id":"cmsg6qeks07otrolg6som6vy0","slug":"cmsg6qeks07otrolg6som6vy0","url":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","title":"AI智能体尚无法开展开放式AI研究","title_en":"AI agents can't yet do open-ended AI research","summary":"普林斯顿大学团队通过\"影子评估\"测试前沿AI智能体的开放式研究能力：让智能体在六天时间内、使用数千美元API额度和算力，回答两篇未发表论文的核心研究问题，结果两篇论文均被原作者明确拒绝。分析显示，智能体缺乏研究判断力、资源意识、创造性反馈应对能力和有效回溯能力，且未遵循具体指令。研究团队认为，开放式研究对前沿AI智能体仍具挑战性，但结果尚属初步，需扩大样本量并持续评估。","source":"AI as Normal Technology（RSS）","sourceUrl":"https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended","aiHotUrl":"https://aihot.virxact.com/items/cmsg6qeks07otrolg6som6vy0","publishedAt":"2026-08-05T13:49:19.000Z","category":"技巧观点","score":70,"selected":true,"articleBody":["The goal of leading ：https://www.anthropic.com/institute/recursive-self-improvement AI labs ：https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts ：https://ai-2027.com/ of explosive ：https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/ AI ：https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion progress. How can we assess if we are close to this milestone?","One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations ：https://cruxevals.com/crux/can-ai-agents-conduct-research#introduction:~:text=AI%20research%20questions.-,Selected,-evaluations%20and%20demonstrations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of RSI.","But while these evaluations are helpful, they are limited to narrow, verifiable tasks. AI research can be much more open-ended. Success is often not immediately clear or verifiable, and to make progress, researchers need to test promising hypotheses, backtrack, or consider new or unconventional approaches. How can we evaluate agents’ ability to conduct open-ended AI research? We take our first step towards answering this question in a new paper ：https://cruxevals.com/crux/can-ai-agents-conduct-research .","We partnered with the authors of two unpublished AI papers and asked them to draft their papers’ main research questions. We then tasked frontier AI agents with conducting research to answer these questions, and gave them thousands of dollars of API credits and compute, and six days of wall-clock time. The original authors reviewed the agents’ papers.","The authors unambiguously rejected both agent papers. To better understand these results, our team spent over a hundred hours analyzing the agents’ logs. Our main takeaways:","The agents lacked the judgment for conducting open-ended research. While the agents proposed directions the expert reviewers found impressive, they quickly rejected their proposed directions based on low-quality or synthetic data.","The agents lacked awareness about the resources available to them. Both runs ended with less than 50% of the API budget spent and with hours left before the deadline, even though the agents could monitor their usage and were encouraged to spend down their budgets.","The agents did not creatively respond to feedback. Despite the agents’ own AI self-reviews surfacing many of the issues that the expert reviewers later raised, the agents did not creatively address these concerns. When faced with negative feedback they responded by adding caveats to existing findings, and doubled down on unpromising research directions.","The agents did not effectively backtrack. They retired their most ambitious research targets within the first day of the experiment, and neither agent fundamentally shifted its approach after that point.","The agents did not follow concrete instructions. They ignored explicit rules about how much time to spend on exploration, how often to get reviews from AI self-review tools, and strict limits on paper length.","We have wanted to evaluate AI’s ability to conduct open-ended research for two years, ever since we released a benchmark ：https://arxiv.org/html/2409.11363v1 to study if agents could be used to improve reproducibility. But we wanted to get our method right. The idea behind our method was suggested by some of the UK AISI coauthors of the paper and refined by our core team at Princeton.","We call these “shadow evaluations” since the agent shadows the original study. In addition to the two of us, the core team comprises Peter Kirgis, Andrew Schwartz, and Stephan Rabanser. The full author list is at the end of this essay.","Shadow evaluations have important advantages: they allow us to test agents on results they haven’t been trained on and can’t access online. They also allow experts who have spent months answering the questions to evaluate agents’ outputs. 1：#footnote-1","But shadow evaluations also have inherent limitations. Expert reviewers know that the paper is AI-generated, and they might prefer the approach they took over the one that the agent took. Because we are conducting in-depth evaluations of each paper, the sample size is small (in our study, we used just two papers). And these evaluations necessarily involve a lot of researcher flexibility in design, execution, and interpretation.","In fact, we are known for a particular position ：https://www.normaltech.ai/p/ai-as-normal-technology in the debate on recursive self-improvement and superintelligence. This could influence how we conduct the research. We have a detailed section in the paper on our potential biases and how we address them. We sought out a team of collaborators who don’t all share our priors, and we explicitly surface the disagreements that resulted. 2：#footnote-2 For future evaluations, we are interested in having “adversarial collaborators” as part of the core team.","Implications for explosive AI progress","Our results suggest that conducting open-ended research remains challenging for frontier AI agents. Still, these findings are tentative, and we are working to address the limitations, such as by increasing the sample size, testing with new models, and through potential scaffold improvements. But if these findings hold up, what are the implications?","First, we need to understand the extent to which frontier AI progress (and RSI) can be achieved simply by hill climbing on verifiable tasks. Our view is that while faster progress is certainly possible on narrow tasks (such as improving efficiency), we don’t think it will lead to broad RSI or explosive progress. Still, we plan to closely follow how AI progress unfolds as a result of AI agents’ capabilities at verifiable tasks.","Second, we need to measure how quickly current limitations of agents at conducting open-ended research (such as the lack of creativity and judgment) can be overcome, such as through more targeted training and scaffold improvements. We plan to continue shadow evaluations on a regular basis to help answer this question.","Finally, even if these limitations can be overcome, there may be further bottlenecks that dampen the pace of AI progress.","Bottlenecks could include compute limits, the necessity of collecting data from real-world experiments, and others that we haven’t recognized yet because they are not currently blocking progress. For example, the importance of high-quality RL environments was not clear before they turned out to be useful for inference scaling. Similarly, the importance of building energy infrastructure for data centers was not realized before companies started investing hundreds of billions on data centers for training and inference.","Whether we encounter further bottlenecks, and how tractable they turn out to be, will be consequential for understanding the pace of progress. In this vein, our paper identifies an unresolved bottleneck, namely, the poor performance of frontier agents on open-ended AI research (though it remains to be seen if it is on the critical path to RSI).","If we’re in a world where the bottlenecks to fully automated research can be easily resolved, we should expect dramatic returns to AI progress from improving AI capabilities. But if we’re in the world with many remaining bottlenecks that are hard to overcome, Amdahl’s law ：https://en.wikipedia.org/wiki/Amdahl%27s_law would kick in: even a hundredfold speedup in the parts amenable to AI would only lead to a small speedup in the overall pace of progress, since progress is bottlenecked by the pace of the slowest component. 3：#footnote-3","Figuring out which world we live in could dramatically impact estimates of the pace of AI progress. We hope our results contribute to a richer understanding of these bottlenecks.","Read the paper here ：https://arxiv.org/pdf/2607.27191 . The authors are Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, and Arvind Narayanan.","This is as opposed to sending papers for blind peer review. Unfortunately, peer review in AI suffers from poor review quality and a mismatch between reviewers’ expertise and the papers they are assigned, perhaps resulting from a dramatic ：https://www.cs.cmu.edu/~nihars/preprints/SurveyPeerReview.pdf increase in submissions.","For example, our coauthors disagree on whether the agents lack creativity, or if they suffered from epistemic lock-in and were unable to productively incorporate feedback.","In practice, it might turn out that some kinds of AI progress have bottlenecks while others don’t. For example, improving the efficiency and speed of existing AI systems is a verifiable task, where progress has been rapid. As a result, companies have productively ：https://x.com/OpenAI/status/2082878156483219672?s=20 utilized agents for improving AI systems’ efficiency. In the near term, we expect quick AI progress in dimensions that have verifiable signals.","Interesting that the agents get \"stuck in a loop\" and fail to reconsider previously set-aside pathways. \"That didn't work before, but maybe this time it will. Maybe if I just tweak something...\" seems to be a fundamentally human course of action (for all that it so often fails). And giving up before exhausting the budget? That ain't right! :-D","Just more evidence that whatever process of \"thinking\" the machines are using, it isn't *human*.","Does research proceed by asking research questions? My (Deweyan) view is that formulating your research as a question distorts it, unless you have already done the research."],"articleImages":[{"sourceUrl":"https://substackcdn.com/image/fetch/$s_!u4S7!,e_trim:10:white/e_trim:10:transparent/h_72,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595a103f-0f5f-421a-a9ed-009cf269b7b4_3024x652.jpeg","alt":"AI as Normal Technology","afterParagraph":0,"url":"/media/articles/cmsg6qeks07otrolg6som6vy0/166d2f38ebb5e94e.jpg"}],"mediaStatus":"ok","articleBodyZh":["领导的目标：https://www.anthropic.com/institute/recursive-self-improvement AI 实验室：https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf 是递归自我改进（RSI）：使用 AI 代理自动化 AI 研究。RSI 也支撑着对 AI 爆炸性进展的预测：https://ai-2027.com/：https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/：https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion。我们如何评估是否接近这一里程碑？","一种方法是使用基准测试，测试代理是否可以进行 AI 研究。鉴于 AI 社区对基准测试的关注，它们一直是评估 RSI 进展的主要方式。在过去一年中，许多此类评估：https://cruxevals.com/crux/can-ai-agents-conduct-research#introduction:~:text=AI%20research%20questions.-,Selected,-evaluations%20and%20demonstrations 发现代理现在能够在成功易于验证的任务上取得进展，这引发了关于我们是否接近 RSI 的猜测。","但尽管这些评估是有帮助的，它们仅限于狭窄且可验证的任务。AI 研究可能更加开放式。成功往往不会立即显现或易于验证，而要取得进展，研究人员需要测试有前景的假设、回溯或考虑新的或非常规的方法。我们如何评估代理进行开放式 AI 研究的能力？我们在一篇新论文中采取了回答这一问题的第一步：https://cruxevals.com/crux/can-ai-agents-conduct-research。","我们与两篇未发表 AI 论文的作者合作，请他们起草各自论文的主要研究问题。然后，我们让前沿 AI 代理进行研究以回答这些问题，并为它们提供了数千美元的 API 积分和计算资源，以及六天的实际时间。原作者审查了代理的论文。","作者明确拒绝了两篇代理论文。为了更好地理解这些结果，我们团队花费了超过一百小时分析代理日志。我们的主要结论:","这些代理缺乏进行开放式研究的判断力。虽然代理提出的方向让专家评审觉得印象深刻，但他们很快因为低质量或合成数据而拒绝了这些提议方向。","代理缺乏对可用资源的意识。两次运行都在 API 预算使用不到 50% 且距离截止时间还有几小时的情况下结束，尽管代理可以监控他们的使用情况，并被鼓励花完预算。","代理没有创造性地回应反馈。尽管代理自身的 AI 自我评审发现了许多专家评审后来提出的问题，代理并没有创造性地解决这些问题。当面对负面反馈时，他们通过在现有发现中增加警告来回应，并在前景不佳的研究方向上加大投入。","代理没有有效地回溯。他们在实验的第一天就放弃了最雄心勃勃的研究目标，而且之后没有任何代理从根本上调整其方法。","代理没有遵循具体的指示。他们忽视了关于探索时间、使用 AI 自我评审工具获取反馈频率以及论文长度严格限制的明确规则。","我们希望评估 AI 进行开放式研究的能力已有两年时间，自从我们发布基准：https://arxiv.org/html/2409.11363v1 以研究是否可以用代理来提高可重复性。但我们希望首先把方法做好。我们的方法理念由论文中一些英国 AISI 合作者提出，并由我们普林斯顿核心团队完善。","我们称这些为“影子评估”，因为代理在原始研究中进行影子跟踪。除了我们两人外，核心团队还包括 Peter Kirgis、Andrew Schwartz 和 Stephan Rabanser。完整作者列表在本文末尾。","影子评估有重要优势：它们允许我们在代理未接受过训练且无法在线获取结果的情况下测试代理。同时，它们也允许花费数月回答问题的专家评估代理的输出。 1：#footnote-1","但影子评估也有其固有的局限性。专家评审知道论文是由 AI 生成的，他们可能会更偏好自己采取的方法，而不是代理采取的方法。由于我们对每篇论文都进行深入评估，样本量较小（在我们的研究中，我们仅使用了两篇论文）。而且这些评估必然涉及研究者在设计、执行和解释方面的诸多灵活性。","事实上，我们以某种特定立场而闻名：https://www.normaltech.ai/p/ai-as-normal-technology 在关于递归自我改进和超级智能的讨论中。这可能会影响我们如何开展研究。我们在论文中有详细的部分介绍我们的潜在偏见以及如何应对它们。我们寻求了一支并非完全共享我们先验观点的合作团队，并明确展示了由此产生的分歧。2：#footnote-2 针对未来的评估，我们希望“对抗性合作者”能成为核心团队的一部分。","对人工智能快速发展的影响","我们的结果表明，对前沿 AI 代理进行开放式研究仍然充满挑战。不过，这些发现仍是初步的，我们正在努力解决这些局限性，例如通过增加样本量、在新模型上进行测试以及可能的支架改进。但如果这些发现成立，其含义是什么？","首先，我们需要了解前沿 AI 进展（及 RSI）在多大程度上仅通过在可验证任务上的逐步优化就能实现。我们的观点是，虽然在狭义任务（例如提高效率）上确实可能加快进展，但我们认为这不会导致广泛的 RSI 或爆发性进展。不过，我们计划密切关注由于 AI 代理在可验证任务上的能力而引发的 AI 进展如何展开。","其次，我们需要衡量当前代理在进行开放式研究时的局限性（例如缺乏创造力和判断能力）能够多快被克服，例如通过更有针对性的训练和支架改进。我们计划定期继续进行影子评估，以帮助回答这个问题。","最后，即使这些局限性可以克服，可能还存在进一步的瓶颈，会减缓 AI 进展的速度。","瓶颈可能包括计算能力的限制、必须从真实世界实验中收集数据的需求，以及其他我们尚未发现的瓶颈，因为它们目前尚未阻碍进展。例如，高质量强化学习环境的重要性在它们对推理扩展有用之前并不明显。同样，为数据中心建设能源基础设施的重要性也没有被认识到，直到公司开始在用于训练和推理的数据中心上投资数千亿美元。","我们是否会遇到进一步的瓶颈，以及这些瓶颈的可解决程度如何，将对理解进展的速度产生重大影响。在这方面，我们的论文指出了一个尚未解决的瓶颈，即前沿智能体在开放性人工智能研究上的表现不佳（尽管是否在通向通用人工智能的关键路径上仍有待观察）。","如果我们处在一个能够轻松解决完全自动化研究瓶颈的世界中，我们应当预期通过提高人工智能能力可以获得人工智能进展的显著回报。但如果我们处在一个仍有许多难以克服的瓶颈的世界，阿姆达尔定律（Amdahl’s law：https://en.wikipedia.org/wiki/Amdahl%27s_law）将发挥作用：即使在易于应用人工智能的部分实现百倍加速，总体进展的速度仍只会略微提高，因为进展受最慢部分的速度所限制。3：#footnote-3","弄清我们所处的世界类型可能会极大影响对人工智能进展速度的估计。我们希望我们的研究结果有助于更深入地理解这些瓶颈。","在这里阅读论文：https://arxiv.org/pdf/2607.27191。作者是 Peter Kirgis、Sayash Kapoor、Andrew Schwartz、Stephan Rabanser、David Africa、Konstantinos Voudouris、Viet Nguyen、Toby Pilditch、Magda Dubois、Harry Coppock、Cozmin Ududec、Nitya Nadgir、Matilda Orona、Tilman Bayer、Derrick Chan-Sew、Yue Ling、Abhishek Shetty、Helen Toner、Gillian Hadfield、Seth Lazar、Steve Newman、Shoshannah Tekofsky、Rishi Bommasani 和 Arvind Narayanan。","这与提交论文进行盲审相反。不幸的是，人工智能领域的同行评审存在评审质量低下以及评审人与所分配论文之间专业匹配度不高的问题，这可能是由于提交数量的急剧增加所导致的：https://www.cs.cmu.edu/~nihars/preprints/SurveyPeerReview.pdf。","例如，我们的合著者对于这些智能体是否缺乏创造力，或者它们是否由于认知锁定而无法有效地整合反馈存在分歧。","在实践中，可能会发现某些类型的人工智能进展存在瓶颈，而其他类型则没有。例如，提高现有人工智能系统的效率和速度是一项可验证的任务，其进展非常迅速。因此，企业已经有效地利用智能体来提升人工智能系统的效率：https://x.com/OpenAI/status/2082878156483219672?s=20。在短期内，我们预计在有可验证信号的方面人工智能会快速进展。","有趣的是，这些智能体会“陷入循环”，并未重新考虑之前搁置的路径。“之前没成功，但或许这次会成功，也许我只需稍作调整……”似乎是根本的人类行为（尽管它经常失败）。而在耗尽预算之前放弃？那可不对！:-D","这只进一步证明，无论机器所使用的“思考”过程是什么，它都不是*人类*的。","研究是否通过提出研究问题来进行？我（杜威式观点）认为，将你的研究表述为问题会扭曲研究，除非你已经完成了研究。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"普林斯顿大学团队让前沿AI智能体用六天时间及数千美元API额度和算力，回答两篇未发表论文的核心问题。原作者明确拒绝了两份智能体论文，团队认为开放式研究仍具挑战。","background":"现有评估多聚焦结果容易验证的狭窄任务，而开放式AI研究需要判断假设是否值得推进，并在必要时回溯或尝试新方法。该团队与两篇未发表AI论文的作者合作，开展了此次影子评估。","viewpoint":"Aioga 判断，这项测试的重要性不在于证明智能体完全不能研究，而在于揭示可验证任务成绩与开放式研究能力之间可能存在差距。方向提议令人印象深刻，也未能弥补研究判断和执行上的不足。","implications":"值得关注的是，智能体未充分使用可用预算和时间，也未能针对自我审查发现的问题创造性调整，反而增加限定说明或继续投入前景不佳的方向。这可能限制其独立承担开放式研究的可靠性。","nextStep":"研究团队已明确将结果界定为初步发现，后续需要扩大样本量并持续评估。Aioga 判断，未来评估还应继续观察智能体能否有效利用资源、响应负面反馈、遵循具体指令，并在方向不佳时及时回溯。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-05T15:47:56.108Z","sourceHash":"4410021c6ff96b9c","review":{"approved":true,"groundedness":93,"clarity":90,"duplicationRisk":12,"blockingIssues":[],"notes":[]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["技巧观点","AI as Normal Technology（RSS）"],"translations":{"zh-CN":{"title":"AI智能体尚无法开展开放式AI研究","summary":"普林斯顿大学团队通过\"影子评估\"测试前沿AI智能体的开放式研究能力：让智能体在六天时间内、使用数千美元API额度和算力，回答两篇未发表论文的核心研究问题，结果两篇论文均被原作者明确拒绝。分析显示，智能体缺乏研究判断力、资源意识、创造性反馈应对能力和有效回溯能力，且未遵循具体指令。研究团队认为，开放式研究对前沿AI智能体仍具挑战性，但结果尚属初步，需扩大样本量并持续评估。","category":"技巧观点","source":"normaltech.ai","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI智能体尚无法开展开放式AI研究 - Aioga AI资讯","description":"普林斯顿大学团队通过\"影子评估\"测试前沿AI智能体的开放式研究能力：让智能体在六天时间内、使用数千美元API额度和算力，回答两篇未发表论文的核心研究问题，结果两篇论文均被原作者明确拒绝。分析显示，智能体缺乏研究判断力、资源意识、创造性反馈应对能力和有效回溯能力，且未遵循具体指令。研究团队认为，开放式研究对前沿AI智能体仍具挑战性，但结果尚属初步，需扩大样本量...","url":"https://www.aioga.com/news/cmsg6qeks07otrolg6som6vy0/","articleBody":["领导的目标：https://www.anthropic.com/institute/recursive-self-improvement AI 实验室：https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf 是递归自我改进（RSI）：使用 AI 代理自动化 AI 研究。RSI 也支撑着对 AI 爆炸性进展的预测：https://ai-2027.com/：https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/：https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion。我们如何评估是否接近这一里程碑？","一种方法是使用基准测试，测试代理是否可以进行 AI 研究。鉴于 AI 社区对基准测试的关注，它们一直是评估 RSI 进展的主要方式。在过去一年中，许多此类评估：https://cruxevals.com/crux/can-ai-agents-conduct-research#introduction:~:text=AI%20research%20questions.-,Selected,-evaluations%20and%20demonstrations 发现代理现在能够在成功易于验证的任务上取得进展，这引发了关于我们是否接近 RSI 的猜测。","但尽管这些评估是有帮助的，它们仅限于狭窄且可验证的任务。AI 研究可能更加开放式。成功往往不会立即显现或易于验证，而要取得进展，研究人员需要测试有前景的假设、回溯或考虑新的或非常规的方法。我们如何评估代理进行开放式 AI 研究的能力？我们在一篇新论文中采取了回答这一问题的第一步：https://cruxevals.com/crux/can-ai-agents-conduct-research。","我们与两篇未发表 AI 论文的作者合作，请他们起草各自论文的主要研究问题。然后，我们让前沿 AI 代理进行研究以回答这些问题，并为它们提供了数千美元的 API 积分和计算资源，以及六天的实际时间。原作者审查了代理的论文。","作者明确拒绝了两篇代理论文。为了更好地理解这些结果，我们团队花费了超过一百小时分析代理日志。我们的主要结论:","这些代理缺乏进行开放式研究的判断力。虽然代理提出的方向让专家评审觉得印象深刻，但他们很快因为低质量或合成数据而拒绝了这些提议方向。","代理缺乏对可用资源的意识。两次运行都在 API 预算使用不到 50% 且距离截止时间还有几小时的情况下结束，尽管代理可以监控他们的使用情况，并被鼓励花完预算。","代理没有创造性地回应反馈。尽管代理自身的 AI 自我评审发现了许多专家评审后来提出的问题，代理并没有创造性地解决这些问题。当面对负面反馈时，他们通过在现有发现中增加警告来回应，并在前景不佳的研究方向上加大投入。","代理没有有效地回溯。他们在实验的第一天就放弃了最雄心勃勃的研究目标，而且之后没有任何代理从根本上调整其方法。","代理没有遵循具体的指示。他们忽视了关于探索时间、使用 AI 自我评审工具获取反馈频率以及论文长度严格限制的明确规则。","我们希望评估 AI 进行开放式研究的能力已有两年时间，自从我们发布基准：https://arxiv.org/html/2409.11363v1 以研究是否可以用代理来提高可重复性。但我们希望首先把方法做好。我们的方法理念由论文中一些英国 AISI 合作者提出，并由我们普林斯顿核心团队完善。","我们称这些为“影子评估”，因为代理在原始研究中进行影子跟踪。除了我们两人外，核心团队还包括 Peter Kirgis、Andrew Schwartz 和 Stephan Rabanser。完整作者列表在本文末尾。","影子评估有重要优势：它们允许我们在代理未接受过训练且无法在线获取结果的情况下测试代理。同时，它们也允许花费数月回答问题的专家评估代理的输出。 1：#footnote-1","但影子评估也有其固有的局限性。专家评审知道论文是由 AI 生成的，他们可能会更偏好自己采取的方法，而不是代理采取的方法。由于我们对每篇论文都进行深入评估，样本量较小（在我们的研究中，我们仅使用了两篇论文）。而且这些评估必然涉及研究者在设计、执行和解释方面的诸多灵活性。","事实上，我们以某种特定立场而闻名：https://www.normaltech.ai/p/ai-as-normal-technology 在关于递归自我改进和超级智能的讨论中。这可能会影响我们如何开展研究。我们在论文中有详细的部分介绍我们的潜在偏见以及如何应对它们。我们寻求了一支并非完全共享我们先验观点的合作团队，并明确展示了由此产生的分歧。2：#footnote-2 针对未来的评估，我们希望“对抗性合作者”能成为核心团队的一部分。","对人工智能快速发展的影响","我们的结果表明，对前沿 AI 代理进行开放式研究仍然充满挑战。不过，这些发现仍是初步的，我们正在努力解决这些局限性，例如通过增加样本量、在新模型上进行测试以及可能的支架改进。但如果这些发现成立，其含义是什么？","首先，我们需要了解前沿 AI 进展（及 RSI）在多大程度上仅通过在可验证任务上的逐步优化就能实现。我们的观点是，虽然在狭义任务（例如提高效率）上确实可能加快进展，但我们认为这不会导致广泛的 RSI 或爆发性进展。不过，我们计划密切关注由于 AI 代理在可验证任务上的能力而引发的 AI 进展如何展开。","其次，我们需要衡量当前代理在进行开放式研究时的局限性（例如缺乏创造力和判断能力）能够多快被克服，例如通过更有针对性的训练和支架改进。我们计划定期继续进行影子评估，以帮助回答这个问题。","最后，即使这些局限性可以克服，可能还存在进一步的瓶颈，会减缓 AI 进展的速度。","瓶颈可能包括计算能力的限制、必须从真实世界实验中收集数据的需求，以及其他我们尚未发现的瓶颈，因为它们目前尚未阻碍进展。例如，高质量强化学习环境的重要性在它们对推理扩展有用之前并不明显。同样，为数据中心建设能源基础设施的重要性也没有被认识到，直到公司开始在用于训练和推理的数据中心上投资数千亿美元。","我们是否会遇到进一步的瓶颈，以及这些瓶颈的可解决程度如何，将对理解进展的速度产生重大影响。在这方面，我们的论文指出了一个尚未解决的瓶颈，即前沿智能体在开放性人工智能研究上的表现不佳（尽管是否在通向通用人工智能的关键路径上仍有待观察）。","如果我们处在一个能够轻松解决完全自动化研究瓶颈的世界中，我们应当预期通过提高人工智能能力可以获得人工智能进展的显著回报。但如果我们处在一个仍有许多难以克服的瓶颈的世界，阿姆达尔定律（Amdahl’s law：https://en.wikipedia.org/wiki/Amdahl%27s_law）将发挥作用：即使在易于应用人工智能的部分实现百倍加速，总体进展的速度仍只会略微提高，因为进展受最慢部分的速度所限制。3：#footnote-3","弄清我们所处的世界类型可能会极大影响对人工智能进展速度的估计。我们希望我们的研究结果有助于更深入地理解这些瓶颈。","在这里阅读论文：https://arxiv.org/pdf/2607.27191。作者是 Peter Kirgis、Sayash Kapoor、Andrew Schwartz、Stephan Rabanser、David Africa、Konstantinos Voudouris、Viet Nguyen、Toby Pilditch、Magda Dubois、Harry Coppock、Cozmin Ududec、Nitya Nadgir、Matilda Orona、Tilman Bayer、Derrick Chan-Sew、Yue Ling、Abhishek Shetty、Helen Toner、Gillian Hadfield、Seth Lazar、Steve Newman、Shoshannah Tekofsky、Rishi Bommasani 和 Arvind Narayanan。","这与提交论文进行盲审相反。不幸的是，人工智能领域的同行评审存在评审质量低下以及评审人与所分配论文之间专业匹配度不高的问题，这可能是由于提交数量的急剧增加所导致的：https://www.cs.cmu.edu/~nihars/preprints/SurveyPeerReview.pdf。","例如，我们的合著者对于这些智能体是否缺乏创造力，或者它们是否由于认知锁定而无法有效地整合反馈存在分歧。","在实践中，可能会发现某些类型的人工智能进展存在瓶颈，而其他类型则没有。例如，提高现有人工智能系统的效率和速度是一项可验证的任务，其进展非常迅速。因此，企业已经有效地利用智能体来提升人工智能系统的效率：https://x.com/OpenAI/status/2082878156483219672?s=20。在短期内，我们预计在有可验证信号的方面人工智能会快速进展。","有趣的是，这些智能体会“陷入循环”，并未重新考虑之前搁置的路径。“之前没成功，但或许这次会成功，也许我只需稍作调整……”似乎是根本的人类行为（尽管它经常失败）。而在耗尽预算之前放弃？那可不对！:-D","这只进一步证明，无论机器所使用的“思考”过程是什么，它都不是*人类*的。","研究是否通过提出研究问题来进行？我（杜威式观点）认为，将你的研究表述为问题会扭曲研究，除非你已经完成了研究。"]},"en":{"title":"AI agents are still unable to conduct open-ended AI research","summary":"A team from Princeton University tested the open-ended research capabilities of cutting-edge AI agents through 'shadow evaluation': allowing the agents to answer the core research questions of two unpublished papers within six days, using thousands of dollars in API credits and computing power. The results showed that both papers were explicitly rejected by the original authors. Analysis indicated that the agents lacked research judgment, resource awareness, the ability to respond with creative feedback, and effective backtracking skills, and did not follow specific instructions. The research team believes that open-ended research remains challenging for cutting-edge AI agents, but the results are still preliminary and require expanding the sample size and continuous evaluation.","category":"Insights","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI agents are still unable to conduct open-ended AI research - Aioga AI News","description":"A team from Princeton University tested the open-ended research capabilities of cutting-edge AI agents through 'shadow evaluation': allowing the agents to answer the core research...","url":"https://www.aioga.com/en/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:23:20.314Z"},"ja":{"title":"AI知能体はまだオープン型AI研究を行うことができない","summary":"プリンストン大学のチームは、「シャドウ評価」を通じて、最先端AIエージェントのオープンリサーチ能力をテストした。エージェントに6日間の時間、数千ドル相当のAPIクレジットと計算資源を使わせ、未発表の2本の論文の核心的研究問題に回答させたところ、両方の論文は原著者によって明確に拒否された。分析の結果、エージェントは研究判断力、リソース意識、創造的フィードバックへの対応能力、効果的な振り返り能力を欠き、具体的な指示には従わなかったことが明らかになった。研究チームは、オープンリサーチは最先端AIエージェントにとってなお挑戦であると考えているが、結果はまだ初期的であり、サンプル数を増やし継続的に評価する必要がある。","category":"ヒントと視点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI知能体はまだオープン型AI研究を行うことができない - Aioga AIニュース","description":"プリンストン大学のチームは、「シャドウ評価」を通じて、最先端AIエージェントのオープンリサーチ能力をテストした。エージェントに6日間の時間、数千ドル相当のAPIクレジットと計算資源を使わせ、未発表の2本の論文の核心的研究問題に回答させたところ、両方の論文は原著者によって明確に拒否された。分析の結果、エージェントは研究判断力、リソース意識、創造的フィードバック...","url":"https://www.aioga.com/ja/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:23:28.938Z"},"ko":{"title":"AI 지능체는 아직 개방형 AI 연구를 진행할 수 없다","summary":"프린스턴 대학 팀은 '그림자 평가'를 통해 최첨단 AI 에이전트의 개방형 연구 능력을 테스트했습니다. 에이전트가 6일 동안 수천 달러의 API 한도와 컴퓨팅 자원을 사용하여 두 편의 미발표 논문의 핵심 연구 질문에 답하도록 했지만, 결과적으로 두 논문 모두 원저자에 의해 명확히 거부되었습니다. 분석 결과, 에이전트는 연구 판단력, 자원 인식, 창의적 피드백 대응 능력, 효과적인 회귀 능력이 부족했으며, 구체적인 지침을 따르지 않았습니다. 연구팀은 개방형 연구가 최첨단 AI 에이전트에게 여전히 도전 과제임을 인정하지만, 이번 결과는 초기 단계에 불과하며, 표본 수를 늘리고 지속적으로 평가할 필요가 있다고 밝혔습니다.","category":"인사이트","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI 지능체는 아직 개방형 AI 연구를 진행할 수 없다 - Aioga AI 뉴스","description":"프린스턴 대학 팀은 '그림자 평가'를 통해 최첨단 AI 에이전트의 개방형 연구 능력을 테스트했습니다. 에이전트가 6일 동안 수천 달러의 API 한도와 컴퓨팅 자원을 사용하여 두 편의 미발표 논문의 핵심 연구 질문에 답하도록 했지만, 결과적으로 두 논문 모두 원저자에 의해 명확히 거부되었습니다. 분석 결과, 에이전트는 연구...","url":"https://www.aioga.com/ko/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:24:20.425Z"},"es":{"title":"Los agentes de inteligencia artificial aún no pueden realizar investigaciones de IA abiertas","summary":"Un equipo de la Universidad de Princeton evaluó la capacidad de investigación abierta de agentes de IA de vanguardia mediante pruebas de \"evaluación sombra\": permitió que los agentes respondieran, en seis días, a las preguntas centrales de dos artículos no publicados, usando miles de dólares en créditos de API y potencia computacional. Como resultado, ambos artículos fueron explícitamente rechazados por los autores originales. El análisis mostró que los agentes carecían de juicio de investigación, conciencia de recursos, capacidad de respuesta a retroalimentación creativa y habilidades de retroceso efectivas, además de no seguir instrucciones específicas. El equipo de investigación considera que la investigación abierta sigue siendo un desafío para los agentes de IA de vanguardia, pero los resultados son preliminares y se necesita aumentar el tamaño de la muestra y continuar evaluando.","category":"Ideas","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Los agentes de inteligencia artificial aún no pueden realizar investigaciones de IA abiertas - Aioga Noticias de IA","description":"Un equipo de la Universidad de Princeton evaluó la capacidad de investigación abierta de agentes de IA de vanguardia mediante pruebas de \"evaluación sombra\": permitió que los agent...","url":"https://www.aioga.com/es/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:24:16.945Z"},"fr":{"title":"Les agents d'IA ne peuvent pas encore mener de recherches en IA ouvertes","summary":"Une équipe de l'université de Princeton a testé la capacité de recherche ouverte des agents d'intelligence artificielle de pointe via une « évaluation fantôme » : les agents devaient, en six jours, en utilisant plusieurs milliers de dollars de crédits API et de puissance de calcul, répondre aux questions centrales de deux articles non publiés, mais les deux articles ont été expressément rejetés par leurs auteurs originaux. L'analyse montre que les agents manquent de jugement en recherche, de conscience des ressources, de capacité à répondre de manière créative aux rétroactions et de capacité de rétroaction efficace, et qu'ils n'ont pas suivi les instructions spécifiques. L'équipe de recherche estime que la recherche ouverte reste un défi pour les agents d'IA de pointe, mais que les résultats sont encore préliminaires et qu'il est nécessaire d'élargir l'échantillon et de poursuivre l'évaluation.","category":"Analyses","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Les agents d'IA ne peuvent pas encore mener de recherches en IA ouvertes - Aioga Actualités IA","description":"Une équipe de l'université de Princeton a testé la capacité de recherche ouverte des agents d'intelligence artificielle de pointe via une « évaluation fantôme » : les agents devaie...","url":"https://www.aioga.com/fr/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:25:15.400Z"},"de":{"title":"KI-Agenten sind noch nicht in der Lage, offene KI-Forschung durchzuführen","summary":"Ein Team der Princeton University testete die Open-Research-Fähigkeiten modernster KI-Agenten durch sogenannte 'Shadow-Evaluation': Die Agenten sollten innerhalb von sechs Tagen unter Nutzung von API-Guthaben im Wert von mehreren Tausend US-Dollar und Rechenleistung die Kernforschungsfragen zweier unveröffentlichter wissenschaftlicher Arbeiten beantworten. Das Ergebnis: Beide Arbeiten wurden von den Originalautoren ausdrücklich abgelehnt. Die Analyse zeigt, dass den Agenten Forschungsurteilskraft, Bewusstsein für Ressourcen, Fähigkeit zu kreativem Feedback und effektiver Rückverfolgbarkeit fehlt und sie den konkreten Anweisungen nicht folgen. Das Forschungsteam ist der Ansicht, dass Open-Research für modernste KI-Agenten nach wie vor eine Herausforderung darstellt, die Ergebnisse bisher jedoch vorläufig sind und eine größere Stichprobe sowie kontinuierliche Bewertung erforderlich machen.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"KI-Agenten sind noch nicht in der Lage, offene KI-Forschung durchzuführen - Aioga KI-News","description":"Ein Team der Princeton University testete die Open-Research-Fähigkeiten modernster KI-Agenten durch sogenannte 'Shadow-Evaluation': Die Agenten sollten innerhalb von sechs Tagen un...","url":"https://www.aioga.com/de/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:25:10.534Z"},"pt-BR":{"title":"Agentes de IA ainda não podem realizar pesquisas de IA abertas","summary":"A equipe da Universidade de Princeton testou a capacidade de pesquisa aberta de agentes de IA avançados por meio de 'avaliação sombra': eles fizeram os agentes responderem, em seis dias, usando milhares de dólares em créditos de API e poder computacional, às questões centrais de duas pesquisas inéditas, e em ambos os casos os trabalhos foram explicitamente rejeitados pelos autores originais. A análise mostrou que os agentes carecem de julgamento de pesquisa, consciência de recursos, capacidade de resposta criativa e habilidade de retroceder de forma eficaz, além de não seguir instruções específicas. A equipe de pesquisa acredita que a pesquisa aberta ainda é desafiadora para agentes de IA avançados, mas os resultados são preliminares, sendo necessário ampliar o tamanho da amostra e continuar a avaliação.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Agentes de IA ainda não podem realizar pesquisas de IA abertas - Aioga Notícias de IA","description":"A equipe da Universidade de Princeton testou a capacidade de pesquisa aberta de agentes de IA avançados por meio de 'avaliação sombra': eles fizeram os agentes responderem, em seis...","url":"https://www.aioga.com/pt-BR/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:26:03.128Z"},"ru":{"title":"ИИ-агенты пока не могут проводить открытые исследования в области ИИ","summary":"Команда Принстонского университета протестировала способность передовых ИИ-агентов к открытому исследованию с помощью «тестирования в тени»: агентам предоставляли шесть дней, тысячи долларов на API и вычислительные ресурсы, чтобы ответить на ключевые исследовательские вопросы двух неопубликованных статей, и в итоге обе статьи были явно отклонены их авторами. Анализ показал, что агентам не хватает исследовательского суждения, осознания ресурсов, способности к творческому реагированию на обратную связь и эффективной ретроспективной оценки, а также они не следовали конкретным инструкциям. Исследовательская команда полагает, что открытые исследования по-прежнему представляют трудность для передовых ИИ-агентов, но результаты пока предварительные, требуется увеличить размер выборки и продолжать оценку.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"ИИ-агенты пока не могут проводить открытые исследования в области ИИ - Aioga Новости ИИ","description":"Команда Принстонского университета протестировала способность передовых ИИ-агентов к открытому исследованию с помощью «тестирования в тени»: агентам предоставляли шесть дней, тысяч...","url":"https://www.aioga.com/ru/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:26:02.045Z"},"ar":{"title":"الوكلاء الذكاء الاصطناعي لا يستطيعون بعد إجراء أبحاث الذكاء الاصطناعي المفتوحة","summary":"قام فريق جامعة برينستون باختبار القدرة البحثية المفتوحة لوكلاء الذكاء الاصطناعي المتقدمين من خلال \"التقييم الظلي\": حيث طلب من الوكلاء، خلال ستة أيام وباستخدام ميزانية API وقوة حوسبة بقيمة عدة آلاف من الدولارات، الإجابة على الأسئلة البحثية الأساسية لمقالتين غير منشورتين، وكانت النتيجة رفض المؤلفان الأصليان للمقالتين بوضوح. أظهرت التحليلات أن الوكلاء يفتقرون إلى حكم البحث، والوعي بالموارد، وقدرة التعامل مع التغذية الراجعة الإبداعية، وقدرة الاسترجاع الفعّالة، ولم يتبعوا التعليمات المحددة. ويعتقد فريق البحث أن البحث المفتوح لا يزال يمثل تحديًا لوكلاء الذكاء الاصطناعي المتقدمين، لكن النتائج لا تزال أولية، وتحتاج إلى زيادة حجم العينة والاستمرار في التقييم.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"الوكلاء الذكاء الاصطناعي لا يستطيعون بعد إجراء أبحاث الذكاء الاصطناعي المفتوحة - Aioga أخبار الذكاء الاصطناعي","description":"قام فريق جامعة برينستون باختبار القدرة البحثية المفتوحة لوكلاء الذكاء الاصطناعي المتقدمين من خلال \"التقييم الظلي\": حيث طلب من الوكلاء، خلال ستة أيام وباستخدام ميزانية API وقوة حوسب...","url":"https://www.aioga.com/ar/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:26:56.465Z"},"hi":{"title":"एआई बुद्धिमान एजेंट अभी खुले एआई शोध को शुरू नहीं कर सकता","summary":"प्रिंसटन विश्वविद्यालय की टीम ने \"छाया मूल्यांकन\" के माध्यम से उन्नत एआई एजेंट की मुक्त-संयोजित शोध क्षमता का परीक्षण किया: एजेंट को छह दिनों में, हजारों डॉलर के एपीआई क्रेडिट और कंप्यूटिंग शक्ति का उपयोग करके, दो अप्रकाशित शोध पत्रों के मुख्य शोध प्रश्नों का उत्तर देना था, परिणामस्वरूप दोनों पेपरों को उनके मूल लेखकों द्वारा स्पष्ट रूप से अस्वीकार कर दिया गया। विश्लेषण से पता चलता है कि एजेंट में शोध निर्णय क्षमता, संसाधन जागरूकता, रचनात्मक प्रतिक्रिया क्षमता और प्रभावी प्रतिपुष्टि क्षमता की कमी है, और उसने विशिष्ट निर्देशों का पालन नहीं किया। शोध टीम का मानना है कि मुक्त-संयोजित शोध उन्नत एआई एजेंट के लिए अभी भी चुनौतीपूर्ण है, लेकिन परिणाम प्रारंभिक हैं, और नमूने के आकार को बढ़ाकर तथा लगातार मूल्यांकन करके आगे बढ़ाने की आवश्यकता है।","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"एआई बुद्धिमान एजेंट अभी खुले एआई शोध को शुरू नहीं कर सकता - Aioga AI समाचार","description":"प्रिंसटन विश्वविद्यालय की टीम ने \"छाया मूल्यांकन\" के माध्यम से उन्नत एआई एजेंट की मुक्त-संयोजित शोध क्षमता का परीक्षण किया: एजेंट को छह दिनों में, हजारों डॉलर के एपीआई क्रेडिट और क...","url":"https://www.aioga.com/hi/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:26:57.766Z"},"it":{"title":"Gli agenti di intelligenza artificiale non sono ancora in grado di condurre ricerche AI aperte","summary":"Un team dell'Università di Princeton ha testato la capacità di ricerca aperta degli agenti AI all'avanguardia attraverso la \"valutazione ombra\": gli agenti hanno dovuto, in sei giorni, utilizzare migliaia di dollari in crediti API e potenza di calcolo per rispondere ai principali quesiti di ricerca di due articoli non ancora pubblicati, ma entrambi gli articoli sono stati chiaramente rifiutati dagli autori originali. L'analisi mostra che gli agenti mancano di giudizio di ricerca, consapevolezza delle risorse, capacità di rispondere in modo creativo ai feedback e capacità di revisione efficace, e non hanno seguito istruzioni specifiche. Il team di ricerca ritiene che la ricerca aperta rimanga una sfida per gli agenti AI all'avanguardia, ma i risultati sono ancora preliminari e sarà necessario aumentare il numero dei campioni e continuare la valutazione.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Gli agenti di intelligenza artificiale non sono ancora in grado di condurre ricerche AI aperte - Aioga Notizie IA","description":"Un team dell'Università di Princeton ha testato la capacità di ricerca aperta degli agenti AI all'avanguardia attraverso la \"valutazione ombra\": gli agenti hanno dovuto, in sei gio...","url":"https://www.aioga.com/it/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:27:47.174Z"},"nl":{"title":"AI-agents zijn nog niet in staat om open AI-onderzoek uit te voeren","summary":"Het team van de Universiteit van Princeton heeft de open onderzoeksvaardigheden van geavanceerde AI-agenten getest via een 'shadow evaluation': de agenten kregen zes dagen de tijd, met gebruik van enkele duizenden dollars aan API-tegoed en rekenkracht, om de kernonderzoeksvragen van twee ongepubliceerde artikelen te beantwoorden. Het resultaat was dat beide artikelen expliciet door de oorspronkelijke auteurs werden afgewezen. Analyse toonde aan dat de agenten gebrek hadden aan onderzoeksvaardigheid, bewustzijn van middelen, het vermogen om creatief op feedback te reageren en effectieve terugkoppelingsmogelijkheden, en dat zij specifieke instructies niet volgden. Het onderzoeksteam is van mening dat open onderzoek voor geavanceerde AI-agenten nog steeds een uitdaging vormt, maar de resultaten zijn voorlopig; er is een grotere steekproef en voortdurende evaluatie nodig.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI-agents zijn nog niet in staat om open AI-onderzoek uit te voeren - Aioga AI-nieuws","description":"Het team van de Universiteit van Princeton heeft de open onderzoeksvaardigheden van geavanceerde AI-agenten getest via een 'shadow evaluation': de agenten kregen zes dagen de tijd,...","url":"https://www.aioga.com/nl/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:27:46.819Z"},"tr":{"title":"AI zekâ ajanı henüz açık uçlu AI araştırmaları yapamıyor","summary":"Princeton Üniversitesi ekibi, öncü yapay zeka ajanlarının açık uçlu araştırma yeteneklerini \"gölge değerlendirme\" yöntemi ile test etti: Ajanın, altı gün içinde, binlerce dolarlık API kredisi ve hesaplama gücü kullanarak yayımlanmamış iki makalenin temel araştırma sorularını yanıtlaması sağlandı; sonuçta her iki makale de orijinal yazarlar tarafından açıkça reddedildi. Analiz, ajanın araştırma değerlendirme yeteneğinden, kaynak bilincinden, yaratıcı geri bildirimle başa çıkabilme kapasitesinden ve etkili geriye dönük izleme yeteneğinden yoksun olduğunu ve belirli talimatlara uymadığını gösterdi. Araştırma ekibi, açık uçlu araştırmanın öncü yapay zeka ajanları için hâlâ zorlu olduğunu ancak sonuçların henüz öncül olduğunu, örneklem hacminin artırılması ve sürekli değerlendirme yapılması gerektiğini belirtti.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"AI zekâ ajanı henüz açık uçlu AI araştırmaları yapamıyor - Aioga AI Haberleri","description":"Princeton Üniversitesi ekibi, öncü yapay zeka ajanlarının açık uçlu araştırma yeteneklerini \"gölge değerlendirme\" yöntemi ile test etti: Ajanın, altı gün içinde, binlerce dolarlık...","url":"https://www.aioga.com/tr/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:28:41.561Z"},"vi":{"title":"Các tác nhân AI vẫn chưa thể tiến hành nghiên cứu AI mở","summary":"Nhóm nghiên cứu tại Đại học Princeton đã kiểm tra khả năng nghiên cứu mở của các tác nhân AI tiên tiến thông qua thử nghiệm “đánh giá bóng tối”: cho các tác nhân trong vòng sáu ngày, sử dụng hàng nghìn đô la hạn mức API và tính toán, trả lời các câu hỏi nghiên cứu cốt lõi của hai bài báo chưa được công bố, kết quả cả hai bài báo đều bị tác giả gốc từ chối rõ ràng. Phân tích cho thấy, các tác nhân thiếu khả năng đánh giá nghiên cứu, nhận thức về nguồn lực, khả năng phản hồi sáng tạo và khả năng truy hồi hiệu quả, và không tuân theo các chỉ dẫn cụ thể. Nhóm nghiên cứu cho rằng, nghiên cứu mở vẫn là thách thức đối với các tác nhân AI tiên tiến, nhưng kết quả vẫn mang tính sơ bộ, cần mở rộng kích thước mẫu và đánh giá liên tục.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Các tác nhân AI vẫn chưa thể tiến hành nghiên cứu AI mở - Tin tức AI Aioga","description":"Nhóm nghiên cứu tại Đại học Princeton đã kiểm tra khả năng nghiên cứu mở của các tác nhân AI tiên tiến thông qua thử nghiệm “đánh giá bóng tối”: cho các tác nhân trong vòng sáu ngà...","url":"https://www.aioga.com/vi/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:28:37.824Z"},"id":{"title":"Agen AI belum bisa melakukan penelitian AI terbuka","summary":"Tim Universitas Princeton menguji kemampuan penelitian terbuka dari agen AI mutakhir melalui 'penilaian bayangan': membiarkan agen menjawab pertanyaan penelitian inti dari dua makalah yang belum diterbitkan dalam waktu enam hari, menggunakan ribuan dolar kuota API dan daya komputasi. Hasilnya, kedua makalah tersebut secara tegas ditolak oleh penulis asli. Analisis menunjukkan bahwa agen kekurangan penilaian penelitian, kesadaran sumber daya, kemampuan menanggapi umpan balik secara kreatif, dan kemampuan penelusuran efektif, serta tidak mengikuti instruksi tertentu. Tim peneliti berpendapat bahwa penelitian terbuka tetap menjadi tantangan bagi agen AI mutakhir, tetapi hasilnya masih bersifat awal, sehingga perlu memperluas jumlah sampel dan evaluasi berkelanjutan.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Agen AI belum bisa melakukan penelitian AI terbuka - Berita AI Aioga","description":"Tim Universitas Princeton menguji kemampuan penelitian terbuka dari agen AI mutakhir melalui 'penilaian bayangan': membiarkan agen menjawab pertanyaan penelitian inti dari dua maka...","url":"https://www.aioga.com/id/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:29:32.294Z"},"th":{"title":"ปัญญาประดิษฐ์ยังไม่สามารถดำเนินการวิจัย AI แบบเปิดได้","summary":"ทีมงานของมหาวิทยาลัยพรินซ์ตันทดสอบความสามารถในการวิจัยแบบเปิดของ AI ขั้นสูงผ่านการ 'ประเมินเงา': ให้เอเจนต์ AI ตอบคำถามวิจัยหลักของบทความวิจัยที่ยังไม่ตีพิมพ์สองบทความภายในเวลา 6 วัน โดยใช้เครดิต API และกำลังประมวลผลมูลค่าหลายพันดอลลาร์ ผลปรากฏว่าทั้งสองบทความถูกผู้เขียนต้นฉบับปฏิเสธโดยชัดเจน การวิเคราะห์แสดงให้เห็นว่าเอเจนต์ขาดความสามารถในการตัดสินใจด้านการวิจัย ความตระหนักถึงทรัพยากร ความสามารถในการตอบสนองต่อความคิดเห็นสร้างสรรค์ และความสามารถในการย้อนกลับอย่างมีประสิทธิภาพ อีกทั้งยังไม่ปฏิบัติตามคำสั่งที่เจาะจง ทีมวิจัยเห็นว่า การวิจัยแบบเปิดยังเป็นความท้าทายสำหรับ AI ขั้นสูง แต่ผลลัพธ์ยังอยู่ในขั้นต้น จำเป็นต้องขยายขนาดตัวอย่างและประเมินอย่างต่อเนื่อง","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"ปัญญาประดิษฐ์ยังไม่สามารถดำเนินการวิจัย AI แบบเปิดได้ - ข่าว AI Aioga","description":"ทีมงานของมหาวิทยาลัยพรินซ์ตันทดสอบความสามารถในการวิจัยแบบเปิดของ AI ขั้นสูงผ่านการ 'ประเมินเงา': ให้เอเจนต์ AI ตอบคำถามวิจัยหลักของบทความวิจัยที่ยังไม่ตีพิมพ์สองบทความภายในเวลา 6 ว...","url":"https://www.aioga.com/th/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:29:40.360Z"},"pl":{"title":"Agent AI nie jest jeszcze w stanie prowadzić otwartych badań nad AI","summary":"Zespół Uniwersytetu Princeton przetestował zdolność do otwartych badań najnowocześniejszych agentów AI poprzez tzw. „shadow evaluation”: pozwolono agentowi w ciągu sześciu dni, korzystając z kilku tysięcy dolarów na API i mocy obliczeniowej, odpowiedzieć na kluczowe pytania badawcze w dwóch nieopublikowanych artykułach, a wynik był taki, że oba artykuły zostały wyraźnie odrzucone przez ich oryginalnych autorów. Analiza wykazała, że agentowi brakowało zdolności badawczych, świadomości zasobów, umiejętności reagowania na kreatywne opinie oraz skutecznej zdolności retrospektywnej, a także nie przestrzegał konkretnych instrukcji. Zespół badawczy uważa, że otwarte badania nadal stanowią wyzwanie dla najnowocześniejszych agentów AI, jednak wyniki są wstępne i wymagają zwiększenia liczebności próby oraz ciągłej oceny.","category":"技巧观点","source":"AI as Normal Technology（RSS）","aggregationSource":"AI as Normal Technology（RSS）","pageTitle":"Agent AI nie jest jeszcze w stanie prowadzić otwartych badań nad AI - Aioga Wiadomości AI","description":"Zespół Uniwersytetu Princeton przetestował zdolność do otwartych badań najnowocześniejszych agentów AI poprzez tzw. „shadow evaluation”: pozwolono agentowi w ciągu sześciu dni, kor...","url":"https://www.aioga.com/pl/news/cmsg6qeks07otrolg6som6vy0/","contentTranslated":true,"sourceHash":"7ee695d1a10908be","translatedAt":"2026-08-05T15:30:44.257Z"}}}}