领导的目标:https://www.anthropic.com/institute/recursive-self-improvement AI 实验室:https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf 是递归自我改进(RSI):使用 AI 代理自动化 AI 研究。RSI 也支撑着对 AI 爆炸性进展的预测:https://ai-2027.com/:https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/:https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion。我们如何评估是否接近这一里程碑?
一种方法是使用基准测试,测试代理是否可以进行 AI 研究。鉴于 AI 社区对基准测试的关注,它们一直是评估 RSI 进展的主要方式。在过去一年中,许多此类评估:https://cruxevals.com/crux/can-ai-agents-conduct-research#introduction:~:text=AI%20research%20questions.-,Selected,-evaluations%20and%20demonstrations 发现代理现在能够在成功易于验证的任务上取得进展,这引发了关于我们是否接近 RSI 的猜测。
但尽管这些评估是有帮助的,它们仅限于狭窄且可验证的任务。AI 研究可能更加开放式。成功往往不会立即显现或易于验证,而要取得进展,研究人员需要测试有前景的假设、回溯或考虑新的或非常规的方法。我们如何评估代理进行开放式 AI 研究的能力?我们在一篇新论文中采取了回答这一问题的第一步:https://cruxevals.com/crux/can-ai-agents-conduct-research。
我们与两篇未发表 AI 论文的作者合作,请他们起草各自论文的主要研究问题。然后,我们让前沿 AI 代理进行研究以回答这些问题,并为它们提供了数千美元的 API 积分和计算资源,以及六天的实际时间。原作者审查了代理的论文。
我们的结果表明,对前沿 AI 代理进行开放式研究仍然充满挑战。不过,这些发现仍是初步的,我们正在努力解决这些局限性,例如通过增加样本量、在新模型上进行测试以及可能的支架改进。但如果这些发现成立,其含义是什么?
首先,我们需要了解前沿 AI 进展(及 RSI)在多大程度上仅通过在可验证任务上的逐步优化就能实现。我们的观点是,虽然在狭义任务(例如提高效率)上确实可能加快进展,但我们认为这不会导致广泛的 RSI 或爆发性进展。不过,我们计划密切关注由于 AI 代理在可验证任务上的能力而引发的 AI 进展如何展开。
The goal of leading :https://www.anthropic.com/institute/recursive-self-improvement AI labs :https://cdn.openai.com/pdf/25752ecb-0e5c-47f9-b9e4-c0f4d76f8d3d/a-blueprint-for-a-federal-framework.pdf is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts :https://ai-2027.com/ of explosive :https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/ AI :https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion progress. How can we assess if we are close to this milestone?
One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations :https://cruxevals.com/crux/can-ai-agents-conduct-research#introduction:~:text=AI%20research%20questions.-,Selected,-evaluations%20and%20demonstrations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of RSI.
But while these evaluations are helpful, they are limited to narrow, verifiable tasks. AI research can be much more open-ended. Success is often not immediately clear or verifiable, and to make progress, researchers need to test promising hypotheses, backtrack, or consider new or unconventional approaches. How can we evaluate agents’ ability to conduct open-ended AI research? We take our first step towards answering this question in a new paper :https://cruxevals.com/crux/can-ai-agents-conduct-research .
We partnered with the authors of two unpublished AI papers and asked them to draft their papers’ main research questions. We then tasked frontier AI agents with conducting research to answer these questions, and gave them thousands of dollars of API credits and compute, and six days of wall-clock time. The original authors reviewed the agents’ papers.
The authors unambiguously rejected both agent papers. To better understand these results, our team spent over a hundred hours analyzing the agents’ logs. Our main takeaways:
The agents lacked the judgment for conducting open-ended research. While the agents proposed directions the expert reviewers found impressive, they quickly rejected their proposed directions based on low-quality or synthetic data.
The agents lacked awareness about the resources available to them. Both runs ended with less than 50% of the API budget spent and with hours left before the deadline, even though the agents could monitor their usage and were encouraged to spend down their budgets.
The agents did not creatively respond to feedback. Despite the agents’ own AI self-reviews surfacing many of the issues that the expert reviewers later raised, the agents did not creatively address these concerns. When faced with negative feedback they responded by adding caveats to existing findings, and doubled down on unpromising research directions.
The agents did not effectively backtrack. They retired their most ambitious research targets within the first day of the experiment, and neither agent fundamentally shifted its approach after that point.
The agents did not follow concrete instructions. They ignored explicit rules about how much time to spend on exploration, how often to get reviews from AI self-review tools, and strict limits on paper length.
We have wanted to evaluate AI’s ability to conduct open-ended research for two years, ever since we released a benchmark :https://arxiv.org/html/2409.11363v1 to study if agents could be used to improve reproducibility. But we wanted to get our method right. The idea behind our method was suggested by some of the UK AISI coauthors of the paper and refined by our core team at Princeton.
We call these “shadow evaluations” since the agent shadows the original study. In addition to the two of us, the core team comprises Peter Kirgis, Andrew Schwartz, and Stephan Rabanser. The full author list is at the end of this essay.
Shadow evaluations have important advantages: they allow us to test agents on results they haven’t been trained on and can’t access online. They also allow experts who have spent months answering the questions to evaluate agents’ outputs. 1:#footnote-1
But shadow evaluations also have inherent limitations. Expert reviewers know that the paper is AI-generated, and they might prefer the approach they took over the one that the agent took. Because we are conducting in-depth evaluations of each paper, the sample size is small (in our study, we used just two papers). And these evaluations necessarily involve a lot of researcher flexibility in design, execution, and interpretation.
In fact, we are known for a particular position :https://www.normaltech.ai/p/ai-as-normal-technology in the debate on recursive self-improvement and superintelligence. This could influence how we conduct the research. We have a detailed section in the paper on our potential biases and how we address them. We sought out a team of collaborators who don’t all share our priors, and we explicitly surface the disagreements that resulted. 2:#footnote-2 For future evaluations, we are interested in having “adversarial collaborators” as part of the core team.
Implications for explosive AI progress
Our results suggest that conducting open-ended research remains challenging for frontier AI agents. Still, these findings are tentative, and we are working to address the limitations, such as by increasing the sample size, testing with new models, and through potential scaffold improvements. But if these findings hold up, what are the implications?
First, we need to understand the extent to which frontier AI progress (and RSI) can be achieved simply by hill climbing on verifiable tasks. Our view is that while faster progress is certainly possible on narrow tasks (such as improving efficiency), we don’t think it will lead to broad RSI or explosive progress. Still, we plan to closely follow how AI progress unfolds as a result of AI agents’ capabilities at verifiable tasks.
Second, we need to measure how quickly current limitations of agents at conducting open-ended research (such as the lack of creativity and judgment) can be overcome, such as through more targeted training and scaffold improvements. We plan to continue shadow evaluations on a regular basis to help answer this question.
Finally, even if these limitations can be overcome, there may be further bottlenecks that dampen the pace of AI progress.
Bottlenecks could include compute limits, the necessity of collecting data from real-world experiments, and others that we haven’t recognized yet because they are not currently blocking progress. For example, the importance of high-quality RL environments was not clear before they turned out to be useful for inference scaling. Similarly, the importance of building energy infrastructure for data centers was not realized before companies started investing hundreds of billions on data centers for training and inference.
Whether we encounter further bottlenecks, and how tractable they turn out to be, will be consequential for understanding the pace of progress. In this vein, our paper identifies an unresolved bottleneck, namely, the poor performance of frontier agents on open-ended AI research (though it remains to be seen if it is on the critical path to RSI).
If we’re in a world where the bottlenecks to fully automated research can be easily resolved, we should expect dramatic returns to AI progress from improving AI capabilities. But if we’re in the world with many remaining bottlenecks that are hard to overcome, Amdahl’s law :https://en.wikipedia.org/wiki/Amdahl%27s_law would kick in: even a hundredfold speedup in the parts amenable to AI would only lead to a small speedup in the overall pace of progress, since progress is bottlenecked by the pace of the slowest component. 3:#footnote-3
Figuring out which world we live in could dramatically impact estimates of the pace of AI progress. We hope our results contribute to a richer understanding of these bottlenecks.
Read the paper here :https://arxiv.org/pdf/2607.27191 . The authors are Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, and Arvind Narayanan.
This is as opposed to sending papers for blind peer review. Unfortunately, peer review in AI suffers from poor review quality and a mismatch between reviewers’ expertise and the papers they are assigned, perhaps resulting from a dramatic :https://www.cs.cmu.edu/~nihars/preprints/SurveyPeerReview.pdf increase in submissions.
For example, our coauthors disagree on whether the agents lack creativity, or if they suffered from epistemic lock-in and were unable to productively incorporate feedback.
In practice, it might turn out that some kinds of AI progress have bottlenecks while others don’t. For example, improving the efficiency and speed of existing AI systems is a verifiable task, where progress has been rapid. As a result, companies have productively :https://x.com/OpenAI/status/2082878156483219672?s=20 utilized agents for improving AI systems’ efficiency. In the near term, we expect quick AI progress in dimensions that have verifiable signals.
Interesting that the agents get "stuck in a loop" and fail to reconsider previously set-aside pathways. "That didn't work before, but maybe this time it will. Maybe if I just tweak something..." seems to be a fundamentally human course of action (for all that it so often fails). And giving up before exhausting the budget? That ain't right! :-D
Just more evidence that whatever process of "thinking" the machines are using, it isn't *human*.
Does research proceed by asking research questions? My (Deweyan) view is that formulating your research as a question distorts it, unless you have already done the research.