散点图揭示了每任务成本与性能的关系:推理系统增加推理时间可带来渐近性能提升,而基础 LLM(如 GPT-4.5、Claude 3.7)的单次推理点代表无扩展推理的原始性能。
ARC-AGI 已从其最初的版本(ARC-AGI-1 和 2)发展而来,这些版本测量的是被动流体智力,而 ARC-AGI-3 则挑战 AI 代理在新颖的互动环境中即时适应的能力。
上图散点图可视化了每任务成本与性能之间的关键关系——这是衡量效率的重要指标。真正的智能不仅在于解决问题,还在于以最少的资源高效地解决问题。
更多信息,请参见我们的测试政策:/policy。
仅显示运行成本低于 10,000 美元的系统。
对于无法完成全部测试输出的模型,其剩余任务将标记为错误。
标记为“预览”的结果为非官方结果,可能基于不完整的测试。
1 ARC-AGI-2 分数估计基于部分测试结果和 o1-pro 定价。
2 基于 Gemini 3 Pro 定价的临时成本估算。模型发布后将重新测试。
立即开始并接收官方竞赛更新和新闻。
ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid intelligence, to ARC-AGI-3 which challenges AI agents to adapt on the fly to novel interactive environments.
The scatter plot above visualizes the critical relationship between cost-per-task and performance - a key measure of efficiency. True intelligence isn't just about solving problems, but solving them efficiently with minimal resources.
For more information, see our testing policy:/policy.
Only systems which required less than $10,000 to run are shown.
For models that were not able to produce full test out puts, remaining tasks were marked as incorrect.
Results marked as "preview" are unofficial and may be based on incomplete testing.
1 ARC-AGI-2 score estimate based on partial testing results and o1-pro pricing.
2 Provisional cost estimates based on Gemini 3 Pro pricing. Model to be retested once released.
Get started and receive official contest updates and news.