Opus 5 设计为日常使用:它的效率比其他模型更高。它是 Claude Max 的新默认模型,也是 Claude Pro 上最强的模型。
Claude Opus 5 在与其前身 Opus 4.8 相同的成本下提供了大幅提升的性能。本节中的图表显示了根据模型的努力设置性能如何变化,客户可以使用这些设置优化智能或节省代币以获得更快、更便宜的结果。
Opus 5 在有价值的软件工程任务中表现出色。例如,在 Frontier-Bench v0.1 上,Opus 5 超过了所有其他模型,同时在每个任务成本更低的情况下,其性能是 Opus 4.8 的两倍以上。在 CursorBench 3.2 上,在最大努力模式下,该模型的表现与 Fable 5 的峰值分数相差不到 0.5%,但每个任务的成本只有一半;此外,在高、极高和最大努力下,它在给定成本下的表现优于所有其他模型。
我们在知识工作和问题解决任务中也看到类似的结果。例如:
它在几个相关评估中也是我们最优且最具成本效益的模型:
Opus 5 在科学研究方面相比 Opus 4.8 有显著提升。在我们所有的生命科学评估中,Opus 5 的表现都优于 Opus 4.8,这些评估涵盖了结构生物学、有机化学和生物信息学等主题。其改进在有机化学任务上尤为明显,例如从光谱数据推断分子结构(在我们内部基准测试中,它比 Opus 4.8 高出 10.2 个百分点),以及蛋白质相关任务,如预测蛋白质序列变化如何影响功能(在此项中,它高出 7.7 个百分点)。
最后,Opus 5 能够生成更强大的视觉输出:
Claude Opus 5 在验证其工作和仔细迭代直到成功方面要强得多。在评估和早期访问测试中,我们和用户发现了许多 Opus 5 展现出的主动性和彻底性的例子:
以下是我们早期访问客户在使用 Opus 5 时的进一步报告:
一致性。在部署前测试中,我们的自动行为审计发现 Opus 5 是迄今为止我们最符合标准的模型(如下图所示)。它遵循 Claude 宪法:https://www.anthropic.com/constitution,比 Opus 4.8、Sonnet 5 或 Fable 5 更好;表现出最低的欺骗行为率;并且最不容易被误导用于不当用途。在避免可能产生难以逆转副作用的鲁莽行为方面,它同样是我们迄今为止最安全的模型。
与其前身 Opus 4.8 一样,我们有意避免将 Opus 5 训练用于网络任务。然而,由于模型变得更通用,它在这些任务上的表现仍有显著提升,并且在寻找网络安全漏洞方面接近 Mythos 5。然而,在漏洞利用方面——即将漏洞转化为实质性网络威胁——它仍明显落后于 Mythos 5。
这一点可从 Opus 5 在 OSS-Fuzz 上的表现看出,这是我们开发的一项评估,用于评估模型在无需大量人工指导的情况下发现并利用漏洞的能力。尽管 Mythos 5 和 Opus 5 在识别漏洞上成功率相似,但 Opus 5 在开发漏洞利用方案方面的得分远远落后于 Mythos 5。
Claude Opus 5 的安全保护措施旨在允许模型在网络安全和生物学中进行有益的使用。它们与我们应用于 Opus 4.8 的措施相似,只是在某些有限的网络任务上设定了更强的防护栏。
网络安全。Opus 5 的网络分类器相较于 Fable 5 更加宽松。它们允许 Opus 5 查找源代码中的漏洞,但会阻止“基于二进制”的漏洞扫描(一种更可能与恶意行为者相关的方法)、渗透测试和漏洞利用生成。
根据我们的测试,我们预计这些分类器介入的频率比在 Fable 5 上低约 85%。在 Claude.ai:http://claude.ai/redirect/website.v1.66aba3b9-1d35-4b13-8282-dd904dd7310f、Claude Code 和 Claude Cowork 中,任何被标记的请求将默认回退到 Opus 4.8。API 上也可以启用回退到 Opus 4.8 的功能。
我们的网络验证计划(CVP):https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet,使那些原本会被模型安全措施阻碍的网络安全工作得以进行。已经加入 CVP 的企业和研究人员可以立即访问安全限制更少的 Opus 5 版本。
生物学。由于 Opus 5 拥有与 Opus 4.8 类似的一系列安全防护,因此它现在是我们在科学研究方面最强大的通用可用模型。不过,该模型在长时间运行的自主研究任务中仍显示出重要的局限性,这些任务正是我们预计 AI 模型在生物学相关风险中可能产生最大影响的领域。(Mythos 5 在这类生物学工作中仍然是更强的模型。)作为此次发布的一部分,在 Fable 5 上被阻止的生物学相关请求现在将路由到 Opus 5,而不是 Opus 4.8。
Claude Opus 5 今天已在所有平台上可用,价格为每百万输入令牌 5 美元,每百万输出令牌 25 美元(与 Opus 4.8 相同)。开发者可以通过 Claude API 开始使用 claude-opus-5。
它还提供快速模式,其运行速度约为默认速度的 2.5 倍。与 Opus 4.8 一样,快速模式在 Claude 平台和通过 Claude Code 的使用信用中,价格是 Opus 5 基础价格的两倍。
除了 Opus 5,我们还发布了两个处于测试阶段的更新:
与之前的 Opus 模型一致,Opus 5 在一般访问中没有数据保留要求。
有关如何最大限度地利用 Opus 5 的更多指导,请参阅我们的提示指南:https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5。
我们正在为 Claude 推出 Anthropic 经济指数连接器,它让任何人都可以探索关于 AI 和工作的真实数据。
Anthropic 正在向 Public First Action 额外捐赠 2000 万美元,使我们的总支持资金达到 4000 万美元。
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench:https://www.frontierbench.ai/ and GDPval-AA:https://artificialanalysis.ai/evaluations/gdpval-aa, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2 , at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.
We see similar results on knowledge work and problem-solving tasks. For example:
It’s also our best and most cost-efficient model on several related evaluations:
Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).
Finally, Opus 5 is capable of producing much stronger visual outputs:
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:
Below are further reports from our early-access customers on their experience of working with Opus 5:
Alignment. During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution:https://www.anthropic.com/constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.
Safety. Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card:https://www.anthropic.com/claude-opus-5-system-card.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.
Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai:http://claude.ai/redirect/website.v1.66aba3b9-1d35-4b13-8282-dd904dd7310f, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.
Our Cyber Verification Program:https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.
Biology. Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.
Alongside Opus 5, we’re releasing two updates in beta:
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
For more guidance on how to get the best out of Opus 5, see our prompting guide:https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5.
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
We’re sharing the research agenda for the Anthropic Economic Futures Research Fund.
We're launching the Anthropic Economic Index connector for Claude, which lets anyone explore real data about AI and work.
Anthropic is contributing an additional $20 million to Public First Action, bringing our total support to $40 million.
情报判断
Aioga 编辑摘要
Anthropic 发布 Claude Opus 5,称其智能水平接近 Claude Fable 5,而价格为后者一半。该模型已成为 Claude Max 默认模型,并被定位为 Claude Pro 当前最强模型。
背景分析
官方材料将 Opus 5 与前代 Opus 4.8、Fable 5 等模型进行比较,并展示其在编程、知识工作和生命科学评测中的表现;用户还可通过 effort 设置权衡智能水平、速度与令牌消耗。
对 Claude Max 和 Claude Pro 用户而言,默认模型与最强模型的调整可能直接改变日常使用选择。对开发者而言,按 effort 设置优化成本与性能,可能成为采用 Opus 5 时的重要考量。 后续值得关注 Opus 5 在 Frontier-Bench、GDPval-AA、CursorBench 及生命科学任务中的外部复核结果,同时观察其在网络安全任务上落后 Mythos 5 的差距是否缩小。