{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T06:40:50.084Z","headline":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","description":"Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","url":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/","mainEntityOfPage":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/","datePublished":"2026-07-14T00:58:47.000Z","dateModified":"2026-07-14T00:58:47.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared","https://aihot.virxact.com/items/cmrjy5jxd01lfbiw28a8ggp6j"],"canonicalUrl":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。 Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/","dateCreated":"2026-07-14T00:58:47.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared","datePublished":"2026-07-14T00:58:47.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrjy5jxd01lfbiw28a8ggp6j","datePublished":"2026-07-14T00:58:47.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrjy5jxd01lfbiw28a8ggp6j"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared"},"article":{"id":"cmrjy5jxd01lfbiw28a8ggp6j","slug":"cmrjy5jxd01lfbiw28a8ggp6j","url":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/","title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","title_en":"Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8： Agentic Coding Benchmarks， API Pricing， and Cost-Performance Tradeoffs Compared","summary":"Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared","aiHotUrl":"https://aihot.virxact.com/items/cmrjy5jxd01lfbiw28a8ggp6j","publishedAt":"2026-07-14T00:58:47.000Z","category":"模型更新","score":58,"selected":false,"articleBody":["Anthropic just shipped Claude Sonnet 5：https://www.anthropic.com/news/claude-sonnet-5 . They call it its most agentic Sonnet model yet. It plans, drives browsers and terminals, and runs autonomously across long tasks.","Sonnet 5 is the default model for Free and Pro plans today. Max, Team, and Enterprise users can select it. It is also live in Claude Code and on the Claude Platform.","Sonnet sits in the middle of Anthropic’s lineup. It is above the cheaper Haiku 4.5 and below the flagship Opus 4.8.","Sonnet 5 is an upgrade to Sonnet 4.6, which launched in February 2026. Anthropic frames this release around agentic reliability, not one headline benchmark.","In practice, that means longer task chains without losing context. It means better self-correction when a tool call fails. It means steadier behavior across extended sessions inside Claude Code or Cowork.","The model exposes effort levels: low, medium, high, and xhigh (extra high). Higher effort spends more tokens on reasoning. That raises both quality and cost.","It is important to note that Sonnet 5 uses an updated tokenizer, the same one introduced with Opus 4.7. The same text can map to roughly 1.0 to 1.35 times more tokens.","Estimate per-task cost across models and compare published benchmarks. All figures from Anthropic’s June 30, 2026 launch.","Anthropic team published a benchmark table comparing Sonnet 5, Sonnet 4.6, and Opus 4.8. Sonnet 5 beats its predecessor in every tested category. It closes much of the gap to Opus 4.8.","On agentic coding (SWE-bench Pro), Sonnet 5 scores 63.2%. Sonnet 4.6 scored 58.1%. Opus 4.8 still leads at 69.2%.","On computer use (OSWorld-Verified), Sonnet 5 posts 81.2% against Sonnet 4.6’s 78.5%. On Terminal-Bench 2.1, it reaches 80.4% versus 67.0%.","On Humanity’s Last Exam with tools, Sonnet 5 hits 57.4%. That nearly matches Opus 4.8 at 57.9%.","There is one place where Sonnet 5 edges ahead. On the GDPval-AA v2 knowledge-work benchmark, it scores 1,618 against Opus 4.8’s 1,615.","The cost-performance story is the most important part for developers. Sonnet 5 is a strict improvement over Sonnet 4.6 across every effort level. The clearest value appears at low and medium effort.","At those levels, Sonnet 5 delivers quality that earlier Sonnet pricing could not buy. Opus 4.8 remains the accuracy leader at the top of the range.","A practical routing policy follows from this. Send most agentic coding, tool use, and knowledge work to Sonnet 5. Reserve Opus 4.8 for accuracy-critical tasks. Keep Haiku 4.5 for high-volume, latency-sensitive calls.","Early access partners described concrete workflows. Their reports map to common engineering jobs.","Sonnet 5’s introductory pricing runs through August 31, 2026. Standard pricing of $3/$15 begins after that date. Standard prompt caching (cache reads at 0.1x input) and the 50% Batch API discount also apply. Per token, Sonnet 5 undercuts GPT-5.5 and Gemini 3.1 Pro, but costs more than Gemini 3.5 Flash. Anthropic lists a 1M-token context window for Sonnet 5 in its launch post. It does not publish context figures for the other models here.","The API call mirrors any other Anthropic model. You change the model string to claude-sonnet-5 .","Early developer reactions from Hacker News and X on launch day, June 30, 2026.","Mixed reception: praise for price-to-value, doubts about standing at full $3/$15 pricing. Manually labeled from the public posts below; the two Reddit links are live threads, not counted here.","Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us ：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/06/Screenshot-2026-06-30-at-2.15.57-PM-1.png","alt":"","afterParagraph":12,"url":"/media/articles/cmrjy5jxd01lfbiw28a8ggp6j/5f1cfa4a75666643.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/06/Screenshot-2026-06-30-at-2.16.26-PM-1.png","alt":"","afterParagraph":12,"url":"/media/articles/cmrjy5jxd01lfbiw28a8ggp6j/2a6fdf5db686c93d.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":21,"url":"/media/articles/cmrjy5jxd01lfbiw28a8ggp6j/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/high-level-description-a-developer-focus_feCO3rqGV4ig6q7LaBcN2w_K5t5TwjZTPWy666HxX0epA-100x70.png","alt":"Patter SDK Guide to Building a Restaurant Booking Phone Agent with Dynamic Variables, Guardrails, Latency Dashboards, and Eval Checks","afterParagraph":22,"url":"/media/articles/cmrjy5jxd01lfbiw28a8ggp6j/c998049732b3b311.webp"}],"mediaStatus":"ok","articleBodyZh":["Anthropic 刚发布了 Claude Sonnet 5：https://www.anthropic.com/news/claude-sonnet-5。他们称其为迄今最具自主性的 Sonnet 模型。它能够规划、驱动浏览器和终端，并在长时间任务中自主运行。","Sonnet 5 目前是免费和专业计划的默认模型。Max、Team 和 Enterprise 用户可以选择它。它也已经在 Claude Code 和 Claude 平台上线。","Sonnet 位于 Anthropic 产品线的中端。它高于价格较低的 Haiku 4.5，但低于旗舰 Opus 4.8。","Sonnet 5 是 Sonnet 4.6 的升级版本，后者于 2026 年 2 月发布。Anthropic 将此次发布定位为围绕自主可靠性，而非单一的标杆性能。","实际应用中，这意味着可以进行更长的任务链而不会丢失上下文。这意味着当工具调用失败时有更好的自我纠正能力。这意味着在 Claude Code 或 Cowork 内的长时间会话中表现更稳定。","该模型显示了不同的努力等级：低、中、高和 x高（超高）。更高的努力等级在推理上消耗更多的 token，从而提高质量和成本。","需要注意的是，Sonnet 5 使用了更新后的分词器，即在 Opus 4.7 中引入的分词器。同样的文本大约映射为 1.0 到 1.35 倍的 token。","可估算不同模型每个任务的成本，并对比已发布的基准数据。所有数据均来自 Anthropic 2026 年 6 月 30 日发布。","Anthropic 团队发布了一张基准表，对比了 Sonnet 5、Sonnet 4.6 和 Opus 4.8。Sonnet 5 在每个测试类别中都超过了其前身，并缩小了与 Opus 4.8 的差距。","在自主编码（SWE-bench Pro）测试中，Sonnet 5 得分为 63.2%。Sonnet 4.6 得分为 58.1%。Opus 4.8 仍以 69.2% 领先。","在计算机使用（OSWorld-Verified）测试中，Sonnet 5 得分 81.2%，而 Sonnet 4.6 为 78.5%。在 Terminal-Bench 2.1 中，它达到 80.4%，而 Sonnet 4.6 为 67.0%。","在 Humanity’s Last Exam 使用工具测试中，Sonnet 5 得分 57.4%，几乎与 Opus 4.8 的 57.9% 相当。","有一个方面 Sonnet 5 略胜一筹。在 GDPval-AA v2 知识工作基准测试中，它得 1,618 分，而 Opus 4.8 得 1,615 分。","对于开发者而言，性价比故事是最重要的部分。Sonnet 5 在每个努力等级上都严格优于 Sonnet 4.6。在低、中努力等级下，其价值最为明显。","在这些层级上，Sonnet 5 提供了以往 Sonnet 定价无法购买的质量。Opus 4.8 仍然是该范围内准确性领先的产品。","由此可以制定一个实用的路由策略。将大部分代理编码、工具使用和知识工作发送给 Sonnet 5。将 Opus 4.8 保留用于对准确性要求极高的任务。将 Haiku 4.5 用于高量、延迟敏感的调用。","早期访问合作伙伴描述了具体的工作流程。他们的报告与常见的工程工作相对应。","Sonnet 5 的入门定价有效期至 2026 年 8 月 31 日。此日期之后，标准定价为 $3/$15。标准提示缓存（以 0.1 倍输入读取缓存）以及 50% 的批量 API 折扣同样适用。按每个令牌计算，Sonnet 5 的价格低于 GPT-5.5 和 Gemini 3.1 Pro，但高于 Gemini 3.5 Flash。Anthropic 在 Sonnet 5 的发布文章中列出了 100 万令牌的上下文窗口，但没有公布其他模型的上下文数据。","API 调用与任何其他 Anthropic 模型相同。只需将模型字符串更改为 claude-sonnet-5。","早期开发者在 2026 年 6 月 30 日上线当天，在 Hacker News 和 X 上的反馈。","反应不一：对性价比表示赞赏，但对 $3/$15 的完整定价存在疑虑。以下根据公开帖子手动标注；两个 Reddit 链接是活跃线程，这里不计入。","需要与我们合作推广您的 GitHub 仓库、Hugging Face 页面、产品发布或网络研讨会等吗？请联系： https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位有远见的企业家和工程师，Asif 致力于利用人工智能的潜力造福社会。他最近的工作是推出人工智能媒体平台 Marktechpost，该平台以对机器学习和深度学习新闻的深入报道而脱颖而出，内容技术扎实，同时易于广大受众理解。该平台月访问量超过 200 万次，显示出其在受众中的受欢迎程度。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。 Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T06:49:19.296Z","sourceHash":"57bd3163c91384f2","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["模型更新","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AI资讯","description":"Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%","url":"https://www.aioga.com/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"en":{"title":"Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8： Agentic Coding Benchmarks， API Pricing， and Cost-Performance Tradeoffs Compared","summary":"Aioga tracks this update from MarkTechPost（RSS） under Models. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"Models","source":"MarkTechPost（RSS）","pageTitle":"Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8： Agentic Coding Benchmarks， API Pricing， and Cost-Performance Tradeoffs Compared - Aioga AI News","description":"Aioga tracks this update from MarkTechPost（RSS） under Models. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-","url":"https://www.aioga.com/en/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"ja":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aiogaは「モデル更新」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"モデル更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AIニュース","description":"Aiogaは「モデル更新」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro ","url":"https://www.aioga.com/ja/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"ko":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga는 MarkTechPost（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"모델 업데이트","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AI 뉴스","description":"Aioga는 MarkTechPost（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro ","url":"https://www.aioga.com/ko/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"es":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Modelos. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"Modelos","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Modelos. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet","url":"https://www.aioga.com/es/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"fr":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Modèles. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"Modèles","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Modèles. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 ","url":"https://www.aioga.com/fr/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"de":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga KI-News","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/de/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"pt-BR":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Notícias de IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/pt-BR/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"ru":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Новости ИИ","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/ru/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"ar":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/ar/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"hi":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AI समाचार","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/hi/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"it":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Notizie IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/it/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"nl":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AI-nieuws","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/nl/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"tr":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga AI Haberleri","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/tr/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"vi":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Tin tức AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/vi/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"id":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Berita AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/id/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"th":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - ข่าว AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/th/news/cmrjy5jxd01lfbiw28a8ggp6j/"},"pl":{"title":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-bench Pro 达 63.2%（Sonnet 4.6 为 58.1%），OSWorld-Verified 达 81.2%（78.5%），HLE 达 57.4%（46.8%），知识工作基准 GDPval-AA v2 得分 1，618，略超 Opus 4.8 的 1，615。API 定价方面，Sonnet 5 输入/输出价格为 $2/$10 每百万 token（8 月 31 日前促销价），之后恢复 $3/$15；Opus 4.8 为 $5/$25。模型支持低、中、高、极高四种推理努力级别，上下文窗口为 100 万 token。Sonnet 5 即日起作为 Free 和 Pro 计划的默认模型，并已在 Claude Code 和 Claude Platform 上线。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Anthropic 发布 Claude Sonnet 5：智能体能力与基准测试对比 - Aioga Wiadomości AI","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Anthropic 发布 Claude Sonnet 5，定位为最具智能体能力的 Sonnet 模型，支持规划、驱动浏览器和终端，并可在长任务中自主运行。Sonnet 5 在每项已发布基准测试中均超越前代 Sonnet 4.6：SWE-be","url":"https://www.aioga.com/pl/news/cmrjy5jxd01lfbiw28a8ggp6j/"}}}}