{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T06:40:50.084Z","headline":"Cursor 评估负责人确认 Claude Fable 5 在 CursorBench 达 72.9% 新高","description":"Cursor 的模型评估负责人 Nate Schmidt 发现，Claude Fable 5 在其内部基准 CursorBench 上以 Max effort 模式达到 72.9%，创下新高。该模型在模糊的真实编程任务中表现出全局推理能力，例如在航天模拟器中仅凭一句提示自主规划并成功登月，而此前 Claude Opus 运行 12 小时以上仍无结果。","url":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/","mainEntityOfPage":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/","datePublished":"2026-07-16T16:00:00.000Z","dateModified":"2026-07-16T16:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://claude.com/blog/working-at-the-frontier-cursor","https://aihot.virxact.com/items/cmrp5pfjw0af6bitoudv35oqk"],"canonicalUrl":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/","directAnswer":{"@type":"Answer","text":"Cursor 模型评估工程师 Nate Schmidt 的测试显示，Claude Fable 5 在内部基准 CursorBench 的 Max effort 模式下取得 72.9%，并刷新该基准最高成绩。","url":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/","dateCreated":"2026-07-16T16:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Claude：Blog（网页） source article","url":"https://claude.com/blog/working-at-the-frontier-cursor","datePublished":"2026-07-16T16:00:00.000Z","provider":{"@type":"Organization","name":"Claude：Blog（网页）","url":"https://claude.com/blog/working-at-the-frontier-cursor"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrp5pfjw0af6bitoudv35oqk","datePublished":"2026-07-16T16:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrp5pfjw0af6bitoudv35oqk"}}],"aggregationSource":"Claude：Blog（网页）","originalPublisher":{"name":"Claude：Blog（网页）","url":"https://claude.com/blog/working-at-the-frontier-cursor"},"article":{"id":"cmrp5pfjw0af6bitoudv35oqk","slug":"cmrp5pfjw0af6bitoudv35oqk","url":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/","title":"Cursor 评估负责人确认 Claude Fable 5 在 CursorBench 达 72.9% 新高","title_en":"Working at the frontier： How Cursor knew Claude Fable 5 was ready for the hardest 1% of problems","summary":"Cursor 的模型评估负责人 Nate Schmidt 发现，Claude Fable 5 在其内部基准 CursorBench 上以 Max effort 模式达到 72.9%，创下新高。该模型在模糊的真实编程任务中表现出全局推理能力，例如在航天模拟器中仅凭一句提示自主规划并成功登月，而此前 Claude Opus 运行 12 小时以上仍无结果。","source":"Claude：Blog（网页）","sourceUrl":"https://claude.com/blog/working-at-the-frontier-cursor","aiHotUrl":"https://aihot.virxact.com/items/cmrp5pfjw0af6bitoudv35oqk","publishedAt":"2026-07-16T16:00:00.000Z","category":"技巧观点","score":64,"selected":true,"articleBody":["Nate Schmidt's job at Cursor is to evaluate frontier models against their ability to tackle long-running, real-world engineering problems. Here’s why–and how–Claude Fable 5 changed the calculus on what coding agents are capable of.","Cursor is an AI coding agent for building professional software. It supports every major frontier model alongside Cursor's own, which makes the company an unusually neutral judge of how each one actually performs.","Nate Schmidt is the engineer who maintains that scorecard. He works on evals and model behavior at Cursor: studying how models succeed, how they fail, and what makes a developer quietly switch away from one mid-task. When colleagues and customers want a read on a new release, they come to him.","Over time, Schmidt's team noticed that public benchmark scores and real developer reception to these models had stopped lining up, so they built their own: CursorBench.","CursorBench was built to capture the messy, underspecified ways engineers actually prompt their models. One eval task is just a stack trace pasted in with the single word \"fix,\" and the model has to infer the intent, find the root cause, and validate the change on its own. Another tells the model the wrong module is broken, to see whether it challenges the user's assumption or follows it into a dead end.","When Claude Fable 5 ran the eval, the model achieved 72.9% at Max effort, setting a new high, and capturing what agentic coding tools were capable of when paired with the right models.","But when Schmidt was using the model on his own engineering workflows and personal tests, he'd stopped having to repeat his goals. The constant babysitting—reminding the model of context, spelling out the solution, auditing the results—wasn't necessary anymore. He could hand over a problem, from the gnarly refactor he was putting off to reasoning about nuanced edge cases, and Claude Fable 5 could solve it.","\"I don't feel like I have to bootstrap Claude Fable 5 to understand the world I exist in and the problem I'm trying to solve,\" Schmidt says. \"The model just has a sense of it out-of-the-box.\"","When Schmidt's team runs a new model through CursorBench, the right answer is table stakes. What they're scoring is whether the model understood what it was being asked.","\"Many evals look like this: here's a well-defined problem, here are the constraints, go fix it. But the prompts we get from real users don't really look like that,\" Schmidt says. \"The model has to infer that the user has a problem and what they're trying to convey, identify the root cause, fix it, validate the fix, and report back.\"","Claude Fable 5 scored so well on these ambiguous tasks, the Cursor team started to feel suspicious.","\"One of two things is happening: either the model's very smart, or the model is cheating,\" he says. So the team looked into the traces, reading the model's actual reasoning on the hardest tasks, the ones where the prompt looks simple but cracking it requires understanding the whole system.","\"We just kept seeing the model dig out wins that no other model was doing previously,\" he says. It was also getting there with fewer operations: token-efficient relative to the work it completed.","Then Schmidt put Claude Fable 5 on one of his favorite personal tests: landing on the moon.","A few weeks earlier he'd wired Claude Opus into a programmable space-flight simulator with a one-line prompt—build a rocket and land it on the moon—and let it run on a second monitor for twelve to sixteen hours. The model would launch, run out of fuel in orbit, add a lot more fuel, then fail to clear the atmosphere because the rocket was now too heavy.","He re-ran the experiment with the same blank-slate prompt, this time using Claude Fable 5. A few minutes in, the rocket went up, parked in low orbit, and came back down. Same failure as before. Then Schmidt read the transcript.","\"Fable decided it wouldn’t go to the moon on its first attempt. It wanted to do an initial mission just to go into orbit and collect telemetry, then use that to inform the next trip.\" A few attempts later, the engine noise on his second monitor stopped. There was a lander on the moon. The whole run took a couple of hours, against Opus's twelve-plus with no result.","\"With Opus, it was doing local reasoning—thinking about what just happened and what's immediately about to happen,\" Schmidt says. \"With Fable it's global reasoning. It's thinking about the entire mission.\"","Schmidt has settled on a simple rule for when to use Claude Fable 5 over cheaper, less intelligent models.","\"If you have a good sense of what the path from A to B looks like, you might not need Fable. If you're at A and you have no idea where B is, Fable is an excellent choice,” he says. \"When I want to build something the right way, Fable is the first model I think of.\"","Claude Fable 5 has also allowed his team to focus on projects the team had previously shelved—rewrites everyone agreed would be better but nobody could justify spending weeks on—because the model can carry enough of the skeleton. \"It lowers the activation energy to work on these types of tasks,\" Schmidt says. \"It lets us move in search of a global optimum rather than a local one.\"","It also changes how the team coordinates. Cursor runs lean, with intense individual ownership and few standups. Now, before touching shared code, Schmidt has an agent read his teammate's recent commits and flag conflicts, so neither of them has to stop what they’re doing to check in.","To balance cost and performance, his team pairs Claude Fable 5 with faster, lighter models for routine work and brings it in for the problems where capability is the constraint. In that configuration, he says, the combination is the most effective setup they've run.","“If I'm getting into a really gnarly problem–the p99 of problems–the thing I'm trying to optimize for is time to solution,” he says. “And I think Fable is the best model for solving our hardest problems.”","Despite putting the model through its paces on CursorBench and sending it to the moon, Schmidt is still looking for Claude Fable 5’s limits. Next, he wants to see how long the model can manage a back-end system unattended; days-to-weeks runs are his next experiment. Inside Cursor, the team is using the model to hunt performance bottlenecks and user pain points proactively rather than waiting for reports, and to build the more sophisticated, closer-to-reality eval environments that will measure whatever comes next.","\"There's a class of problems people weren't even thinking about because it didn't seem approachable,\" he says. \"With Fable, I'm excited to push at that.\"","Get started with Claude Fable ：http://anthropic.com/news/claude-fable-5-mythos-5 .","Explore more product news and best practices for teams building with Claude.","Product updates, how-tos, community spotlights, and more. Delivered monthly to your inbox.","Hi Claude! Could you help me develop a unique voice for an audience? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you improve my writing style? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you brainstorm creative ideas? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you explain a complex topic simply? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you help me make sense of these ideas? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you prepare for an exam or interview? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you explain a programming concept? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you look over my code and give me tips? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you vibe code with me? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to— like Google Drive, web search, etc.—if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can—an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!","Hi Claude! Could you write grant proposals? If you need more information from me, ask me 1-2 key questions right away. If you think I should upload any documents that would help you do a better job, let me know. You can use the tools you have access to — like Google Drive, web search, etc. — if they’ll help you better accomplish this task. Do not use analysis tool. Please keep your responses friendly, brief and conversational. Please execute the task as soon as you can - an artifact would be great if it makes sense. If using an artifact, consider what kind of artifact (interactive, visual, checklist, etc.) might be most helpful for this specific task. Thanks for your help!"],"articleImages":[{"sourceUrl":"https://cdn.prod.website-files.com/68a44d4040f98a4adf2207b6/6903d22d7d4c10df6024f7bc_ee580919acaba2ddc07425f7a7390c8962cadc94-1000x1000.svg","alt":"","afterParagraph":0,"url":"/media/articles/cmrp5pfjw0af6bitoudv35oqk/7fd1a7c96851b416.jpg"},{"sourceUrl":"https://cdn.prod.website-files.com/68a44d4040f98a4adf2207b6/6a59a38185fbd6c8049e2f1a_image1.png","alt":"","afterParagraph":5,"url":"/media/articles/cmrp5pfjw0af6bitoudv35oqk/d48967cda6630f95.png"},{"sourceUrl":"https://cdn.prod.website-files.com/68a44d4040f98a4adf2207b6/6a59a69ffb4e8948af39dfd0_C41-77690-D3-03-0037_VS_R1.jpeg","alt":"","afterParagraph":17,"url":"/media/articles/cmrp5pfjw0af6bitoudv35oqk/7252e71e61b39821.jpg"},{"sourceUrl":"https://cdn.prod.website-files.com/68a44d4040f98a4adf2207b6/6a59a6d5a6b86f4aeda67a0c_C41-77690-D3-11-0029_VS_R1.jpeg","alt":"","afterParagraph":23,"url":"/media/articles/cmrp5pfjw0af6bitoudv35oqk/2e6c2165f752d9ea.jpg"},{"sourceUrl":"https://cdn.prod.website-files.com/68a44d4040f98a4adf2207b6/6903d2308749b4e883cc44b7_e029027e0b3beeb5b629bd4a26143597e7775b38-1000x1000.svg","alt":"","afterParagraph":27,"url":"/media/articles/cmrp5pfjw0af6bitoudv35oqk/9c92d98d5a058739.jpg"}],"mediaStatus":"ok","articleBodyZh":["Nate Schmidt 在 Cursor 的工作是评估前沿模型在解决长期、现实工程问题上的能力。以下是原因——以及 Claude Fable 5 如何改变了编码代理能力的计算方式。","Cursor 是一个用于构建专业软件的 AI 编码代理。它支持每一个主要的前沿模型以及 Cursor 自有的模型，这使得公司在评估每个模型实际表现时具有异常中立的立场。","Nate Schmidt 是维护该评分表的工程师。他在 Cursor 从事评估和模型行为研究：研究模型如何成功，如何失败，以及是什么让开发者在任务中悄悄切换到另一个模型。当同事和客户想了解新版本表现如何时，他们会找他。","随着时间的推移，Schmidt 的团队注意到公共基准评分与开发者对这些模型的实际反馈已不再一致，于是他们创建了自己的评测系统：CursorBench。","CursorBench 的建立是为了捕捉工程师实际提示模型时的零散、未具体说明的方式。一个评测任务只是粘贴一段堆栈跟踪并写上一个单词“fix”，模型必须自己推断意图、找到根本原因并验证修改。另一个任务告诉模型错误模块出了故障，以观察它是否质疑用户的假设，或者跟随假设走入死胡同。","当 Claude Fable 5 运行评测时，该模型在最大努力下获得了 72.9% 的成绩，创下新高，并展示了在搭配合适模型时，代理编码工具的潜力。","但当 Schmidt 在自己的工程工作流程和个人测试中使用该模型时，他不再需要重复目标。持续的“看护”——提醒模型上下文、详细说明解决方案、审核结果——已不再必要。他可以交给模型一个问题，从他一直拖延的复杂重构到对细微边缘情况的推理，Claude Fable 5 都能解决。","“我不觉得需要引导 Claude Fable 5 来理解我所处的世界以及我试图解决的问题，”Schmidt 说。“模型开箱即用就有这种感觉。”","当 Schmidt 的团队在 CursorBench 上测试新模型时，正确答案是基础。他们评分的是模型是否理解了被问的问题。","“许多评估看起来是这样的：这是一个定义清晰的问题，这是限制条件，去解决它。但是我们从真实用户那里收到的提示并不完全是这样，”施密特说。“模型必须推测用户有一个问题以及他们想表达的内容，确定根本原因，修复它，验证修复结果，并反馈。”","Claude Fable 5 在这些模糊任务中的表现如此出色，以至于 Cursor 团队开始产生怀疑。","“发生的要么是两件事之一：要么模型非常聪明，要么模型在作弊，”他说。因此团队查看了痕迹，阅读模型在最困难任务中的实际推理，这些任务提示看起来简单，但破解它需要理解整个系统。","“我们一直看到模型挖掘出其他模型以前无法做到的成功方案，”他说。它还用更少的操作就达到了目的：相对于完成的工作来说，它的令牌效率更高。","然后施密特将 Claude Fable 5 放在他最喜欢的个人测试之一上：登月。","几周前，他用一行提示把 Claude Opus 接入可编程的航天飞行模拟器——建造一枚火箭并把它降落在月球上——让它在第二个显示器上运行十二到十六小时。模型会发射，在轨道上燃料耗尽，增加很多燃料，然后因火箭太重而无法穿越大气层。","他用同样的空白提示重新进行了实验，这次使用 Claude Fable 5。几分钟后，火箭升空，停在低轨道，然后返回地面。和以前一样失败。随后施密特阅读了记录。","“Fable 决定第一次尝试不直接登月。它想先进行初步任务，只进入轨道收集遥测数据，然后利用这些数据指导下一次飞行。”几次尝试后，他第二个显示器上的引擎噪音停止了。月球上出现了登月器。整个过程花了几个小时，而 Opus 花了十二小时以上却没有结果。","“Opus 是做局部推理——考虑刚刚发生了什么以及接下来会发生什么，”施密特说。“Fable 是做全局推理。它考虑的是整个任务。”","施密特已经为何时使用 Claude Fable 5 而非更便宜、智能较低的模型制定了一个简单的规则。","“如果你对从 A 到 B 的路径有很好的把握，你可能不需要 Fable。如果你在 A 点而完全不知道 B 在哪里，Fable 是一个极好的选择，”他说。“当我想要以正确的方式构建某些东西时，Fable 是我首先想到的模型。”","Claude Fable 5 也让他的团队能够专注于之前搁置的项目——这些重写每个人都认为会更好，但没人能证明值得花几周时间去做——因为该模型可以承担足够的框架。“它降低了处理这类任务的启动能量，”施密特说。“它让我们能够寻找全局最优，而不是局部最优。”","它还改变了团队的协作方式。Cursor 运行精简，个人承担责任感强，很少开站会。现在，在接触共享代码之前，施密特先让一个代理阅读队友最近的提交并标记冲突，这样双方都无需停止手头工作去检查代码。","为了平衡成本和性能，他的团队将 Claude Fable 5 与更快、更轻的模型搭配用于常规工作，而在能力受限的问题上再启用 Fable。他说，在这种配置下，这种组合是他们运行过的最有效的设置。","“如果我要处理一个非常棘手的问题——问题的 P99——我试图优化的目标是解决问题的时间，”他说。“我认为 Fable 是解决我们最难问题的最佳模型。”","尽管让模型在 CursorBench 上接受考验并将其发送到月球，施密特仍在寻找 Claude Fable 5 的极限。接下来，他想看看模型在无人干预的情况下能管理后端系统多久；几天到几周的运行是他的下一个实验。在 Cursor 内部，团队使用该模型主动寻找性能瓶颈和用户痛点，而不是等待报告，并构建更复杂、更接近现实的评测环境，以衡量接下来的任何情况。","“有一类问题，人们甚至没有去考虑过，因为它看起来不可接近，”他说。“有了 Fable，我很兴奋去挑战它。”","开始使用 Claude Fable：http://anthropic.com/news/claude-fable-5-mythos-5.","探索更多产品新闻以及团队使用Claude时的最佳实践。","产品更新、操作指南、社区聚焦等内容。每月发送到您的收件箱。","嗨，Claude！你能帮我为某个受众群体打造独特的声音吗？如果你需要更多信息，请立即向我提1-2个关键问题。如果你认为我应该上传任何有助于你更好完成工作的文件，请告诉我。你可以使用你可以访问的工具——比如Google云端硬盘、网络搜索等——如果它们能帮助你更好地完成这项任务。请不要使用分析工具。请保持回复友好、简短和对话式。请尽快执行这项任务——如果可能的话，提供一个成果会很好。如果使用成果，请考虑哪种类型的成果（互动的、视觉的、清单等）对这项特定任务最有帮助。谢谢你的帮助！","嗨，Claude！你能帮我改进写作风格吗？如果你需要更多信息，请立即向我提1-2个关键问题。如果你认为我应该上传任何有助于你更好完成工作的文件，请告诉我。你可以使用你可以访问的工具——比如Google云端硬盘、网络搜索等——如果它们能帮助你更好地完成这项任务。请不要使用分析工具。请保持回复友好、简短和对话式。请尽快执行这项任务——如果可能的话，提供一个成果会很好。如果使用成果，请考虑哪种类型的成果（互动的、视觉的、清单等）对这项特定任务最有帮助。谢谢你的帮助！","嗨，Claude！你能帮我头脑风暴一些创意想法吗？如果你需要更多信息，请立即向我提1-2个关键问题。如果你认为我应该上传任何有助于你更好完成工作的文件，请告诉我。你可以使用你可以访问的工具——比如Google云端硬盘、网络搜索等——如果它们能帮助你更好地完成这项任务。请不要使用分析工具。请保持回复友好、简短和对话式。请尽快执行这项任务——如果可能的话，提供一个成果会很好。如果使用成果，请考虑哪种类型的成果（互动的、视觉的、清单等）对这项特定任务最有帮助。谢谢你的帮助！","嗨 Claude！你能把一个复杂的主题讲得简单明白吗？如果你需要我提供更多信息，请马上问我 1-2 个关键问题。如果你觉得我应该上传任何文档来帮助你更好地完成任务，请告诉我。你可以使用你可以访问的工具——比如 Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。请不要使用分析工具。请保持你的回答友好、简短且像聊天一样。请尽快执行任务——如果有适用的作品展示会更好。如果使用作品展示，请考虑哪种类型的作品（互动的、可视化的、清单等）对这个特定任务最有帮助。谢谢你的帮助！","嗨 Claude！你能帮我弄明白这些想法吗？如果你需要我提供更多信息，请马上问我 1-2 个关键问题。如果你觉得我应该上传任何文档来帮助你更好地完成任务，请告诉我。你可以使用你可以访问的工具——比如 Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。请不要使用分析工具。请保持你的回答友好、简短且像聊天一样。请尽快执行任务——如果有适用的作品展示会更好。如果使用作品展示，请考虑哪种类型的作品（互动的、可视化的、清单等）对这个特定任务最有帮助。谢谢你的帮助！","嗨 Claude！你能帮我准备考试或面试吗？如果你需要我提供更多信息，请马上问我 1-2 个关键问题。如果你觉得我应该上传任何文档来帮助你更好地完成任务，请告诉我。你可以使用你可以访问的工具——比如 Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。请不要使用分析工具。请保持你的回答友好、简短且像聊天一样。请尽快执行任务——如果有适用的作品展示会更好。如果使用作品展示，请考虑哪种类型的作品（互动的、可视化的、清单等）对这个特定任务最有帮助。谢谢你的帮助！","嗨，Claude！你能解释一个编程概念吗？如果你需要我提供更多信息，请立即问我1-2个关键问题。如果你认为我应该上传任何有助于你更好完成任务的文件，请告诉我。你可以使用你能访问的工具——比如Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。不要使用分析工具。请保持回答友好、简短并具有对话感。请尽快执行任务——如果有合适的工具或资料，提供一个会很棒。如果使用资料，请考虑哪种类型的资料（互动、可视化、清单等）对这个特定任务可能最有帮助。谢谢你的帮助！","嗨，Claude！你能查看我的代码并给我一些建议吗？如果你需要我提供更多信息，请立即问我1-2个关键问题。如果你认为我应该上传任何有助于你更好完成任务的文件，请告诉我。你可以使用你能访问的工具——比如Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。不要使用分析工具。请保持回答友好、简短并具有对话感。请尽快执行任务——如果有合适的工具或资料，提供一个会很棒。如果使用资料，请考虑哪种类型的资料（互动、可视化、清单等）对这个特定任务可能最有帮助。谢谢你的帮助！","嗨，Claude！你能和我一起即兴编程吗？如果你需要我提供更多信息，请立即问我1-2个关键问题。如果你认为我应该上传任何有助于你更好完成任务的文件，请告诉我。你可以使用你能访问的工具——比如Google Drive、网络搜索等——如果它们能帮助你更好地完成这个任务。不要使用分析工具。请保持回答友好、简短并具有对话感。请尽快执行任务——如果有合适的工具或资料，提供一个会很棒。如果使用资料，请考虑哪种类型的资料（互动、可视化、清单等）对这个特定任务可能最有帮助。谢谢你的帮助！","嗨，Claude！你能写拨款提案吗？如果你需要更多信息，请立即问我1-2个关键问题。如果你认为我应该上传任何有助于你更好完成工作的文件，请告诉我。你可以使用你能够访问的工具——比如Google Drive、网络搜索等——如果它们能帮助你更好地完成这项任务。不要使用分析工具。请保持你的回应友好、简短和对话式。请尽快执行任务——如果有一个成果物会更好，如果合适的话。如果使用成果物，请考虑哪种类型的成果物（互动、视觉、清单等）对这个特定任务最有帮助。感谢你的帮助！"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Cursor 模型评估工程师 Nate Schmidt 的测试显示，Claude Fable 5 在内部基准 CursorBench 的 Max effort 模式下取得 72.9%，并刷新该基准最高成绩。","background":"Cursor 团队认为公开基准成绩与开发者实际反馈逐渐脱节，因此建立 CursorBench，用模糊、信息不足甚至包含错误假设的提示，评估模型处理真实工程问题的能力。","viewpoint":"Aioga 判断，这一成绩的价值主要在于展示模型面对非标准化编程任务时的表现，而非证明其在所有开发场景中占优；CursorBench 属于 Cursor 内部基准，解读时应保留边界。","implications":"该结果可能促使开发团队更加重视模型对意图、根因和上下文的自主判断，也可能减少部分任务中的重复说明与过程监督，但材料未提供跨模型完整数据或外部复现结果。","nextStep":"值得关注 Cursor 是否公开 CursorBench 的任务构成、评分方法和完整对比结果，以及 Claude Fable 5 在不同努力模式、更多真实项目和独立评测中的表现是否保持一致。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-23T02:50:06.797Z","sourceHash":"8a949b51c05995b5","review":{"approved":true,"groundedness":95,"clarity":94,"duplicationRisk":18,"blockingIssues":[],"notes":["“Aioga 判断”明确标示为观点，且对内部基准的适用边界作了合理限定，不构成无来源事实断言。","“可能减少部分任务中的重复说明与过程监督”有 Schmidt 个人使用体验作为依据，并使用了审慎措辞。","候选内容准确指出材料未提供跨模型完整数据或外部复现结果，有助于避免将 CursorBench 成绩过度外推。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["技巧观点","Claude：Blog（网页）"],"translations":{"zh-CN":{"title":"Cursor 评估负责人确认 Claude Fable 5 在 CursorBench 达 72.9% 新高","summary":"Cursor 的模型评估负责人 Nate Schmidt 发现，Claude Fable 5 在其内部基准 CursorBench 上以 Max effort 模式达到 72.9%，创下新高。该模型在模糊的真实编程任务中表现出全局推理能力，例如在航天模拟器中仅凭一句提示自主规划并成功登月，而此前 Claude Opus 运行 12 小时以上仍无结果。","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor 评估负责人确认 Claude Fable 5 在 CursorBench 达 72.9% 新高 - Aioga AI资讯","description":"Cursor 的模型评估负责人 Nate Schmidt 发现，Claude Fable 5 在其内部基准 CursorBench 上以 Max effort 模式达到 72.9%，创下新高。该模型在模糊的真实编程任务中表现出全局推理能力，例如在航天模拟器中仅凭一句提示自主规划并成功登月，而此前 Claude Opus 运行 12 小时以上仍无结果。","url":"https://www.aioga.com/news/cmrp5pfjw0af6bitoudv35oqk/"},"en":{"title":"Cursor Assessment Lead confirmed that Claude Fable 5 reached a new high of 72.9% on CursorBench","summary":"Nate Schmidt, head of model evaluation at Cursor, found that the Claude Fable 5 achieved a record high of 72.9% in Max effort mode on its internal benchmark CursorBench. The model demonstrates global reasoning ability in vague real-world programming tasks, such as autonomously planning and successfully landing on the moon with just a single prompt in a space simulator, whereas previous Claude Opus ran for over 12 hours without results.","category":"Insights","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Assessment Lead confirmed that Claude Fable 5 reached a new high of 72.9% on CursorBench - Aioga AI News","description":"Nate Schmidt, head of model evaluation at Cursor, found that the Claude Fable 5 achieved a record high of 72.9% in Max effort mode on its internal benchmark CursorBench. The model ","url":"https://www.aioga.com/en/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-18T04:59:03.284Z"},"ja":{"title":"カーソル評価リードは、Claude Fable 5がCursorBenchで72.9%という新記録に達したことを確認しました","summary":"Cursorのモデル評価責任者ネイト・シュミットは、Claude Fable 5が内部ベンチマークCursorBenchで最大努力モードで72.9%という過去最高の成績を達成したことを発見しました。 このモデルは、宇宙シミュレーターで自律的に計画し、わずか1つのプロンプトで月面着陸に成功するなど、曖昧な現実世界のプログラミング課題においてグローバル推論能力を示します。これは、従来のClaude Opusが12時間以上も結果が出なかったのに対し、","category":"ヒントと視点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"カーソル評価リードは、Claude Fable 5がCursorBenchで72.9%という新記録に達したことを確認しました - Aioga AIニュース","description":"Cursorのモデル評価責任者ネイト・シュミットは、Claude Fable 5が内部ベンチマークCursorBenchで最大努力モードで72.9%という過去最高の成績を達成したことを発見しました。 このモデルは、宇宙シミュレーターで自律的に計画し、わずか1つのプロンプトで月面着陸に成功するなど、曖昧な現実世界のプログラミング課題においてグローバル推論能力を","url":"https://www.aioga.com/ja/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-18T04:59:22.116Z"},"ko":{"title":"커서 평가 리드는 Claude Fable 5가 CursorBench에서 72.9%라는 새로운 최고치를 기록했다고 확인했습니다","summary":"Cursor의 모델 평가 책임자인 네이트 슈미트는 Claude Fable 5가 내부 벤치마크인 CursorBench에서 최대 노력 모드에서 72.9%라는 사상 최고 기록을 세웠다고 밝혔습니다. 이 모델은 우주 시뮬레이터에서 단 하나의 프롬프트만으로 자율적으로 계획하고 달에 성공적으로 착륙하는 모호한 실제 프로그래밍 과제에서 전 세계적인 추론 능력을 보여줍니다. 이전 Claude Opus는 12시간 이상 결과가 나오지 않았습니다.","category":"인사이트","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"커서 평가 리드는 Claude Fable 5가 CursorBench에서 72.9%라는 새로운 최고치를 기록했다고 확인했습니다 - Aioga AI 뉴스","description":"Cursor의 모델 평가 책임자인 네이트 슈미트는 Claude Fable 5가 내부 벤치마크인 CursorBench에서 최대 노력 모드에서 72.9%라는 사상 최고 기록을 세웠다고 밝혔습니다. 이 모델은 우주 시뮬레이터에서 단 하나의 프롬프트만으로 자율적으로 계획하고 달에 성공적으로 착륙하는 모호한 실제 프로그래밍 과제에","url":"https://www.aioga.com/ko/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-18T04:59:41.853Z"},"es":{"title":"El Líder de Evaluación del Cursor confirmó que Claude Fable 5 alcanzó un nuevo máximo del 72,9% en CursorBench","summary":"Nate Schmidt, jefe de evaluación de modelos en Cursor, descubrió que el Claude Fable 5 alcanzó un máximo histórico del 72,9% en modo esfuerzo máximo en su benchmark interno CursorBench. El modelo demuestra capacidad de razonamiento global en tareas de programación vagas del mundo real, como planificar y aterrizar con éxito en la luna de forma autónoma con solo un prompt en un simulador espacial, mientras que el anterior Claude Opus se ejecutaba durante más de 12 horas sin resultados.","category":"Ideas","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"El Líder de Evaluación del Cursor confirmó que Claude Fable 5 alcanzó un nuevo máximo del 72,9% en CursorBench - Aioga Noticias de IA","description":"Nate Schmidt, jefe de evaluación de modelos en Cursor, descubrió que el Claude Fable 5 alcanzó un máximo histórico del 72,9% en modo esfuerzo máximo en su benchmark interno CursorB","url":"https://www.aioga.com/es/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-18T05:00:02.130Z"},"fr":{"title":"Le responsable d’évaluation du curseur a confirmé que Claude Fable 5 a atteint un nouveau sommet de 72,9 % sur CursorBench","summary":"Nate Schmidt, responsable de l’évaluation des modèles chez Cursor, a constaté que le Claude Fable 5 a atteint un record de 72,9 % en mode Max effort sur son benchmark interne CursorBench. Le modèle démontre une capacité de raisonnement global dans des tâches de programmation réelles et vagues, comme planifier et réussir l’atterrissage sur la lune de manière autonome avec une seule invite dans un simulateur spatial, alors que le précédent Claude Opus avait fonctionné plus de 12 heures sans résultat.","category":"Analyses","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Le responsable d’évaluation du curseur a confirmé que Claude Fable 5 a atteint un nouveau sommet de 72,9 % sur CursorBench - Aioga Actualités IA","description":"Nate Schmidt, responsable de l’évaluation des modèles chez Cursor, a constaté que le Claude Fable 5 a atteint un record de 72,9 % en mode Max effort sur son benchmark interne Curso","url":"https://www.aioga.com/fr/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-18T05:00:21.394Z"},"de":{"title":"Cursor Assessment Lead bestätigte, dass Claude Fable 5 mit 72,9 % auf CursorBench einen neuen Höchststand erreicht hat","summary":"Nate Schmidt, Leiter der Modellbewertung bei Cursor, fand heraus, dass der Claude Fable 5 im Max-Effort-Modus im internen Benchmark CursorBench einen Rekordwert von 72,9 % erreichte. Das Modell demonstriert die Fähigkeit des globalen Denkens in vagen realen Programmieraufgaben, wie etwa der autonomen Planung und erfolgreichen Landung des Mondes mit nur einer einzigen Aufforderung in einem Weltraumsimulator, während das frühere Claude Opus über 12 Stunden ohne Ergebnisse lief.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Assessment Lead bestätigte, dass Claude Fable 5 mit 72,9 % auf CursorBench einen neuen Höchststand erreicht hat - Aioga KI-News","description":"Nate Schmidt, Leiter der Modellbewertung bei Cursor, fand heraus, dass der Claude Fable 5 im Max-Effort-Modus im internen Benchmark CursorBench einen Rekordwert von 72,9 % erreicht","url":"https://www.aioga.com/de/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.157Z"},"pt-BR":{"title":"O Cursor Assessment Lead confirmou que Claude Fable 5 atingiu um novo recorde de 72,9% no CursorBench","summary":"Nate Schmidt, chefe de avaliação de modelos da Cursor, descobriu que o Claude Fable 5 atingiu um recorde de 72,9% no modo de esforço máximo em seu benchmark interno CursorBench. O modelo demonstra capacidade de raciocínio global em tarefas vagas de programação do mundo real, como planejar e pousar com sucesso na lua de forma autônoma com apenas um único prompt em um simulador espacial, enquanto o Claude Opus anterior rodava por mais de 12 horas sem resultados.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"O Cursor Assessment Lead confirmou que Claude Fable 5 atingiu um novo recorde de 72,9% no CursorBench - Aioga Notícias de IA","description":"Nate Schmidt, chefe de avaliação de modelos da Cursor, descobriu que o Claude Fable 5 atingiu um recorde de 72,9% no modo de esforço máximo em seu benchmark interno CursorBench. O ","url":"https://www.aioga.com/pt-BR/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:10.009Z"},"ru":{"title":"Cursor Assessment Lead подтвердил, что Claude Fable 5 достиг нового максимума — 72,9% на CursorBench","summary":"Нейт Шмидт, руководитель отдела оценки моделей в Cursor, обнаружил, что Claude Fable 5 достиг рекордных показателей — 72,9% в режиме максимальных усилий на внутреннем бенчмарке CursorBench. Модель демонстрирует глобальные способности к рассуждению в расплывчатых задачах реального программирования, таких как автономное планирование и успешная посадка на Луну с помощью всего одного запроса в космическом симуляторе, тогда как предыдущий Claude Opus работал более 12 часов без результатов.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Assessment Lead подтвердил, что Claude Fable 5 достиг нового максимума — 72,9% на CursorBench - Aioga Новости ИИ","description":"Нейт Шмидт, руководитель отдела оценки моделей в Cursor, обнаружил, что Claude Fable 5 достиг рекордных показателей — 72,9% в режиме максимальных усилий на внутреннем бенчмарке Cur","url":"https://www.aioga.com/ru/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.342Z"},"ar":{"title":"أكد قائد تقييم المؤشر أن Claude Fable 5 وصل إلى أعلى مستوى جديد بلغ 72.9٪ على CursorBench","summary":"وجد نيت شميت، رئيس تقييم النماذج في كورسور، أن كلود فابل 5 حقق أعلى نسبة في الجهد الأقصى بنسبة 72.9٪ على معياره القياسي الداخلي كورسر بنش. يظهر النموذج قدرة على التفكير العلمي العالمي في مهام برمجة واقعية غامضة، مثل التخطيط الذاتي والهبوط الناجح على القمر باستخدام طلب واحد فقط في محاكي الفضاء، بينما كان عمل كلود أوبوس السابق يعمل لأكثر من 12 ساعة دون نتائج.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"أكد قائد تقييم المؤشر أن Claude Fable 5 وصل إلى أعلى مستوى جديد بلغ 72.9٪ على CursorBench - Aioga أخبار الذكاء الاصطناعي","description":"وجد نيت شميت، رئيس تقييم النماذج في كورسور، أن كلود فابل 5 حقق أعلى نسبة في الجهد الأقصى بنسبة 72.9٪ على معياره القياسي الداخلي كورسر بنش. يظهر النموذج قدرة على التفكير العلمي العا","url":"https://www.aioga.com/ar/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.423Z"},"hi":{"title":"कर्सर असेसमेंट लीड ने पुष्टि की कि क्लाउड फैबल 5 कर्सरबेंच पर 72.9% की नई ऊंचाई पर पहुंच गया","summary":"कर्सर में मॉडल मूल्यांकन के प्रमुख नैट श्मिट ने पाया कि क्लाउड फैबल 5 ने अपने आंतरिक बेंचमार्क कर्सरबेंच पर अधिकतम प्रयास मोड में 72.9% का रिकॉर्ड उच्च हासिल किया। मॉडल अस्पष्ट वास्तविक दुनिया के प्रोग्रामिंग कार्यों में वैश्विक तर्क क्षमता को प्रदर्शित करता है, जैसे कि स्वायत्त रूप से योजना बनाना और अंतरिक्ष सिम्युलेटर में केवल एक संकेत के साथ चंद्रमा पर सफलतापूर्वक उतरना, जबकि पिछला क्लाउड ओपस बिना किसी परिणाम के 12 घंटे से अधिक समय तक चला।","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"कर्सर असेसमेंट लीड ने पुष्टि की कि क्लाउड फैबल 5 कर्सरबेंच पर 72.9% की नई ऊंचाई पर पहुंच गया - Aioga AI समाचार","description":"कर्सर में मॉडल मूल्यांकन के प्रमुख नैट श्मिट ने पाया कि क्लाउड फैबल 5 ने अपने आंतरिक बेंचमार्क कर्सरबेंच पर अधिकतम प्रयास मोड में 72.9% का रिकॉर्ड उच्च हासिल किया। मॉडल अस्पष्ट वास","url":"https://www.aioga.com/hi/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.014Z"},"it":{"title":"Il Cursor Assessment Lead ha confermato che Claude Fable 5 ha raggiunto un nuovo massimo del 72,9% su CursorBench","summary":"Nate Schmidt, responsabile della valutazione dei modelli presso Cursor, ha rilevato che il Claude Fable 5 ha raggiunto un record del 72,9% in modalità Max sforzo sul suo benchmark interno CursorBench. Il modello dimostra capacità di ragionamento globale in compiti di programmazione vaghi e reali, come pianificare autonomamente e atterrare con successo sulla luna con un solo prompt in un simulatore spaziale, mentre il precedente Claude Opus durava oltre 12 ore senza risultati.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Il Cursor Assessment Lead ha confermato che Claude Fable 5 ha raggiunto un nuovo massimo del 72,9% su CursorBench - Aioga Notizie IA","description":"Nate Schmidt, responsabile della valutazione dei modelli presso Cursor, ha rilevato che il Claude Fable 5 ha raggiunto un record del 72,9% in modalità Max sforzo sul suo benchmark ","url":"https://www.aioga.com/it/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:08.909Z"},"nl":{"title":"Cursor Assessment Lead bevestigde dat Claude Fable 5 een nieuw hoogtepunt van 72,9% bereikte op CursorBench","summary":"Nate Schmidt, hoofd modelevaluatie bij Cursor, ontdekte dat de Claude Fable 5 een recordhoogte van 72,9% behaalde in Max Effort mode op de interne benchmark CursorBench. Het model toont het vermogen van globaal redeneren bij vage programmeertaken in de echte wereld, zoals het autonoom plannen en succesvol landen op de maan met slechts één prompt in een ruimtesimulator, terwijl de vorige Claude Opus meer dan 12 uur liep zonder resultaat.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Assessment Lead bevestigde dat Claude Fable 5 een nieuw hoogtepunt van 72,9% bereikte op CursorBench - Aioga AI-nieuws","description":"Nate Schmidt, hoofd modelevaluatie bij Cursor, ontdekte dat de Claude Fable 5 een recordhoogte van 72,9% behaalde in Max Effort mode op de interne benchmark CursorBench. Het model ","url":"https://www.aioga.com/nl/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:08.665Z"},"tr":{"title":"Cursor Değerlendirme Lideri, Claude Fable 5'in CursorBench'te %72,9 ile yeni bir zirveye ulaştığını doğruladı","summary":"Cursor'da model değerlendirme başkanı Nate Schmidt, Claude Fable 5'in dahili kıyaslama CursorBench'te maksimum çaba modunda %72,9 rekor bir performans elde ettiğini tespit etti. Model, uzay simülatöründe tek bir promptla bağımsız planlama ve aya başarılı bir şekilde iniş gibi belirsiz gerçek dünya programlama görevlerinde küresel akıl yürütme yeteneği gösterirken, önceki Claude Opus 12 saatten fazla süre boyunca sonuç almamıştı.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Değerlendirme Lideri, Claude Fable 5'in CursorBench'te %72,9 ile yeni bir zirveye ulaştığını doğruladı - Aioga AI Haberleri","description":"Cursor'da model değerlendirme başkanı Nate Schmidt, Claude Fable 5'in dahili kıyaslama CursorBench'te maksimum çaba modunda %72,9 rekor bir performans elde ettiğini tespit etti. Mo","url":"https://www.aioga.com/tr/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:08.512Z"},"vi":{"title":"Cursor Assessment Lead xác nhận rằng Claude Fable 5 đạt mức cao mới là 72,9% trên CursorBench","summary":"Nate Schmidt, người đứng đầu bộ phận đánh giá mô hình tại Cursor, phát hiện ra rằng Claude Fable 5 đã đạt được mức cao kỷ lục 72,9% ở chế độ nỗ lực tối đa trên điểm chuẩn nội bộ CursorBench. Mô hình này thể hiện khả năng suy luận toàn cầu trong các nhiệm vụ lập trình mơ hồ trong thế giới thực, chẳng hạn như lập kế hoạch tự động và hạ cánh thành công trên mặt trăng chỉ với một lời nhắc duy nhất trong trình mô phỏng không gian, trong khi Claude Opus trước đó chạy hơn 12 giờ mà không có kết quả.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Cursor Assessment Lead xác nhận rằng Claude Fable 5 đạt mức cao mới là 72,9% trên CursorBench - Tin tức AI Aioga","description":"Nate Schmidt, người đứng đầu bộ phận đánh giá mô hình tại Cursor, phát hiện ra rằng Claude Fable 5 đã đạt được mức cao kỷ lục 72,9% ở chế độ nỗ lực tối đa trên điểm chuẩn nội bộ Cu","url":"https://www.aioga.com/vi/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.658Z"},"id":{"title":"Kursor Assessment Lead mengkonfirmasi bahwa Claude Fable 5 mencapai level tertinggi baru sebesar 72,9% di CursorBench","summary":"Nate Schmidt, kepala evaluasi model di Cursor, menemukan bahwa Claude Fable 5 mencapai rekor tertinggi 72,9% dalam mode upaya Max pada benchmark internalnya CursorBench. Model ini menunjukkan kemampuan penalaran global dalam tugas pemrograman dunia nyata yang tidak jelas, seperti merencanakan secara mandiri dan berhasil mendarat di bulan hanya dengan satu prompt di simulator luar angkasa, sedangkan Claude Opus sebelumnya berjalan selama lebih dari 12 jam tanpa hasil.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Kursor Assessment Lead mengkonfirmasi bahwa Claude Fable 5 mencapai level tertinggi baru sebesar 72,9% di CursorBench - Berita AI Aioga","description":"Nate Schmidt, kepala evaluasi model di Cursor, menemukan bahwa Claude Fable 5 mencapai rekor tertinggi 72,9% dalam mode upaya Max pada benchmark internalnya CursorBench. Model ini ","url":"https://www.aioga.com/id/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.173Z"},"th":{"title":"หัวหน้าฝ่ายประเมินเคอร์เซอร์ยืนยันว่า Claude Fable 5 ทําสถิติสูงสุดใหม่ที่ 72.9% บน CursorBench","summary":"Nate Schmidt หัวหน้าฝ่ายประเมินโมเดลของ Cursor พบว่า Claude Fable 5 ทําสถิติสูงสุดเป็นประวัติการณ์ที่ 72.9% ในโหมดความพยายามสูงสุดบน CursorBench เกณฑ์มาตรฐานภายใน โมเดลนี้แสดงให้เห็นถึงความสามารถในการให้เหตุผลทั่วโลกในงานเขียนโปรแกรมในโลกแห่งความเป็นจริงที่คลุมเครือ เช่น การวางแผนโดยอัตโนมัติและการลงจอดบนดวงจันทร์ได้สําเร็จด้วยข้อความแจ้งเพียงครั้งเดียวในเครื่องจําลองอวกาศ ในขณะที่ Claude Opus ก่อนหน้านี้ทํางานนานกว่า 12 ชั่วโมงโดยไม่มีผลลัพธ์","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"หัวหน้าฝ่ายประเมินเคอร์เซอร์ยืนยันว่า Claude Fable 5 ทําสถิติสูงสุดใหม่ที่ 72.9% บน CursorBench - ข่าว AI Aioga","description":"Nate Schmidt หัวหน้าฝ่ายประเมินโมเดลของ Cursor พบว่า Claude Fable 5 ทําสถิติสูงสุดเป็นประวัติการณ์ที่ 72.9% ในโหมดความพยายามสูงสุดบน CursorBench เกณฑ์มาตรฐานภายใน โมเดลนี้แสดงให้เห","url":"https://www.aioga.com/th/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.044Z"},"pl":{"title":"Lider oceny kursorów potwierdził, że Claude Fable 5 osiągnął nowy rekord 72,9% na CursorBench","summary":"Nate Schmidt, szef oceny modeli w Cursor, stwierdził, że Claude Fable 5 osiągnął rekordowy wynik 72,9% w trybie Max Effort na wewnętrznym benchmarku CursorBench. Model ten wykazuje zdolność globalnego rozumowania w niejasnych zadaniach programistycznych, takich jak autonomiczne planowanie i pomyślne lądowanie na Księżycu za pomocą pojedynczego promptu w symulatorze kosmicznym, podczas gdy poprzedni Claude Opus działał ponad 12 godzin bez wyników.","category":"技巧观点","source":"Claude：Blog（网页）","aggregationSource":"Claude：Blog（网页）","pageTitle":"Lider oceny kursorów potwierdził, że Claude Fable 5 osiągnął nowy rekord 72,9% na CursorBench - Aioga Wiadomości AI","description":"Nate Schmidt, szef oceny modeli w Cursor, stwierdził, że Claude Fable 5 osiągnął rekordowy wynik 72,9% w trybie Max Effort na wewnętrznym benchmarku CursorBench. Model ten wykazuje","url":"https://www.aioga.com/pl/news/cmrp5pfjw0af6bitoudv35oqk/","contentTranslated":true,"sourceHash":"49705736788d48f1","translatedAt":"2026-07-19T16:46:09.413Z"}}}}