Mike 让它将博客从 Webflow 迁移出去。它一次提示就完成了。接着他要求它构建斯坦福的 AI 城镇模拟版本:https://arxiv.org/abs/2304.03442,该模拟跟随 AI 角色的日常生活,观察他们如何建立关系和传播信息。一次性提示中,它创建了25个带有记忆和日常作息的角色,这些角色交流并传播关于镇上派对和市长选举的消息。“在这一点上,我甚至不确定模型有什么不能做的,”Mike 说道。
丹(Dan)处理了六月份 Fable 拒绝的与安全相关的编码工作,当时 Fable 的防护措施将这些请求发送给了 Opus 4.8。丹指出,以安全风险为由的拒绝“明显减少”——模型在判断什么构成安全风险时更多地信任人的判断。凯兰(Kieran)在常规工程工作中也出现了更少的误报,而 Fable 5 有时会将常规编码任务误判为安全工作并阻止它们。凯兰唯一遇到的拒绝情况是:在自动测试循环中无法使用测试密码登录,他称其为“有点过头”。
这是过去一年中我们团队的作者希望使用的首个 Claude 模型。它的文字顺序清晰,并且能够接受修改而不争论。但在长工作中仍需要事实校对,它尚未取代我们自动化编辑流程中的 Opus 5:/context-window/benchmarks-don-t-know-your-job,至于它的风格是否胜过 GPT-5.6 Sol,则取决于个人喜好。
在 Every 的写作测试台上,该平台安排模型完成 TK 编辑任务,新模型表现最好的任务是长篇任务,例如从零开始写介绍或填补论文中缺失的部分。在缺失段落任务中,该任务要求模型从丹的文章《自动化之后》中定义 AI slop:/p/after-automation,它的得分超过我们运行过的所有模型,包括 Fable 5。它的段落更多地使用了周围文章的细节,使论证清晰,每段表达一个观点,并且忠实于原始材料,而 Fable 5 和 Opus 5 往往会进行推测。Fable 5.1 的输出同样比其他模型的 AI 特征更少,阅读水平为七年级,比 Fable 5 低一个年级,比 Sol 低两个年级。
X 帖子是唯一一个落后的工作,几乎排在最后,仅高于 Sonnet 5,而 Sol 在每个 Anthropic 模型中都领先。基准测试一直沿着这条线分化:Anthropic 模型喜欢扩展论点,GPT 模型则希望压缩它。新模型是我们测试过的最强扩展器,而压缩能力低于平均水平。它的 X 帖子在简短要求一条论点时却堆叠了第二和第三个想法。
当你提高努力水平时,模型在判断如何删减内容上表现更好。High 和 xHigh 得分最高。在段落填充任务中,中等和高之间的差距最大,在这个任务中,模型必须阅读周围的论点并决定插入内容以使论点完整,而高设置下写得更短:同一篇介绍在 Medium 下为 805 字,在 xHigh 下为 657 字。与以往模型一样,我选择 High 作为起草的最佳选项。
作为编辑,它速度快但谨慎。Jannik 让它编辑了 20 篇过去 Every 草稿,方式和 Kate Lee(Every 的主编)编辑它们的方式一致。它找到了 Kate 的精确修改次数与 Opus 5 相差不大,但用时不到一半,却丢掉了大约一半找到的修改;Opus 5 保留了大部分自己的修改,Sol 几乎保留全部。它能看见编辑,但没有足够的自信展示给你,这与它起草时的问题正好相反。Jannik 的结论是 Opus 5 仍然适合该流程;Dan 的结论是 Fable 5.1 略差但快得多。
Marcus 在 Every Agent(https://agent.every.to/,我们的基于 Slack 的 AI 助手)中测试了两种模型,使用内部检查评估它们的写作能力以及使用工具完成任务的能力。两者均设置为中等努力,指令最初为 Opus 5 调整。Marcus 认为这两种模型的输出可比,但 Fable 5.1 使用的 token 不到一半,响应速度更快。在 Jannik 的编辑测试中,它完成的时间不到 Opus 5 的一半。
Marcus 的注意事项:因为它行动迅速且使用的 token 很少,所以“在复杂多工具任务中可能表现较差”。如果你的流程需要模型小心地调用八个工具,Opus 5 仍然是更安全的选择。
任务是重型多工具工作或者调优的编辑流程。Marcus 和 Jannik 都发现 Opus 5 在这些方面更稳定。
没人设定预算。在 xHigh 模式下,如果放任不管,可能要花整整一天时间。
这是否是一个 4.5 级时刻——那种发布能让一整批新用户意识到 AI 潜力的时刻——还是仅仅是一个非常好的 Fable,要取决于 Anthropic 目前尚未公开的包装和定价决策。无论如何,团队再次使用了 Claude。
披露:Anthropic 在 Fable 5.1 发布前向 Every 提供了预先访问权限。该公司没有参与本次评测。
现在每个人都是建设者。Every All Access:https://every.to/builder-pack 为你提供完整会员资格加 Builder Pack——即我们构建工具所需的 $9,000+ 代币。
Fast, friendly, and even more powerful than its predecessor
We called the first Fable:/vibe-check/anthropic-mythos-our-fable-vibe-check a “warp drive” because it was the most powerful coding model we’d ever tested. But it was slow and hard to understand—it felt like a power tool only for the most AI-pilled of users.
In our testing it’s more powerful than Fable 5, running for days at a time to tackle coding challenges that stumped its predecessor. But it’s also fast and friendly, and, importantly, it speaks English like a normal person, er, AI. Unlike the original Fable and Anthropic’s recent Sonnet 5:/vibe-check/sonnet-5 and Opus 5:/vibe-check/opus-5 models, Anthropic has returned to form and created a model we actually enjoy talking to.
Oh, and there’s more: In our testing, it uses about half the tokens of Opus 5 for similar tasks. It should make Fable-level work available within an Opus budget. So it’s priced for everyone too.
Even more interesting, “everyone” now includes enterprises. Fable 5.1 comes with a zero-data retention option, removing a barrier that precluded most large companies from using it.
Fable 5.1 is a rare combination: the strongest coding model we’ve used and a collaborator we can recommend to people who don’t write code. It should be the first glimpse for most knowledge workers of truly delegated AI work—previously only the domain of programmers. You can give it a task like building a slide deck and move on to other things, confident that it’ll come back great in an hour or two.
We’ve been testing it for about a week, across coding, writing, and knowledge work tasks. Here’s our day-zero Vibe Check.
Anthropic says Fable 5.1 makes its strongest capabilities easier and cheaper to use. Its main claims:
Anthropic says the model can pursue a goal over extended runs and make better decisions about how to solve the underlying problem. Kieran’s experience supports that: Compared with Fable 5 on a task to build a clone of Every’s document editor Proof from scratch, it went deeper, added useful details he hadn’t requested, and showed better judgment about what worked.
Anthropic promises less stock phrasing, shorter updates, and closer attention to writing instructions. We saw this too: Its writing had fewer AI tells, and the team found its explanations easier to follow. It still routinely exceeded requested word counts.
Anthropic says Low and Medium can deliver performance at least as good as Fable 5 at substantially lower cost. High is the default in Claude Code; Medium is the default in Claude.ai and Cowork. Kieran liked the coding results at all three settings, though our writing tests favored High and Extra-high.
Anthropic estimates costs about 25 percent below Fable 5 for typical API workloads and Claude Code extra usage, with savings approaching 45 percent for extensive autonomous work. Our tests offer evidence of token efficiency: In Every’s Slack assistant, it used less than half as many tokens as Opus 5 with comparable results. We haven’t measured the cost savings against Fable 5.
Anthropic says it is less likely than Mythos 5 to disregard constraints, rationalize its decisions, or game an evaluation. We didn’t run those safety tests, but ordinary instruction-following still had gaps: At Extra-high, Kieran sometimes found it kept working when he interrupted to ask what it was doing.
Standard API rates remain $10 per million input tokens and $50 per million output tokens, matching Fable 5. Reading previously cached input now costs $0.25 per million tokens, a 75 percent reduction. Anthropic estimates overall savings of about 25 percent for typical API workloads and Claude Code extra usage, rising to roughly 45 percent for extensive autonomous work.
At launch, Fable 5.1 will be available in Claude.ai, Claude Code, and Cowork, through Anthropic’s API, and through Amazon Web Services, Google Cloud, and Microsoft Azure.
It handles the same work Fable did, from one-prompt app builds to coding runs that last all day, and it’s easier to work with while it does it. It answers faster, explains what it’s doing in plain language, and changes course when you tell it to instead of arguing.
Dan’s first message to the team about it: “I can actually understand what it’s saying.” The outputs I measured read at a lower grade level than Opus 5 or GPT-5.6 Sol, meaning the text is accessible to more people, with fewer AI tells.
Asked for 1,000 words, it wrote 1,288. Asked for three to six themes, it gave eight. Asked for eight to 12 quotes, it pulled 43, and five of 27 quotes weren’t in the source. At the highest effort setting, it spun up subagents it didn’t need and kept going when Dan asked it to stop and explain.
AI-pilled writer by day, vibe coder by night
For coding, Fable 5.1 does what Fable 5 did, and it’s easier to work with while it does it. Dan moved all of his coding tasks over from Codex. Mike sends all of his new projects to it.
Kieran asked it to rebuild Proof:https://www.proofeditor.ai/ from a single prompt—the same test Fable 5 completed in June. Fable 5.1 also built a working version in one shot, but Kieran found it went deeper, adding useful details he hadn’t specified. He also saw more style and better judgment about what worked. As he put it, “What a time to be alive to get this in one shot.”
It also handled his LFG workflow, where the model plans, writes, reviews, and tests without a person in the loop. By the end of the week, he’d updated the compound engineering plugin:https://github.com/everyinc/compound-engineering-plugin with instructions tailored to it. “It’s the right balance of having an opinion and getting shit done versus pushing back,” he said.
Dan still uses ChatGPT far more for everyday work, but he’s delegating gigantic coding jobs to Fable 5.1 and letting them run. After he got access, Claude Code was making roughly five times as many requests to the model each day.
Mike asked it to move a blog off Webflow. It did it in one prompt. He then asked it to build a version of Stanford’s AI-town simulation:https://arxiv.org/abs/2304.03442, which follows AI characters through their daily routines to see how they form relationships and spread information. In one shot, it created 25 characters with memories and daily routines, who chatted and spread news about a party and a mayoral race happening in the town. “At this point, I’m not even sure what the model can’t do,” Mike said.
Dan ran security-related coding work that Fable had refused in June, when Fable’s safeguards sent those requests to Opus 4.8. Dan noted that refusals on the basis of security risks are “way down”—the model trusts the human’s judgment more on what does or does not constitute a security risk. Kieran also hit fewer false positives during normal engineering work, whereas Fable 5 sometimes mistook routine coding tasks for security work and blocked them. The one refusal Kieran hit: It wouldn’t log in with a test password during an automated testing loop, which he called “a bit extra.”
As for effort levels, Kieran ran the same kinds of tasks at every effort setting and found that Low, Medium, and High were all “super good.” Extra-high “rips” on long runs but hands off too much work to subagents. He’d rather not pick at all: “I just wish they don’t have effort levels and it will do this automatically.” Most of the team settled on High for work with which they stay in the loop.
This is the first Claude model in a year that the writers on our team want to draft with. It writes clear prose in the right order, and it takes an edit without arguing. Still, it needs a fact-checker on long jobs, it hasn’t displaced Opus 5 in our automated editing pipeline:/context-window/benchmarks-don-t-know-your-job, and whether its style beats GPT-5.6 Sol’s is a matter of taste.
On Every’s writing bench, which tasks models with TK editorial tasks, the new model’s best work was the long-form ones, such as writing an introduction from scratch or filling in a missing section of an essay. On the missing passage assignment, which asks the model to define AI slop from Dan’s essay “After Automation,”:/p/after-automation it outscored every model we’ve run, Fable 5 included. Its passages used more of the particulars in the surrounding essay, built the argument so that each paragraph made one point, and stayed faithful to the source material where Fable 5 and Opus 5 tended to infer. Fable 5.1 outputs also carried fewer AI tells than any other model’s, and they read at a seventh-grade level, a grade below Fable 5 and two below Sol.
The X post is the one job where it trailed, landing near the bottom, above only Sonnet 5 while Sol led every Anthropic model. The bench has always split along this line: Anthropic models like to extend an argument, GPT models want to compress it. The new model is the strongest extender we’ve tested and a below-average compressor. Its X posts stack a second and third idea where the brief asked for one.
When you kick the effort level up, the model exercises better judgment about what to leave out. High and xHigh scored best. The gap between medium and high was widest on the paragraph-filling assignment, where the model has to read the surrounding argument and decide what belongs in the gap to make the argument complete, and higher settings wrote shorter: The same introduction ran 805 words at Medium and 657 at xHigh. As with previous models, I settled on High as the best option for drafting.
The same instinct that makes it a good drafter makes it a risky reporter. Given a tight brief and set of sources, it’s faithful. But if you give it long source documents and room to elaborate, it can get creative in the wrong ways: Mike asked it to produce a Vibe Check using eight to 12 exact quotes from the Every team. It produced 43 quotes, including several that were missing from the supplied text.
As an editor it’s fast and timid. Jannik had it edit 20 past Every drafts the way Kate Lee:/@kate_1767, Every’s editor in chief, had edited them. It found Kate’s exact edits about as often as Opus 5 in less than half the time, but threw away about half of what it had found; Opus 5 kept most of its own and Sol nearly all. It can see the edit; it just doesn’t trust itself enough to show you, which is the reverse of its drafting problem. Jannik’s call is that Opus 5 stays the model for that pipeline; Dan’s is that the Fable 5.1 is a little worse and a lot faster.
Give the new model a big consulting assignment and it finishes the whole thing, and on judgment-heavy tasks like slide decks and persona interviews, it beats GPT-5.6 Sol. Give it a brief with a word count, a theme count, or a quote count, and it goes over.
Mike’s EC (executive consulting) Bench covers 12 consulting tasks: dashboards, a training-needs synthesis, a workshop run sheet, a branded slide deck, a roleplayed interview, and others. Sol scored 100 on six of the structured building tasks, then 44 on the blog post and 33 on the roleplayed interview. The new model did the reverse. It scored 88 on persona answers, 93 on the slide deck, and 53 on the roleplay, which Mike still called “basically a fail” but which beat Sol. It lost points where it ignored explicit limits.
Mike asked the model to design three hands-on projects for an executive learning to build with AI, complete with prompts and realistic practice data. One was a dashboard tracking four company acquisitions. The model supplied spreadsheets, status emails, and Slack conversations, including warnings about a project delay that hadn’t yet reached the tracker. In just under 23 minutes, it produced a training pack Mike called “close to perfect.”
Limits are where it loses points. A hotel recommendation with a 1,000-word cap came back at 1,288 words. A training-needs synthesis that called for three to six themes came back with eight. A brief that allowed eight to 12 supporting quotes got 43. A request for markdown got HTML. Opus 5, on the same tasks, stayed inside every limit and timed out twice.
Design is the other weak spot. Mike asked it to build an app that lets people fill out forms by talking to an AI interviewer. The interview worked, but the app had a generic purple-on-black design, large stretches of empty space, and a stray “false” displayed beneath the controls. Fable produced the same purple screen in June. Opus 5 didn’t.
In automated pipelines like Slack assistants designed to execute longform tasks that require tool use, Fable 5.1 does the same work as Opus 5 on about half the tokens and in about 60 percent of the time. Left alone at the highest effort setting, it keeps working, keeps spawning subagents, and doesn’t always stop when you ask. The moral of the story: Set a budget before you start.
Marcus tested both models inside the Every Agent:https://agent.every.to/, our Slack-based AI assistant, using internal checks of its writing and ability to complete tasks with tools. Both ran at Medium effort, with instructions originally tuned for Opus 5. Marcus judged the models’ outputs to be comparable, but Fable 5.1 used less than half as many tokens and responded faster. In Jannik’s editing test, it finished in less than half the time of Opus 5.
Marcus’s caveat: Because it acts fast and spends few tokens, it “can perform worse on heavy multi-tool tasks.” If your pipeline needs a model to chain eight tool calls carefully, Opus 5 is still the safer pick.
Kieran ran long tasks at xHigh and found it “extremely thorough, maybe too much, especially delegating to subagents,” and said that “sometimes [it] does ignore me too.” His xHigh sessions “go for like a day at a time.” When he interrupted one to ask, “Can you explain what you’re doing?” he said it “just ignored me and continued to use a billion subagents.” Mike’s exercise pack ran more than twice his normal time. Kieran put 1.8 billion tokens through it in a single day. It is cheap per step and will take as many steps as you let it, and the effort setting is what decides how many that is.
Fable 5.1 is Fable with less waiting and better manners. It’s faster, writes more clearly, uses fewer tokens per step, argues less, and handles long jobs as well as Fable 5 did. Mike moved his new work to it. Dan moved his coding to it. Katie moved her writing to it, two months after deciding Fable wasn’t for her.
Dan and Kieran give it a gold star for the turnaround and the token efficiency. Mike is holding their gold for a price that lets more people afford Fable-grade work, and Anthropic hadn’t shared a price when we filed this piece. Kieran’s caveat on the comeback story: “Fable never sucked.” Anthropic’s bad months were about Fable being unavailable and two other models missing the mark, while GPT-5.6 came out.
You work with the model in the loop. Writing, editing alongside it, and coding where you read as it goes. High effort is where most of the team landed.
You want Fable-grade results on long jobs without Fable’s wait. Delegated builds, big refactors, and whole-project rebuilds finished for Kieran and Mike at the same level as Fable, faster.
You have Opus 5 agent prompts and a token bill. Marcus got the same pass rate on task completion in the Every Agent at Medium on half the tokens, with a better writing score, without rewriting a prompt.
The brief has hard limits. Word counts, theme counts, quote counts, and output formats all got exceeded in our testing. Opus 5 stayed inside every one.
Quotes have to be right the first time. Five of 27 checkable quotes on one task weren’t in the source. Verify every quotation and every added detail before it leaves your hands.
The task is heavy multi-tool work or a tuned editing pipeline. Marcus and Jannik both found Opus 5 steadier there.
Nobody has set a budget. At xHigh it will take the whole day if you let it.
Whether this is a 4.5 moment—the kind of release that helps a whole new group of people realize what AI makes possible—or just a very good Fable depends on packaging and pricing decisions Anthropic hadn’t shared as of this writing. Either way, the team is using Claude again.
Disclosure: Anthropic provided Every with pre-launch access to Fable 5.1. The company had no input on this review.
Everyone’s a builder now. Every All Access:https://every.to/builder-pack gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.
情报判断
Aioga 编辑摘要
Every 在约一周的测试后发布 Fable 5.1 体验评测,称其在编程、写作和知识工作任务中比 Fable 5 更强,同时更快、更易交流;文章还称,相似任务下其令牌用量约为 Opus 5 的一半。
背景分析
Every 此前将初代 Fable 称为其测试过最强的编程模型,但认为该模型速度慢且难以理解。本次文章称 Fable 5.1 提供零数据保留选项,并表示 Anthropic 的主要主张是让模型能力更易用、成本更低。
Aioga 观点
Aioga 判断:这篇文章的核心价值在于提供早期实际使用感受,并将能力、交互体验、令牌用量和数据保留并列考察;但其结论来自 Every 的短期测试,适用范围仍需更多材料验证。