{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"Fireworks AI 分析：前沿不是单个模型，而是路由器，FireRouter 发布","description":"Fireworks AI 发布 FireRouter 并分析 DeepSWE v1.1 上 113 个任务、18 个模型的路由潜力。","url":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","mainEntityOfPage":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","datePublished":"2026-09-20T16:00:00.000Z","dateModified":"2026-09-20T16:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://fireworks.ai/blog/the-frontier-isnt-a-model-its-a-router","https://aihot.news/items/cmud2u29803s0rov6uljj8bm2"],"canonicalUrl":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","directAnswer":{"@type":"Answer","text":"Fireworks AI 发布 FireRouter，并以 DeepSWE v1.1 的 113 项工程任务、18 个模型分析路由潜力。正文称，单模型 GPT-6 Astra 的通过率为 74.1%、每项成本为 6.52 美元；事后择优的路由结果为 97.6%、每项 1.88 美元。","url":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","dateCreated":"2026-09-20T16:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Fireworks AI（网页） source article","url":"https://fireworks.ai/blog/the-frontier-isnt-a-model-its-a-router","datePublished":"2026-09-20T16:00:00.000Z","provider":{"@type":"Organization","name":"Fireworks AI（网页）","url":"https://fireworks.ai/blog/the-frontier-isnt-a-model-its-a-router"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmud2u29803s0rov6uljj8bm2","datePublished":"2026-09-20T16:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmud2u29803s0rov6uljj8bm2"}}],"aggregationSource":"Fireworks AI（网页）","originalPublisher":{"name":"Fireworks AI（网页）","url":"https://fireworks.ai/blog/the-frontier-isnt-a-model-its-a-router"},"geoDeepAnswer":null,"article":{"id":"cmud2u29803s0rov6uljj8bm2","slug":"cmud2u29803s0rov6uljj8bm2","url":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","title":"Fireworks AI 分析：前沿不是单个模型，而是路由器，FireRouter 发布","title_en":"","summary":"Fireworks AI 发布 FireRouter 并分析 DeepSWE v1.1 上 113 个任务、18 个模型的路由潜力。","source":"Fireworks AI（网页）","sourceUrl":"https://fireworks.ai/blog/the-frontier-isnt-a-model-its-a-router","aiHotUrl":"https://aihot.news/items/cmud2u29803s0rov6uljj8bm2","publishedAt":"2026-09-20T16:00:00.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["Join us for our inaugural conference, Forge 2026","How much better could a coding agent perform if it used the best model for each task?","The best single model, GPT-6 Astra, gets 74.1% of DeepSWE tasks at $6.52 each. Pick the right model for each task and the same eighteen models get 97.6% at $1.88. 23 points better, at under a third of the cost.","That number comes from hindsight. We ran all eighteen models on every task first and picked the winner for each one. What it measures is the capability already sitting in the pool, but it's split across models that nobody uses together.","Putting them together is a router's job. It picks which model handles each task before the work starts, and before is the hard part. Looking back, it's easy to point at a task and name the model that would have done it better. A router has to choose before it sees the outcome, and a wrong choice costs far more than the few dollars it saved.","We analyzed DeepSWE v1.1：https://deepswe.datacurve.ai/, an agentic coding benchmark where the unit of work is an engineering task: the agent has to understand an issue, inspect a repository, use tools, edit code, execute it, and get the task to pass.","The policy is deliberately simple. Pick one model at the start of a task and keep it for the whole run, with no switching mid-session.","Then we name the winner for each task by measured pass rate, breaking ties on cost. That's the oracle router. the same method we used in our Kimi K3 and Fable analysis：/blog/kimik3-fable.","The oracle scores on the same 113 tasks it picks from, using four rollouts per model-task pair, and taking a maximum over 18 noisy estimates biases it upward.","The best models score around 70% and spend $6.46 to $13.41 a task getting there:","That's the best a fixed-model policy does. Now pick per task:","The oracle router across all eighteen models reaches 97.6% at $1.88 a task . That is 23 points above GPT-6 Astra, at under a third of its cost. Restrict it to open-weight models only (DeepSeek V4 Flash and Pro, GLM-5.3 and GLM-5.3 Flash, Kimi K3, Qwen3.8 Max), and it still reaches 90.3% at $1.45 a task , which beats every closed model here by 16 points while spending under a quarter of what Astra does.","These results make \"open versus closed\" a less interesting debate. The emerging race is to move from the theoretical oracle router to building a system of models with collectively better intelligence than any single model. A system of open models can in principle already far surpass the closed frontier.","There is substantially more capability in the pool than any individual model exposes.","In our prior work, we found different models are sufficient (even exceptional) on different tasks ：/blog/kimik3-fable. The DeepSWE analysis makes the cost implications concrete. At that 97.6% point, the oracle still sends 94 of the 113 tasks to a model costing under $3 .","The three most expensive models in the field, all above $11.50 a task, are the sole best choice on only three tasks .","On 79 of the 113 tasks , at least one of those expensive models ties the top score and loses the task on price alone. A strong general-purpose model can be excellent across a broad distribution without being uniquely necessary on most individual tasks.","A fixed-model policy pays for broad capability on every task. A system can ask a narrower question:","What capability does this task actually require?","How many models does it take to capture the effect?","The best pair adds 13.1 points over the best single model, and the best trio reaches 91.2%. Expanding from three models to all eighteen adds another 6.4 percentage points. The useful object is not a catalog of hundreds of nearly interchangeable models. It is a portfolio with complementary coverage.","The value is capability coverage, not model count.","LLMRouterBench：https://aclanthology.org/2026.findings-acl.1881/ evaluates routing across 33 models and more than 400,000 instances. It finds that a handful of models covers most of what the full set can do, and that bigger pools add little without careful curation.","An oracle is easy to love because it never gets to be wrong. A production router does. We measured it strictly: we use pass@1, the probability that a single attempt passes, rather than a \"did this model ever succeed across four attempts\" rule. That second rule would make the ceiling look far more impressive while meaning much less.","The gap is a product problem and the literature is blunt about it. LLMRouterBench finds that several recent routing approaches, including commercial ones, fail to reliably beat simple baselines, and traces much of that to model recall: even when a model with the right capability exists in the pool, the router has to recognize when to reach for it.","So sticking with one model you know isn't conservative, it's rational: a stable error distribution beats a router that unpredictably picks the wrong specialist. The bar for a routing system is to make model specialization predictable enough that changing models improves the system without making its behavior less trustworthy .","Routing is usually introduced as a cost optimization: send easy work to a more cost-optimized model, reserve the expensive one for hard work, and keep the difference. At Fireworks：/nexus, we take a broader view.","If different models are genuinely complementary, then selecting among them moves you up the capability curve, not merely left along the cost curve.","That's what FireRouter is built for. It routes at the task level across both open and closed models, and it's cache aware, so switching models doesn't silently throw away the context you already paid for.","Over four weeks of our own production coding traffic, sessions routed through FireRouter cost $7.42 against $15.81 for Opus 5 alone, a 53% reduction across 2,334 sessions.","The useful unit of AI work is already larger than the single model call. A coding agent is a model inside a harness that supplies context, tools, execution, tests, state, and feedback.","Once several models have complementary strengths, the selection policy becomes a component of the system：https://www.faros.ai/blog/open-models-vs-frontier-models, alongside context, tools, and tests. Choosing and composing those components is the job. That's what AI engineering is.","Our experiment measures only the simplest version of that system: pick one model at the start of a task and leave it there. The selection policy is the part we can actually build.","We serve every frontier open model in production, which is where a real understanding of each model's strengths comes from. You do not learn what a model is uniquely good at from benchmark averages. You learn it by running all of them, on real work, at scale. That is where FireRouter's model choices come from, and that bar is the one we intend to clear. We will go into our own router：https://docs.fireworks.ai/ecosystem/firerouter/overview and how to hill-climb on your own specialized intelligence：/training in future posts.","Define your frontier on FireRouter：https://app.fireworks.ai/fire-router.","On costs. All cost figures in this analysis come from the DeepSWE leaderboard's published per-model numbers. The raw cost_usd in the public trials file does not match what the board displays, and for the DeepSeek family it differs by several times over, so each model’s per-task costs are scaled so its mean matches the published figure. Accuracy comes from the four raw rollouts of each task-model pair, cost from the board.","Source: DeepSWE v1.1 trials, refreshed 17 September 2026. 113 tasks, 18 models each at its best available configuration, 2,034 model-task cells."],"articleImages":[{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2F995c9bcc8c8ddb3be4015380119b11e78dd9c516-1600x900.png%3Fauto%3Dformat&w=3840&q=75","alt":"The frontier isn't a model. It's a router headline image","afterParagraph":0,"url":"/media/articles/cmud2u29803s0rov6uljj8bm2/6aa076e62377c634.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2F8da320b7d174b7f055a4cf8c1860a9b6daf9b732-2160x556.png%3Fauto%3Dformat&w=3840&q=75","alt":"one model per task","afterParagraph":6,"url":"/media/articles/cmud2u29803s0rov6uljj8bm2/aada3963d1c28a4c.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2Fd4fe4230f76f2a0abb136cb729b489e853f1362f-2160x1252.png%3Fauto%3Dformat&w=3840&q=75","alt":"hero results image","afterParagraph":10,"url":"/media/articles/cmud2u29803s0rov6uljj8bm2/551ec73b55d8c1ad.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2Fa35cfef1d64c5097d120c7d89b7fd0e2ba23c567-2160x896.png%3Fauto%3Dformat&w=3840&q=75","alt":"coverage ladder","afterParagraph":19,"url":"/media/articles/cmud2u29803s0rov6uljj8bm2/3103173b06bddb84.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2Fcd6c566e3bef7e937cad5b3a6bccdada224b7ef1-2160x896.png%3Fauto%3Dformat&w=3840&q=75","alt":"harness","afterParagraph":30,"url":"/media/articles/cmud2u29803s0rov6uljj8bm2/960e696478e16a99.webp"}],"mediaStatus":"ok","articleBodyZh":["加入我们，参加我们的首届会议，Forge 2026","如果编程代理对每个任务都使用最佳模型，它的表现会提升多少？","最佳的单一模型 GPT-6 Astra 在 DeepSWE 任务中达到 74.1%，每个任务花费 $6.52。为每个任务选择合适的模型，同样的十八个模型可以达到 97.6%，每个任务仅花费 $1.88。提高了 23 个百分点，成本不足三分之一。","这个数字来自事后分析。我们首先在每个任务上运行了所有十八个模型，然后选出每个任务的获胜者。它衡量的是已经存在于模型池中的能力，但这些能力分散在没有人一起使用的模型中。","将它们组合在一起是路由器的工作。在工作开始之前，它选择哪个模型处理每个任务，而“之前”才是难点。回头看很容易指出某个任务并说明哪款模型会做得更好。路由器必须在看到结果之前做出选择，而错误的选择所造成的成本远高于它节省的几美元。","我们分析了 DeepSWE v1.1：https://deepswe.datacurve.ai/，这是一个自主编码基准，其工作单元是一个工程任务：代理需要理解问题、检查代码仓库、使用工具、编辑代码、执行代码，并使任务通过。","策略故意保持简单。任务开始时选择一个模型，并在整个任务过程中保持该模型，不在中途切换。","然后我们按实际通过率为每个任务指定获胜者，成本相同时打平。这就是所谓的“神谕路由器”。我们在 Kimi K3 和 Fable 分析中使用了相同方法：/blog/kimik3-fable。","神谕路由器在同样的 113 个任务上评分，每个模型-任务组合进行四次尝试，并取 18 个不稳定估计的最大值，这使得评分偏高。","最佳模型得分约为 70%，每个任务花费 $6.46 至 $13.41：","这就是固定模型策略的最佳表现。现在按任务选择模型：","十八个模型的神谕路由器平均在每个任务 $1.88 时达到 97.6% 的通过率。这比 GPT-6 Astra 高 23 个百分点，成本不到三分之一。如果仅限于开源权重模型（DeepSeek V4 Flash 和 Pro，GLM-5.3 和 GLM-5.3 Flash，Kimi K3，Qwen3.8 Max），它仍能在每个任务 $1.45 时达到 90.3%，比这里的任何封闭模型高 16 个百分点，且花费不足 Astra 的四分之一。","这些结果使得“开放与封闭”的辩论不再那么有趣。新兴的竞争是从理论上的神谕路由器转向构建一个模型系统，其集体智能超过任何单一模型。原则上，一个开放模型系统已经可以远远超越封闭前沿。","池中的能力远超过任何单一模型所显示的能力。","在我们之前的工作中，我们发现不同的模型在不同的任务上是足够的（甚至是卓越的）：/blog/kimik3-fable。DeepSWE分析使成本影响变得具体。在97.6%的点上，神谕仍将113个任务中的94个分配给每个任务成本低于3美元的模型。","该领域三个最昂贵的模型，每个任务的成本都超过11.50美元，仅在三个任务中是唯一最佳选择。","在113个任务中的79个任务上，至少有一个昂贵模型的得分与最高得分持平，但仅因价格而失去任务。一种强大的通用模型可以在广泛的分布上表现出色，而在大多数单独任务上并非必不可少。","固定模型策略为每个任务支付广泛的能力。系统可以提出一个更狭窄的问题：","这个任务实际上需要什么能力？","需要多少模型才能捕捉到效果？","最佳组合的双模型比最佳单模型增加13.1个百分点，最佳三模型达到91.2%。从三模型扩展到全部十八模型，再增加6.4个百分点。真正有用的对象不是数百个几乎可互换模型的目录，而是具有互补覆盖的组合。","价值在于能力覆盖，而不是模型数量。","LLMRouterBench：https://aclanthology.org/2026.findings-acl.1881/ 评估了33个模型和超过400,000个实例的路由能力。它发现少量模型覆盖了完整集合的大部分能力，而且在没有精心策划的情况下，增大模型池几乎没有增加价值。","神谕容易让人喜爱，因为它永远不会出错。而生产路由器会出错。我们严格地衡量它：我们使用pass@1，即单次尝试成功的概率，而不是“这个模型在四次尝试中是否成功过”的规则。第二种规则会让上限看起来更令人印象深刻，但意义远小。","差距是一个产品问题，文献对此直接了当。LLMRouterBench发现在若干近期的路由方法中，包括商业方法，都无法可靠地击败简单的基准测试，且大部分原因归结于模型的记忆能力：即使池中存在具备合适能力的模型，路由器也必须识别何时调用它。","因此，坚持使用你熟悉的单一模型并不是保守，而是理性的：一个稳定的错误分布胜过一个不可预测地选择错误专家的路由器。路由系统的标准是使模型专业化足够可预测，这样更换模型能够提升系统性能而不降低其行为的可信度。","通常，路由被引入为成本优化：将简单任务发送给更优化的低成本模型，将昂贵的模型保留给复杂任务，并保留差额。在Fireworks：/nexus，我们采取更广的视角。","如果不同模型真正互补，那么在它们之间选择会让你沿能力曲线向上移动，而不仅仅是在成本曲线上向左移动。","这正是FireRouter的构建目标。它在任务层面跨开放模型和封闭模型进行路由，同时具有缓存感知能力，因此切换模型不会悄无声息地丢失你已经支付的上下文。","在我们自己四周的生产编码流量实验中，通过FireRouter路由的会话成本为7.42美元，而仅使用Opus 5的成本为15.81美元，2334个会话整体减少了53%。","有用的AI工作单位已经大于单次模型调用。一个编码代理是在提供上下文、工具、执行、测试、状态和反馈的套件中的模型。","一旦多个模型具有互补优势，选择策略就成为系统的一个组成部分：https://www.faros.ai/blog/open-models-vs-frontier-models，和上下文、工具及测试一起。选择和组合这些组成部分就是工作内容。这就是AI工程学。","我们的实验只衡量该系统的最简单版本：在任务开始时选择一个模型并保持使用。选择策略是我们实际上可以构建的部分。","我们在生产中服务每一个前沿开放模型，这正是对每个模型优势进行真实理解的来源。你无法从基准平均分中了解一个模型的独特优势。你需要通过运行所有模型，在实际工作中大规模地测试，才能得出结论。这就是 FireRouter 模型选择的来源，也是我们打算达到的标准。我们将在未来的文章中介绍我们自己的路由器：https://docs.fireworks.ai/ecosystem/firerouter/overview，以及如何在你自己的专用智能上进行爬坡优化：/training。","在 FireRouter 上定义你的前沿：https://app.fireworks.ai/fire-router。","关于成本。本分析中所有成本数据均来自 DeepSWE 排行榜上发布的每个模型的数据。公共试验文件中的原始 cost_usd 与排行榜显示的数据不一致，对于 DeepSeek 系列更是有几倍的差异，因此，每个模型的每任务成本会按比例调整，使其均值与发布数据一致。准确率来自每个任务-模型对的四次原始运行数据，成本来自排行榜。","来源：DeepSWE v1.1 试验，刷新于 2026 年 9 月 17 日。共有 113 个任务，每个任务包含 18 个模型，均采用其最佳可用配置，总共 2,034 个模型-任务单元。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Fireworks AI 发布 FireRouter，并以 DeepSWE v1.1 的 113 项工程任务、18 个模型分析路由潜力。正文称，单模型 GPT-6 Astra 的通过率为 74.1%、每项成本为 6.52 美元；事后择优的路由结果为 97.6%、每项 1.88 美元。","background":"该分析以单次工程任务为单位，任务涉及理解问题、检查代码库、使用工具、修改并执行代码，以使任务通过。策略是在任务开始时选定一个模型，整个运行期间不切换；oracle 则按每项任务的实测通过率选优，同分时比较成本。","viewpoint":"Aioga 判断：结果显示，18 个模型的事后最优组合在该基准上优于单一模型；但它是在运行全部模型后才选出赢家，不能直接视为任务开始前的路由器已达到这一表现，两者衡量条件不同。","implications":"可能影响：该测试使路由方案值得评估，但事后 oracle 的结果不足以证明实际路由器能在任务开始前选中合适模型，也不代表其他任务或使用条件会有相同通过率与成本；评估时需要区分事后上限和实际策略表现。","nextStep":"后续观察：可关注 FireRouter 在任务开始前如何选择模型，以及后续是否披露实际路由表现、成本口径和与 oracle 的差距；当前材料说明了事后择优方法及其数据，未提供这些比较结果。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-22T22:17:44.485Z","sourceHash":"9c34596c6f94353c","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":10,"blockingIssues":[],"notes":[]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Fireworks AI（网页）"],"translations":{"zh-CN":{"title":"Fireworks AI 分析：前沿不是单个模型，而是路由器，FireRouter 发布","summary":"Fireworks AI 发布 FireRouter 并分析 DeepSWE v1.1 上 113 个任务、18 个模型的路由潜力。","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI 分析：前沿不是单个模型，而是路由器，FireRouter 发布 - Aioga AI资讯","description":"Fireworks AI 发布 FireRouter 并分析 DeepSWE v1.1 上 113 个任务、18 个模型的路由潜力。","url":"https://www.aioga.com/news/cmud2u29803s0rov6uljj8bm2/","articleBody":["加入我们，参加我们的首届会议，Forge 2026","如果编程代理对每个任务都使用最佳模型，它的表现会提升多少？","最佳的单一模型 GPT-6 Astra 在 DeepSWE 任务中达到 74.1%，每个任务花费 $6.52。为每个任务选择合适的模型，同样的十八个模型可以达到 97.6%，每个任务仅花费 $1.88。提高了 23 个百分点，成本不足三分之一。","这个数字来自事后分析。我们首先在每个任务上运行了所有十八个模型，然后选出每个任务的获胜者。它衡量的是已经存在于模型池中的能力，但这些能力分散在没有人一起使用的模型中。","将它们组合在一起是路由器的工作。在工作开始之前，它选择哪个模型处理每个任务，而“之前”才是难点。回头看很容易指出某个任务并说明哪款模型会做得更好。路由器必须在看到结果之前做出选择，而错误的选择所造成的成本远高于它节省的几美元。","我们分析了 DeepSWE v1.1：https://deepswe.datacurve.ai/，这是一个自主编码基准，其工作单元是一个工程任务：代理需要理解问题、检查代码仓库、使用工具、编辑代码、执行代码，并使任务通过。","策略故意保持简单。任务开始时选择一个模型，并在整个任务过程中保持该模型，不在中途切换。","然后我们按实际通过率为每个任务指定获胜者，成本相同时打平。这就是所谓的“神谕路由器”。我们在 Kimi K3 和 Fable 分析中使用了相同方法：/blog/kimik3-fable。","神谕路由器在同样的 113 个任务上评分，每个模型-任务组合进行四次尝试，并取 18 个不稳定估计的最大值，这使得评分偏高。","最佳模型得分约为 70%，每个任务花费 $6.46 至 $13.41：","这就是固定模型策略的最佳表现。现在按任务选择模型：","十八个模型的神谕路由器平均在每个任务 $1.88 时达到 97.6% 的通过率。这比 GPT-6 Astra 高 23 个百分点，成本不到三分之一。如果仅限于开源权重模型（DeepSeek V4 Flash 和 Pro，GLM-5.3 和 GLM-5.3 Flash，Kimi K3，Qwen3.8 Max），它仍能在每个任务 $1.45 时达到 90.3%，比这里的任何封闭模型高 16 个百分点，且花费不足 Astra 的四分之一。","这些结果使得“开放与封闭”的辩论不再那么有趣。新兴的竞争是从理论上的神谕路由器转向构建一个模型系统，其集体智能超过任何单一模型。原则上，一个开放模型系统已经可以远远超越封闭前沿。","池中的能力远超过任何单一模型所显示的能力。","在我们之前的工作中，我们发现不同的模型在不同的任务上是足够的（甚至是卓越的）：/blog/kimik3-fable。DeepSWE分析使成本影响变得具体。在97.6%的点上，神谕仍将113个任务中的94个分配给每个任务成本低于3美元的模型。","该领域三个最昂贵的模型，每个任务的成本都超过11.50美元，仅在三个任务中是唯一最佳选择。","在113个任务中的79个任务上，至少有一个昂贵模型的得分与最高得分持平，但仅因价格而失去任务。一种强大的通用模型可以在广泛的分布上表现出色，而在大多数单独任务上并非必不可少。","固定模型策略为每个任务支付广泛的能力。系统可以提出一个更狭窄的问题：","这个任务实际上需要什么能力？","需要多少模型才能捕捉到效果？","最佳组合的双模型比最佳单模型增加13.1个百分点，最佳三模型达到91.2%。从三模型扩展到全部十八模型，再增加6.4个百分点。真正有用的对象不是数百个几乎可互换模型的目录，而是具有互补覆盖的组合。","价值在于能力覆盖，而不是模型数量。","LLMRouterBench：https://aclanthology.org/2026.findings-acl.1881/ 评估了33个模型和超过400,000个实例的路由能力。它发现少量模型覆盖了完整集合的大部分能力，而且在没有精心策划的情况下，增大模型池几乎没有增加价值。","神谕容易让人喜爱，因为它永远不会出错。而生产路由器会出错。我们严格地衡量它：我们使用pass@1，即单次尝试成功的概率，而不是“这个模型在四次尝试中是否成功过”的规则。第二种规则会让上限看起来更令人印象深刻，但意义远小。","差距是一个产品问题，文献对此直接了当。LLMRouterBench发现在若干近期的路由方法中，包括商业方法，都无法可靠地击败简单的基准测试，且大部分原因归结于模型的记忆能力：即使池中存在具备合适能力的模型，路由器也必须识别何时调用它。","因此，坚持使用你熟悉的单一模型并不是保守，而是理性的：一个稳定的错误分布胜过一个不可预测地选择错误专家的路由器。路由系统的标准是使模型专业化足够可预测，这样更换模型能够提升系统性能而不降低其行为的可信度。","通常，路由被引入为成本优化：将简单任务发送给更优化的低成本模型，将昂贵的模型保留给复杂任务，并保留差额。在Fireworks：/nexus，我们采取更广的视角。","如果不同模型真正互补，那么在它们之间选择会让你沿能力曲线向上移动，而不仅仅是在成本曲线上向左移动。","这正是FireRouter的构建目标。它在任务层面跨开放模型和封闭模型进行路由，同时具有缓存感知能力，因此切换模型不会悄无声息地丢失你已经支付的上下文。","在我们自己四周的生产编码流量实验中，通过FireRouter路由的会话成本为7.42美元，而仅使用Opus 5的成本为15.81美元，2334个会话整体减少了53%。","有用的AI工作单位已经大于单次模型调用。一个编码代理是在提供上下文、工具、执行、测试、状态和反馈的套件中的模型。","一旦多个模型具有互补优势，选择策略就成为系统的一个组成部分：https://www.faros.ai/blog/open-models-vs-frontier-models，和上下文、工具及测试一起。选择和组合这些组成部分就是工作内容。这就是AI工程学。","我们的实验只衡量该系统的最简单版本：在任务开始时选择一个模型并保持使用。选择策略是我们实际上可以构建的部分。","我们在生产中服务每一个前沿开放模型，这正是对每个模型优势进行真实理解的来源。你无法从基准平均分中了解一个模型的独特优势。你需要通过运行所有模型，在实际工作中大规模地测试，才能得出结论。这就是 FireRouter 模型选择的来源，也是我们打算达到的标准。我们将在未来的文章中介绍我们自己的路由器：https://docs.fireworks.ai/ecosystem/firerouter/overview，以及如何在你自己的专用智能上进行爬坡优化：/training。","在 FireRouter 上定义你的前沿：https://app.fireworks.ai/fire-router。","关于成本。本分析中所有成本数据均来自 DeepSWE 排行榜上发布的每个模型的数据。公共试验文件中的原始 cost_usd 与排行榜显示的数据不一致，对于 DeepSeek 系列更是有几倍的差异，因此，每个模型的每任务成本会按比例调整，使其均值与发布数据一致。准确率来自每个任务-模型对的四次原始运行数据，成本来自排行榜。","来源：DeepSWE v1.1 试验，刷新于 2026 年 9 月 17 日。共有 113 个任务，每个任务包含 18 个模型，均采用其最佳可用配置，总共 2,034 个模型-任务单元。"]},"en":{"title":"Fireworks AI Analysis: The frontier is not a single model, but a router, FireRouter released","summary":"Fireworks AI releases FireRouter and analyzes the routing potential of 113 tasks and 18 models on DeepSWE v1.1.","category":"Industry","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI Analysis: The frontier is not a single model, but a router, FireRouter released - Aioga AI News","description":"Fireworks AI releases FireRouter and analyzes the routing potential of 113 tasks and 18 models on DeepSWE v1.1.","url":"https://www.aioga.com/en/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:03:28.442Z"},"ja":{"title":"Fireworks AI 分析：最先端は単一のモデルではなくルーター、FireRouter 発表","summary":"Fireworks AI は FireRouter をリリースし、DeepSWE v1.1 上の 113 のタスクと 18 のモデルのルーティング可能性を分析しました。","category":"業界動向","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI 分析：最先端は単一のモデルではなくルーター、FireRouter 発表 - Aioga AIニュース","description":"Fireworks AI は FireRouter をリリースし、DeepSWE v1.1 上の 113 のタスクと 18 のモデルのルーティング可能性を分析しました。","url":"https://www.aioga.com/ja/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:03:30.227Z"},"ko":{"title":"Fireworks AI 분석: 최전선은 단일 모델이 아니라 라우터, FireRouter 발표","summary":"Fireworks AI가 FireRouter를 발표하고 DeepSWE v1.1에서 113개의 작업과 18개의 모델의 라우팅 잠재력을 분석했습니다.","category":"업계 동향","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI 분석: 최전선은 단일 모델이 아니라 라우터, FireRouter 발표 - Aioga AI 뉴스","description":"Fireworks AI가 FireRouter를 발표하고 DeepSWE v1.1에서 113개의 작업과 18개의 모델의 라우팅 잠재력을 분석했습니다.","url":"https://www.aioga.com/ko/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:04:08.063Z"},"es":{"title":"Análisis de Fireworks AI: La vanguardia no es un solo modelo, sino un enrutador, FireRouter lanzado","summary":"Fireworks AI lanza FireRouter y analiza el potencial de enrutamiento de 113 tareas y 18 modelos en DeepSWE v1.1.","category":"Industria","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Análisis de Fireworks AI: La vanguardia no es un solo modelo, sino un enrutador, FireRouter lanzado - Aioga Noticias de IA","description":"Fireworks AI lanza FireRouter y analiza el potencial de enrutamiento de 113 tareas y 18 modelos en DeepSWE v1.1.","url":"https://www.aioga.com/es/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:04:09.240Z"},"fr":{"title":"Analyse de Fireworks AI : l'avant-garde n'est pas un seul modèle, mais un routeur, FireRouter publié","summary":"Fireworks AI a publié FireRouter et analysé le potentiel de routage de 113 tâches et 18 modèles sur DeepSWE v1.1.","category":"Industrie","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Analyse de Fireworks AI : l'avant-garde n'est pas un seul modèle, mais un routeur, FireRouter publié - Aioga Actualités IA","description":"Fireworks AI a publié FireRouter et analysé le potentiel de routage de 113 tâches et 18 modèles sur DeepSWE v1.1.","url":"https://www.aioga.com/fr/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:04:46.682Z"},"de":{"title":"Fireworks AI Analyse: Die Spitze ist nicht ein einzelnes Modell, sondern ein Router, FireRouter veröffentlicht","summary":"Fireworks AI veröffentlicht FireRouter und analysiert das Routing-Potenzial von 113 Aufgaben und 18 Modellen auf DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI Analyse: Die Spitze ist nicht ein einzelnes Modell, sondern ein Router, FireRouter veröffentlicht - Aioga KI-News","description":"Fireworks AI veröffentlicht FireRouter und analysiert das Routing-Potenzial von 113 Aufgaben und 18 Modellen auf DeepSWE v1.1.","url":"https://www.aioga.com/de/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:04:46.993Z"},"pt-BR":{"title":"Análise do Fireworks AI: a vanguarda não é um único modelo, mas um roteador, FireRouter lançado","summary":"Fireworks AI lançou o FireRouter e analisou o potencial de roteamento de 113 tarefas e 18 modelos no DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Análise do Fireworks AI: a vanguarda não é um único modelo, mas um roteador, FireRouter lançado - Aioga Notícias de IA","description":"Fireworks AI lançou o FireRouter e analisou o potencial de roteamento de 113 tarefas e 18 modelos no DeepSWE v1.1.","url":"https://www.aioga.com/pt-BR/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:05:25.508Z"},"ru":{"title":"Анализ Fireworks AI: передний край — это не отдельная модель, а маршрутизатор, выпущен FireRouter","summary":"Fireworks AI выпустила FireRouter и проанализировала потенциал маршрутизации 113 задач и 18 моделей на DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Анализ Fireworks AI: передний край — это не отдельная модель, а маршрутизатор, выпущен FireRouter - Aioga Новости ИИ","description":"Fireworks AI выпустила FireRouter и проанализировала потенциал маршрутизации 113 задач и 18 моделей на DeepSWE v1.1.","url":"https://www.aioga.com/ru/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:05:25.533Z"},"ar":{"title":"تحليل Fireworks AI: الطليعة ليست نموذجًا فرديًا، بل هي الموجه، تم إصدار FireRouter","summary":"أصدرت Fireworks AI FireRouter وحللت إمكانيات التوجيه لـ 113 مهمة و18 نموذجًا على DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"تحليل Fireworks AI: الطليعة ليست نموذجًا فرديًا، بل هي الموجه، تم إصدار FireRouter - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت Fireworks AI FireRouter وحللت إمكانيات التوجيه لـ 113 مهمة و18 نموذجًا على DeepSWE v1.1.","url":"https://www.aioga.com/ar/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:06:08.559Z"},"hi":{"title":"Fireworks AI विश्लेषण: अग्रिम सीमा कोई एकल मॉडल नहीं है, बल्कि राउटर है, FireRouter जारी","summary":"फायरवर्क्स एआई ने फायरराउटर जारी किया और डीपएसडब्ल्यूई v1.1 पर 113 कार्यों और 18 मॉडलों की रूटिंग क्षमता का विश्लेषण किया।","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI विश्लेषण: अग्रिम सीमा कोई एकल मॉडल नहीं है, बल्कि राउटर है, FireRouter जारी - Aioga AI समाचार","description":"फायरवर्क्स एआई ने फायरराउटर जारी किया और डीपएसडब्ल्यूई v1.1 पर 113 कार्यों और 18 मॉडलों की रूटिंग क्षमता का विश्लेषण किया।","url":"https://www.aioga.com/hi/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:06:11.152Z"},"it":{"title":"Analisi di Fireworks AI: l'avanguardia non è un singolo modello, ma un router, FireRouter rilasciato","summary":"Fireworks AI ha rilasciato FireRouter e ha analizzato il potenziale di instradamento di 113 compiti e 18 modelli su DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Analisi di Fireworks AI: l'avanguardia non è un singolo modello, ma un router, FireRouter rilasciato - Aioga Notizie IA","description":"Fireworks AI ha rilasciato FireRouter e ha analizzato il potenziale di instradamento di 113 compiti e 18 modelli su DeepSWE v1.1.","url":"https://www.aioga.com/it/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:06:47.333Z"},"nl":{"title":"Fireworks AI-analyse: de voorhoede is niet een enkel model, maar een router, FireRouter uitgebracht","summary":"Fireworks AI heeft FireRouter uitgebracht en analyseert het routeringspotentieel van 113 taken en 18 modellen op DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI-analyse: de voorhoede is niet een enkel model, maar een router, FireRouter uitgebracht - Aioga AI-nieuws","description":"Fireworks AI heeft FireRouter uitgebracht en analyseert het routeringspotentieel van 113 taken en 18 modellen op DeepSWE v1.1.","url":"https://www.aioga.com/nl/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:06:49.592Z"},"tr":{"title":"Fireworks AI Analizi: Ön cephe tek bir model değil, bir yönlendiricidir, FireRouter yayınlandı","summary":"Fireworks AI, FireRouter'ı yayınladı ve DeepSWE v1.1 üzerindeki 113 görev ve 18 modelin yönlendirme potansiyelini analiz etti.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks AI Analizi: Ön cephe tek bir model değil, bir yönlendiricidir, FireRouter yayınlandı - Aioga AI Haberleri","description":"Fireworks AI, FireRouter'ı yayınladı ve DeepSWE v1.1 üzerindeki 113 görev ve 18 modelin yönlendirme potansiyelini analiz etti.","url":"https://www.aioga.com/tr/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:07:29.759Z"},"vi":{"title":"Phân tích Fireworks AI: Tiên tiến không phải là một mô hình đơn lẻ, mà là bộ định tuyến, FireRouter ra mắt","summary":"Fireworks AI phát hành FireRouter và phân tích tiềm năng định tuyến của 113 nhiệm vụ và 18 mô hình trên DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Phân tích Fireworks AI: Tiên tiến không phải là một mô hình đơn lẻ, mà là bộ định tuyến, FireRouter ra mắt - Tin tức AI Aioga","description":"Fireworks AI phát hành FireRouter và phân tích tiềm năng định tuyến của 113 nhiệm vụ và 18 mô hình trên DeepSWE v1.1.","url":"https://www.aioga.com/vi/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:07:27.955Z"},"id":{"title":"Analisis Fireworks AI: Garis depan bukan model tunggal, melainkan router, FireRouter dirilis","summary":"Fireworks AI merilis FireRouter dan menganalisis potensi routing dari 113 tugas dan 18 model pada DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Analisis Fireworks AI: Garis depan bukan model tunggal, melainkan router, FireRouter dirilis - Berita AI Aioga","description":"Fireworks AI merilis FireRouter dan menganalisis potensi routing dari 113 tugas dan 18 model pada DeepSWE v1.1.","url":"https://www.aioga.com/id/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:08:05.617Z"},"th":{"title":"การวิเคราะห์ Fireworks AI: ขอบเขตล้ำหน้าไม่ใช่โมเดลเดียว แต่คือเราเตอร์, FireRouter เปิดตัว","summary":"Fireworks AI เปิดตัว FireRouter และวิเคราะห์ศักยภาพการกำหนดเส้นทางของ 113 งานและ 18 โมเดลบน DeepSWE v1.1","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"การวิเคราะห์ Fireworks AI: ขอบเขตล้ำหน้าไม่ใช่โมเดลเดียว แต่คือเราเตอร์, FireRouter เปิดตัว - ข่าว AI Aioga","description":"Fireworks AI เปิดตัว FireRouter และวิเคราะห์ศักยภาพการกำหนดเส้นทางของ 113 งานและ 18 โมเดลบน DeepSWE v1.1","url":"https://www.aioga.com/th/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:08:10.273Z"},"pl":{"title":"Analiza Fireworks AI: granica nie leży w pojedynczym modelu, lecz w routerze, FireRouter wydany","summary":"Fireworks AI wprowadza FireRouter i analizuje potencjał routingu 113 zadań i 18 modeli na DeepSWE v1.1.","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Analiza Fireworks AI: granica nie leży w pojedynczym modelu, lecz w routerze, FireRouter wydany - Aioga Wiadomości AI","description":"Fireworks AI wprowadza FireRouter i analizuje potencjał routingu 113 zadań i 18 modeli na DeepSWE v1.1.","url":"https://www.aioga.com/pl/news/cmud2u29803s0rov6uljj8bm2/","contentTranslated":true,"sourceHash":"dcc4fa479ae30413","translatedAt":"2026-09-22T20:08:47.929Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}