代理可以生成代码。对于你的系统、团队规范以及过去的决策来说,做到正确是最难的部分。你最终会在修正循环中浪费时间和令牌。
更多的多代理协作协议(MCP)、规则和更大的上下文窗口可以让代理访问信息,但不能保证理解。领先的团队拥有一个上下文层,为代理提供任务所需的精确信息。
加入我们于 9 月 23 日举办的免费网络研讨会:https://go.bytebytego.com/Unblocked_090926,了解:
团队在 AI 成熟度曲线上遇到的障碍,以及为什么常见的解决办法不足
上下文层如何解决质量、效率和成本问题
实况演示:同一个编码任务在有无上下文层下的表现
如果你想最大化 AI 代理带来的价值,这场研讨会绝对值得你的时间。
立即注册:https://go.bytebytego.com/Unblocked_090926
当应用程序采用大型语言模型(LLM)时,通常会选择最强大的模型。这意味着每一个请求都将发送到那个昂贵的模型。
虽然这种方法更易于实施,但从长远来看可能相当昂贵。例如,一个 “将此支持工单分类为账单、技术或账户相关” 的请求不需要像 “调查为什么这些财务记录不匹配,并解释可能的原因” 这样高水平的推理。
通过智能模型路由,我们可以解决这个问题。在这种路由方法中,我们为每个请求选择特定的模型。换句话说,简单的工作被发送到可能更便宜的小模型,而复杂的工作则路由到更强大的模型。如果大多数请求都很简单,这种方法可以大幅降低总成本,有时甚至可降低约 10 倍。同时,响应质量不会明显下降。
然而,成本降低并非理所当然。它还取决于应用程序接收的请求类型、模型之间的价格差异以及路由系统的性能。在本文中,我们将探讨各个方面。涵盖内容如下:
为什么 LLM 应用程序会变得昂贵?
模型路由如何提供成本节省?
在回答之前如何判断一个请求?
级联:首先尝试更便宜的模型
使用大型语言模型(LLM)API 的总成本通常取决于处理的 tokens 数量。
明确来说,token 是文本的一个小单位。一个短词可能是一个 token,但一个长词可能会被拆分为多个 token。
通常有两个重要的标记计数:
输入 tokens 包括用户的消息、系统指令、对话历史以及提供给模型的任何文档。
输出标记是响应中生成的标记。
根据不同的 LLM 提供商,输入和输出 tokens 的价格可能不同。更大、更强大的模型通常费用更高,因为它们需要更多的计算资源。它们可能还会花费额外的计算来进行推理。这种额外的能力对于解决复杂问题非常重要。但当任务很简单时,这种能力就被浪费了。
例如,想象一个每月需要处理一百万条请求的客户支持应用程序。在这些请求中,有些用户可能会询问退款政策,有些可能希望从电子邮件中提取地址,还有些可能有需要仔细分析的复杂账户问题。如果每个请求都使用最强大的模型处理,公司即使对于这些对该模型来说很简单的工作也必须支付高价。你可以把这想象成雇佣一个高级软件架构师来重命名文件、整理支持工单和格式化日期。当然,架构师技术上可以完成这些工作,但这会浪费架构师的能力,也是一种资源管理不当。
使用$5积分免费试用Crusoe:https://go.bytebytego.com/Crusoe_090826
模型路由是检查传入请求以决定哪个模型最适合处理它的过程。
通过模型路由,我们不会编写应用代码一直盲目调用某个模型。我们在多个模型前放置一个路由器。路由器可以访问小模型、中等模型以及高能力模型。它的任务是评估每个请求,并将其发送到最合适的模型。
例如,路由器可能会收到一个简单的分类请求,并将其发送到最小的模型。或者,路由器可能会收到包含复杂法律比较的请求,并将其发送到最强大的模型。
你可以将模型路由看作负载均衡。但它有一个重要的区别。负载均衡器通常在大体相当的服务器之间分配流量。然而,模型路由器必须在能力、成本和特性差异巨大的模型之间做出选择。
模型路由也与专家混合(MoE)模型有很大不同。在MoE设置中,路由发生在单个模型的内部部分之间。相比之下,应用层模型路由发生在模型之外。它处理的是决定哪个模型应接收请求,而不是处理该模型的内部结构。
考虑一个每次平均请求成本为1美分的强大模型。如果一个应用处理一百万个请求,使用该强大模型处理所有请求的费用约为10,000美元。
现在假设一个较小的模型的成本仅为强大模型的1/20,而中等模型的成本是强大模型的1/5。经过工作量分析,我们发现85%的请求可以由小模型处理,10%需要中等模型,而只有5%需要强大模型。
在这种情况下,每次请求的平均成本变为:
(0.85×0.05) + (0.10×0.20) + (0.05×1.00) = 0.1125
这意味着,使用模型路由构建的系统的成本可能仅为使用同一强大模型处理每个请求的系统成本的11%。这几乎将成本降低了10倍。
更有利的流量模式或价格差异可能会将节省推高到十倍以上。例如,如果超过90%的工作量由抽取、分类、格式化和简单摘要生成组成,则昂贵的模型可能只需偶尔使用。
最终,最佳节省发生在满足三个条件时:模型之间存在较大价格差异,大多数请求相对简单,以及路由器能够可靠地识别简单请求。
模型路由的最大难点在于如何在不回答请求的情况下确定请求的难度级别。
如果一个请求很短,并不一定意味着这个请求很简单。例如,“合同有效吗?”只有4个单词。但要安全地回答这个问题,模型可能需要法律专业知识和广泛的背景信息。另一方面,一个长请求并不总是困难的。用户可能粘贴了一份长文档,并要求模型提取每一个电子邮件地址。从概念上来说,这相当直接。
因此,一个好的路由器不能仅依靠消息长度来判断难度等级。它需要在做出公平决定时检查多个信号。
例如,模型路由器可能会考虑用户请求的任务类型。这是因为分类、提取、翻译、改写和格式化等任务通常需要较少的推理。然而,涉及计划、调试、数学证明或比较冲突文档的任务需要更高水平的推理能力。
模型路由器还应考虑风险因素。例如,医学、法律、金融或安全相关的问题,即使查询看起来简单,也可能被路由到更强的模型。这是因为答案不准确的代价非常高。
模型路由器可以使用的另一个信号是整体上下文的数量。例如,如果模型需要检查多个文档、理解长对话或连接不同的来源,它就需要更大的上下文窗口或更强的指令能力。
最后,模型路由器在选择合适模型之前,也可能需要检查输出要求。例如,生成带有几个已知字段的有效 JSON 对象可能是一项简单的任务。然而,生成符合一系列关键约束的详细技术设计则要困难得多。
换句话说,单一信号不足以决定。一个智能的模型路由方法通常会结合多个信号来做出正确选择。
最灵活的模型路由方法是使用较小的模型对请求进行分类。
所谓的路由模型可以按以下指示工作:
然后,路由器可以返回一个小型的结构化结果:
由于路由提示和由此产生的响应都比较短,因此分类调用的成本不会太高。基于响应,应用程序随后将完整请求发送到所选模型。
虽然这种方法比编写固定规则更适用于自然语言,但它可能有另一种出错的原因。较小的路由模型可能误解请求,并将难度较大的工作发送给能力较低的模型。这就是为什么生产系统通常将基于模型的分类与固定安全规则结合起来。某些特定规则可能明确规定,某些医疗或金融查询无论路由模型建议如何,都应始终发送到最强大的模型。
现在让我们看看另一种有用的模型路由策略,称为模型级联。
在这种策略中,我们并不试图完全预测请求的难度。相反,系统首先将请求发送到成本较低的模型。然后检查答案是否足够好。如果答案未通过检查,系统将请求发送到更强的模型。
当答案可以自动检查时,这种方法效果很好。例如,假设应用程序要求模型从发票中提取日期、客户ID和总金额。程序随后可以验证所有必填字段是否存在,日期是否有效,金额是否为数字。如果小模型产生了格式错误的数据,第二次尝试将发送给强大的模型。
在代码生成的情况下,我们也有类似的机会。应用程序可以对生成的代码运行测试。如果测试通过,它接受成本较低模型的答案。如果测试失败,它可以将任务升级到能力更强的模型。
然而,当质量判断具有主观性时,级联方法就变得困难。可能没有简单的自动化测试来判断商业策略是否有用,或者解释是否真正清楚。在这种情况下,应用程序可能会使用单独的评估模型。然而,这样的模型也会有其自身的成本,并且也可能出错。
最后,级联过程必须仔细设计,因为失败的尝试也会消耗资金和时间。如果小模型的大部分尝试都失败了,应用程序最终只会为小模型和高能力模型都付费。路由最终会使系统变慢且成本更高。
在语义路由中,我们根据请求的意义而不是特定关键词选择模型。
例如,考虑一个应用程序,它为计费、技术支持、产品推荐和账户安全提供了专门的模型或提示。但是,用户可能以许多不同方式描述同一个计费问题:
同一个订单在我的信用卡上出现了两次。
典型的基于关键词的系统可能无法支持许多这些变体。但是,语义路由器会将请求转换为嵌入表示。嵌入表示是一种对请求意义的数值化表示。
然后,路由器可以将该嵌入与已知请求类别的示例进行比较。如果请求接近计费示例,它将发送到计费模型。如果它类似于账户安全示例,它将发送到安全模型。
Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.
More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.
Join us for a FREE webinar on Sep 23 :https://go.bytebytego.com/Unblocked_090926 to see:
Where teams get stuck on the AI maturity curve and why common fixes fall short
How a context layer solves for quality, efficiency, and cost
Live demo: the same coding task with and without a context layer
If you want to maximize the value you get from AI agents, this one is worth your time.
Register now :https://go.bytebytego.com/Unblocked_090926
When an application adopts a large language model (LLM), they generally choose the most capable model possible. This means that every single request is sent to that expensive model.
While this approach is easier to implement, it can become quite expensive in the long run. For example, a request such as “classify this support ticket as billing, technical, or account-related” doesn’t require the same level of reasoning as “investigate why these financial records don’t match properly and explain the likely cause.”
With smart model routing, we can solve this problem. In such a routing approach, we choose a specific model for each request. In other words, simple work is sent to a small model that might be less expensive, and difficult work is routed to a more capable model. If most requests are simple, this approach can reduce the total cost in a big way, sometimes by even around 10 times. Also, the quality of the response doesn’t go down noticeably.
However, cost reduction isn’t a given. It also depends on the types of requests the application receives, the price difference between models, and how well the routing system performs. In this article, we are going to look at various aspects. Here’s what we will cover:
Why do LLM applications become expensive?
How can model routing provide cost savings?
How to judge a request before answering?
Cascading: Trying the cheaper model first
The total cost of using an LLM API usually depends on the number of tokens processed.
To be clear, a token is a small unit of text. A short word might be one token. But a longer word can be split into multiple tokens.
There are usually two important token counts:
Input tokens include the user’s message, system instructions, conversation history, and any documents supplied to the model.
Output tokens are the tokens generated within the response.
Depending on the LLM provider, input and output tokens can have different price points. Larger and more capable models generally cost more because they require more computing resources. They may also spend additional computation for reasoning. This extra capability is very important for solving complex problems. But this capability is wasted when the task is simple.
For example, imagine a customer-support application that has to process a million requests per month. Within those requests, some users may ask for refund policies. Others may want an address extracted from an email. Some might have complicated account problems that need careful analysis. If each request goes to the most powerful model, the company has to pay a premium price even for work that is quite simple for this capable model. You could think of this as hiring a senior software architect to rename files, sort support tickets, and format dates. Sure, the architect can technically do those things. But it would be a waste of the architect’s capability and a case of poor resource management.
Try Crusoe for free with $5 in credits :https://go.bytebytego.com/Crusoe_090826
Model routing is the process of checking an incoming request to decide which model is the best choice for handling it.
With model routing, we don’t write application code that always calls one model blindly. We place a router in front of several models. The router can access a small model, a medium-sized model, and a highly capable one. Its job is to evaluate each request and send it to the most suitable model.
For example, the router might receive a simple classification request and send it to the smallest model. Or the router might receive a request that contains a complicated legal comparison and send it to the most powerful model.
You can think of model routing as load balancing. But it has an important difference. A load balancer normally distributes traffic between largely equivalent servers. However, a model router has to choose between models with vastly different capabilities, costs, and characteristics.
Model routing is also quite different from a mixture-of-experts (MoE) model. In an MoE setup, routing happens internally between parts of a single model. In contrast, application-level model routing happens outside the models. It deals with deciding which model should receive the request and doesn’t deal with the internals of that model.
Consider a powerful model that costs 1 cent per average request. If an application handles a million requests, using that powerful model for everything would cost around $10,000.
Now imagine a smaller model costs only 1/20th as much, while a medium model costs 1/5th as much as the powerful model. After studying the workload, we discover that 85% of requests can be handled by the small model, 10% need the medium model, and just 5% require the powerful model.
In this case, the average cost per request becomes:
(0.85×0.05) + (0.10×0.20) + (0.05×1.00) = 0.1125
This means that a system built with model routing can potentially cost just 11% as much as the system that uses the same powerful model for handling every request. This is almost a 10X reduction in costs.
Even more favourable traffic patterns or price differences could push the savings beyond tenfold. For example, if more than 90% of the workload consists of extraction, classification, formatting, and straightforward summary generation, the expensive model may be needed only occasionally.
Ultimately, the best savings happen when three conditions are met: a large price difference between models, most requests being relatively simpler, and the router being able to identify the simple requests reliably.
The greatest difficulty in model routing is around determining the difficulty level of a request without answering it.
If a request is short, it doesn’t necessarily mean that the request is simple. For example, “Is the contract valid?” contains just 4 words. But to answer this query safely, the model might need legal expertise and extensive context. On the other hand, a long request is not always difficult. A user may have pasted a long document and asked the model to extract every email address. It is conceptually quite straightforward.
Therefore, a good router cannot rely only on message length to determine the difficulty level. It needs to check several signals while making a fair decision.
For example, the model router might consider what kind of task the user is requesting. This is because tasks like classification, extraction, translation, rewriting, and formatting often require less reasoning. However, tasks that involve planning, debugging, mathematical proofs, or comparing conflicting documents require much higher levels of reasoning.
The model router should also consider the risk factor. For example, a medical, legal, financial, or security-related question may be routed to a stronger model even if the query appears simple. This is because the cost of an inaccurate answer matters a lot.
Another signal the model router could use is the amount of overall context. For example, if a model needs to inspect several documents, make sense of a long conversation, or connect different sources, it needs a larger context window or stronger instruction ability.
Lastly, the model router may also need to check the output requirements before selecting the right model. For example, producing a valid JSON object with a few known fields may be an easy task. However, producing a detailed technical design that adheres to a bunch of critical constraints is much harder.
In other words, no single signal is sufficient. A smart model routing approach normally combines several signals to make the right choice.
The most flexible approach to model routing is to use a smaller model to classify the request.
This so-called router model can work on instructions as follows:
The router can then return a small structured result:
Since the routing prompt and the resulting response are quite short, the classification call won’t be too costly. Based on the response, the application then sends the full request to the selected model.
While this approach deals better with natural language rather than coding fixed rules, it can have another cause of error. The smaller router model can misunderstand the request and send difficult work to a less-capable model. This is why production systems often combine model-based classification with fixed safety rules. A specific rule might clearly specify that certain medical or financial queries should always be sent to the strongest model, irrespective of what the router model suggests.
Let us now look at another useful model routing strategy known as model cascading.
In this strategy, we don’t try to predict the difficulty perfectly. Instead, the system first sends the request to a cheaper model. It then checks whether the answer appears good enough. If the answer fails the check, the system sends the request to a stronger model.
This approach works quite well when answers can be checked automatically. For example, let’s say the application asks the model to extract a date, customer ID, and total amount from an invoice. The program can then verify that all required fields exist, the date is valid, and the amount is numeric. If the small model has produced malformed data, the second attempt goes to the powerful model.
We get similar opportunities in the case of code generation. The application can run tests against the generated code. If the tests pass, it accepts the cheaper model’s answer. If they fail, it can escalate the task to the more capable model.
However, cascading gets difficult when quality judgement is subjective. There may be no simple automated test to find out if a business strategy is useful or whether an explanation is actually clear. In those cases, the application may use a separate evaluator model. However, such a model would have its own cost and can also make mistakes.
Lastly, the cascade process must be designed carefully because failed attempts also consume money and time. If most attempts made by the small model end up in failure, the application only ends up paying for both the small and the capable model. Routing ends up making the system slower and more expensive.
In semantic routing, we choose the model based on the meaning of the request rather than specific keywords.
For example, consider an application that has specialized models or prompts for billing, technical support, product recommendations, and account security. However, users may describe the same billing problem in many different ways:
The same order appears twice on my card.
A typical keyword-based system might not be able to support many of these variations. But a semantic router converts the request into an embedding. For reference, an embedding is a numerical representation of a request’s meaning.
The router can then compare the embedding with examples of known request categories. If the request is close to billing examples, it goes to the billing model. If it resembles account-security examples, it goes to the security model.