Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.

More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.

Join us for a FREE webinar on Sep 23 :https://go.bytebytego.com/Unblocked_090926 to see:

Where teams get stuck on the AI maturity curve and why common fixes fall short

How a context layer solves for quality, efficiency, and cost

Live demo: the same coding task with and without a context layer

If you want to maximize the value you get from AI agents, this one is worth your time.

Register now :https://go.bytebytego.com/Unblocked_090926

When an application adopts a large language model (LLM), they generally choose the most capable model possible. This means that every single request is sent to that expensive model.

While this approach is easier to implement, it can become quite expensive in the long run. For example, a request such as “classify this support ticket as billing, technical, or account-related” doesn’t require the same level of reasoning as “investigate why these financial records don’t match properly and explain the likely cause.”

With smart model routing, we can solve this problem. In such a routing approach, we choose a specific model for each request. In other words, simple work is sent to a small model that might be less expensive, and difficult work is routed to a more capable model. If most requests are simple, this approach can reduce the total cost in a big way, sometimes by even around 10 times. Also, the quality of the response doesn’t go down noticeably.

However, cost reduction isn’t a given. It also depends on the types of requests the application receives, the price difference between models, and how well the routing system performs. In this article, we are going to look at various aspects. Here’s what we will cover:

Why do LLM applications become expensive?

How can model routing provide cost savings?

How to judge a request before answering?

Cascading: Trying the cheaper model first

The total cost of using an LLM API usually depends on the number of tokens processed.

To be clear, a token is a small unit of text. A short word might be one token. But a longer word can be split into multiple tokens.

There are usually two important token counts:

Input tokens include the user’s message, system instructions, conversation history, and any documents supplied to the model.

Output tokens are the tokens generated within the response.

Depending on the LLM provider, input and output tokens can have different price points. Larger and more capable models generally cost more because they require more computing resources. They may also spend additional computation for reasoning. This extra capability is very important for solving complex problems. But this capability is wasted when the task is simple.

For example, imagine a customer-support application that has to process a million requests per month. Within those requests, some users may ask for refund policies. Others may want an address extracted from an email. Some might have complicated account problems that need careful analysis. If each request goes to the most powerful model, the company has to pay a premium price even for work that is quite simple for this capable model. You could think of this as hiring a senior software architect to rename files, sort support tickets, and format dates. Sure, the architect can technically do those things. But it would be a waste of the architect’s capability and a case of poor resource management.

Try Crusoe for free with $5 in credits :https://go.bytebytego.com/Crusoe_090826

Model routing is the process of checking an incoming request to decide which model is the best choice for handling it.

With model routing, we don’t write application code that always calls one model blindly. We place a router in front of several models. The router can access a small model, a medium-sized model, and a highly capable one. Its job is to evaluate each request and send it to the most suitable model.

For example, the router might receive a simple classification request and send it to the smallest model. Or the router might receive a request that contains a complicated legal comparison and send it to the most powerful model.

You can think of model routing as load balancing. But it has an important difference. A load balancer normally distributes traffic between largely equivalent servers. However, a model router has to choose between models with vastly different capabilities, costs, and characteristics.

Model routing is also quite different from a mixture-of-experts (MoE) model. In an MoE setup, routing happens internally between parts of a single model. In contrast, application-level model routing happens outside the models. It deals with deciding which model should receive the request and doesn’t deal with the internals of that model.

Consider a powerful model that costs 1 cent per average request. If an application handles a million requests, using that powerful model for everything would cost around $10,000.

Now imagine a smaller model costs only 1/20th as much, while a medium model costs 1/5th as much as the powerful model. After studying the workload, we discover that 85% of requests can be handled by the small model, 10% need the medium model, and just 5% require the powerful model.

In this case, the average cost per request becomes:

(0.85×0.05) + (0.10×0.20) + (0.05×1.00) = 0.1125

This means that a system built with model routing can potentially cost just 11% as much as the system that uses the same powerful model for handling every request. This is almost a 10X reduction in costs.

Even more favourable traffic patterns or price differences could push the savings beyond tenfold. For example, if more than 90% of the workload consists of extraction, classification, formatting, and straightforward summary generation, the expensive model may be needed only occasionally.

Ultimately, the best savings happen when three conditions are met: a large price difference between models, most requests being relatively simpler, and the router being able to identify the simple requests reliably.

The greatest difficulty in model routing is around determining the difficulty level of a request without answering it.

If a request is short, it doesn’t necessarily mean that the request is simple. For example, “Is the contract valid?” contains just 4 words. But to answer this query safely, the model might need legal expertise and extensive context. On the other hand, a long request is not always difficult. A user may have pasted a long document and asked the model to extract every email address. It is conceptually quite straightforward.

Therefore, a good router cannot rely only on message length to determine the difficulty level. It needs to check several signals while making a fair decision.

For example, the model router might consider what kind of task the user is requesting. This is because tasks like classification, extraction, translation, rewriting, and formatting often require less reasoning. However, tasks that involve planning, debugging, mathematical proofs, or comparing conflicting documents require much higher levels of reasoning.

The model router should also consider the risk factor. For example, a medical, legal, financial, or security-related question may be routed to a stronger model even if the query appears simple. This is because the cost of an inaccurate answer matters a lot.

Another signal the model router could use is the amount of overall context. For example, if a model needs to inspect several documents, make sense of a long conversation, or connect different sources, it needs a larger context window or stronger instruction ability.

Lastly, the model router may also need to check the output requirements before selecting the right model. For example, producing a valid JSON object with a few known fields may be an easy task. However, producing a detailed technical design that adheres to a bunch of critical constraints is much harder.

In other words, no single signal is sufficient. A smart model routing approach normally combines several signals to make the right choice.

The most flexible approach to model routing is to use a smaller model to classify the request.

This so-called router model can work on instructions as follows:

The router can then return a small structured result:

Since the routing prompt and the resulting response are quite short, the classification call won’t be too costly. Based on the response, the application then sends the full request to the selected model.

While this approach deals better with natural language rather than coding fixed rules, it can have another cause of error. The smaller router model can misunderstand the request and send difficult work to a less-capable model. This is why production systems often combine model-based classification with fixed safety rules. A specific rule might clearly specify that certain medical or financial queries should always be sent to the strongest model, irrespective of what the router model suggests.

Let us now look at another useful model routing strategy known as model cascading.

In this strategy, we don’t try to predict the difficulty perfectly. Instead, the system first sends the request to a cheaper model. It then checks whether the answer appears good enough. If the answer fails the check, the system sends the request to a stronger model.

This approach works quite well when answers can be checked automatically. For example, let’s say the application asks the model to extract a date, customer ID, and total amount from an invoice. The program can then verify that all required fields exist, the date is valid, and the amount is numeric. If the small model has produced malformed data, the second attempt goes to the powerful model.

We get similar opportunities in the case of code generation. The application can run tests against the generated code. If the tests pass, it accepts the cheaper model’s answer. If they fail, it can escalate the task to the more capable model.

However, cascading gets difficult when quality judgement is subjective. There may be no simple automated test to find out if a business strategy is useful or whether an explanation is actually clear. In those cases, the application may use a separate evaluator model. However, such a model would have its own cost and can also make mistakes.

Lastly, the cascade process must be designed carefully because failed attempts also consume money and time. If most attempts made by the small model end up in failure, the application only ends up paying for both the small and the capable model. Routing ends up making the system slower and more expensive.

In semantic routing, we choose the model based on the meaning of the request rather than specific keywords.

For example, consider an application that has specialized models or prompts for billing, technical support, product recommendations, and account security. However, users may describe the same billing problem in many different ways:

The same order appears twice on my card.

A typical keyword-based system might not be able to support many of these variations. But a semantic router converts the request into an embedding. For reference, an embedding is a numerical representation of a request’s meaning.

The router can then compare the embedding with examples of known request categories. If the request is close to billing examples, it goes to the billing model. If it resembles account-security examples, it goes to the security model.