Back to blog
EvalsSep 29, 2026

The Model Router Is Not the Real Problem

Model routing is only as good as the decision behind it, and that decision depends on the context the system can see.

7 min read

Not every AI task needs the most capable model available. Some tasks are straightforward. Others require deeper reasoning, more context, or a stronger model. Using the same model for both can mean paying more and waiting longer than necessary.

That makes model routing an attractive idea: send simpler tasks to smaller, faster models and reserve the more capable ones for work that actually needs them. The logic is sound.

But there is a question that comes before routing:

THE QUESTION THAT COMES FIRST

Who decides which task is simple?

Because before a system can route a task to the right model, it has to understand enough about the task to make that decision. And that is the harder problem.

The Rise of Model Routing

AI systems increasingly have a choice. A small model may be sufficient for classification, extraction, formatting, or other predictable tasks. A more capable model may be needed when the work involves deeper reasoning, ambiguity, or several pieces of information that need to be understood together.

Using the strongest model for everything is therefore not always the most sensible architecture. If a smaller model can complete a task reliably, there is little reason to pay the additional cost and latency of a larger one.

Task
Routerclassify complexity
Simplesmaller model
Complexstronger model

The router sits in front and makes the choice. At first, this seems like an optimization problem: identify the complexity of the request, choose the appropriate model, and send the work on its way.

But there is an assumption hidden inside that architecture:

A HIDDEN ASSUMPTION

The router already knows how difficult the task is.

And in many real systems, that is not obvious from the request itself.

Routing Is Itself a Decision

Consider a software engineering task:

"Update the validation logic for this field."

From the instruction alone, it looks small. A few lines of code may need to change. A smaller model might appear perfectly capable of handling it.

But now add some context. That field is consumed by three other services. One of them expects the existing format. A condition inside the validation logic reflects a business rule that was introduced years ago but never formally captured outside the code. Another downstream process behaves differently when that condition is triggered.

The task is still:

"Update the validation logic for this field."

Nothing about the wording changed. What changed is our understanding of what the task actually involves.

Request alone
Update the validation logic
classified: simple
Request + system context
Update the validation logic
3 downstream services
Undocumented business rule
classified: not simple

A router looking only at the request might classify it as simple. A router with a better understanding of the surrounding system might reach a very different conclusion.

The Real Problem Is Decision-Making

Once you look at model routing this way, choosing a model starts to look like only one of several decisions an AI system may need to make.

Before beginning the work, it might need to determine whether it has enough information. Perhaps more context needs to be retrieved. Perhaps what initially looked like one task needs to be broken into several. Perhaps a tool needs to be called before any model should attempt the task.

And after the work is completed, the system may still need to decide whether the result is sufficient, whether another step is necessary, or whether it should try a different approach.

The simple picture
Task
Model
Result
selection = the process
What actually happens
Task
Understand
Decide
Act
Check
selection is one step inside

Model selection sits somewhere inside that process. It is not the process itself.

This matters because it is easy to focus on making the router faster, cheaper, or smarter while overlooking the more fundamental question:

What is the router making its decision from?

A sophisticated routing mechanism will not compensate for a poor understanding of the task. It may simply make the wrong decision more efficiently.

Decisions Depend on Context

Take a few seemingly simple software tasks:

  • "Remove this condition."
  • "Change this API."
  • "Refactor this method."
  • "Update this database field."

Viewed locally, each could be a small piece of work. But enterprise software is rarely local. A condition may represent a policy that has been encoded into the system over years of changes. An API may have consumers that are not obvious from the file being edited. A database field may participate in reporting, reconciliation, or another workflow several steps away.

The difficulty of the task is therefore not always contained in the task description. It exists partly in the relationships surrounding it.

That means an AI system deciding how to handle the task may need to understand more than the immediate file. It may need relationships between components, dependencies, architecture, database structures, application behaviour, and business rules that live inside the system.

We ran into a related distinction when we measured adherence and accuracy in AI-generated answers. An answer can be factually correct and still fail if it does not respond to the specific context in which the question was asked. The same principle applies here. A routing decision can appear reasonable in isolation and still be wrong for the system it is operating in.

MORE CONTEXT IS NOT AUTOMATICALLY BETTER CONTEXT

The goal is not to provide every piece of available information for every request. The system needs the right context to understand what the task means inside this particular application.

This is one reason we treat codebase context as a persistent part of the engineering environment at Vyazen. An AI system should not have to rediscover the relationships, dependencies, and business rules of an application every time it receives a task. The better it understands the system, the better foundation it has for deciding what should happen next.

Model routing is only as good as the decision behind it, and that decision is only as good as the context it sees.

The Model Can Change. The Context Shouldn't Have To.

There is another reason to separate context from the model making use of it.

Models are going to keep changing. A task that requires a large model today may be handled reliably by a much smaller one later. A new model may become better at software engineering. Another may offer similar capability with lower latency or cost.

An enterprise may also use several models at the same time. That is not a problem. The model should be replaceable as better options become available. What should not have to be recreated every time is the organization's understanding of its own software.

Its architecture does not reset when a new model is released. Its dependencies do not disappear. The reasons behind years of engineering decisions remain. The business behaviour embedded across the application remains.

That knowledge belongs to the organization, not to the model currently operating on it. This creates a useful separation. The models can evolve quickly. The context they operate on should remain durable.

And if routing changes because a newer model becomes faster, cheaper, or more capable, the system should still be able to make that routing decision from the same underlying understanding of the application.

Routing can change with the models. Context has to outlive them.

The Decision Comes Before the Route

Model routing solves a real problem. As AI systems gain access to models with different costs, speeds, and capabilities, using the same model for every task makes less and less sense.

But the difficult part does not begin with choosing between Model A and Model B. It begins one step earlier. The system has to understand what it is looking at well enough to decide where the work should go. For enterprise software, that understanding rarely comes from the prompt alone.

So instead of beginning with:

Which model should handle this task?

there is a more fundamental question to ask:

What does the system need to understand before it can decide?

As models become cheaper, faster, and more capable, the act of routing between them may become relatively easy. Making the right routing decision may remain the harder part. And that decision will depend on how well the system understands the context in which it is operating.