Choosing a model for an in-app AI agent, and routing the turns that don't need it
The short answer
Classify each message before loading anything expensive. Greetings, thanks and out-of-scope requests go to a small model with no tools and a short prompt; real work goes to the capable model with the full tool schema. Make the classifier fail toward the expensive path, because misrouting real work breaks the product while misrouting a greeting costs a fraction of a cent. Measure the split for a week before building it, and keep every model setting in one object that changes together.
Choosing a model for an in-app agent usually starts as a benchmark question: which one calls tools most accurately? That matters. But in production the bigger cost decision is often not which model you picked. It is how many turns reach it that never needed it.
A greeting costs as much as real work
One production agent handled small talk correctly. Its system prompt said that greetings and thanks are never a request to do anything, so call no tools and reply briefly. The model obeyed. It still cost almost the same as a real request.
The reason is order of operations. Every message was sent to the capable model with the full system prompt and every tool definition attached, and a logged turn came to more than 90,000 prompt tokens. The instruction about greetings sat inside that payload. By the time the model read it, the cost had been paid. The part that varies between “hi” and a complex request is the user’s message, and “hi” is two tokens.
Prompt engineering cannot fix a routing problem. A prompt changes what the model does. It cannot change what you sent it.
Classify before you load
Put a cheap step in front. A small model with no tools and a short prompt reads the message and sorts it:
- Greeting, thanks, small talk: the small model replies, no tools, minimal prompt.
- Out of scope: the small model gives a one-line redirect.
- Real work: the capable model, full prompt, full tools.
Make it fail toward the expensive path
The two ways to be wrong are not symmetrical. Send a greeting to the capable model and you spend a fraction of a cent. Send real work to the small model and the product fails in front of a user. So bias the classifier hard: anything uncertain counts as real work.
And measure before you build. Log the classification for a week without acting on it. If greetings are three percent of traffic, the router is not worth its complexity yet. The agent above never measured, which is why nobody noticed the expensive model was answering “ok”.
Scope is a routing question too
An agent with no stated scope will help with anything, and for a B2B product that means your customers’ users can spend your inference on their homework. Put a short scope block near the top of the prompt, what it is for and a one-sentence redirect for everything else, and name the trap: a request phrased in your product’s vocabulary is still out of scope if the task is. Routing makes the out-of-scope reply cheap as well as correct.
Model swaps break things quietly
Every one of these happened in that same production agent during a model change:
- The price table stayed on the old model. Cost reporting was wrong and nothing errored.
- Sampling settings stopped applying. The new model supported none of the temperature and penalty parameters still being sent. They were silently ignored while the code read as though they worked.
- A provider pin outlived its model. It pointed at a provider that did not serve the new model at all.
- Settings were exported one by one. A swap needed four call-site edits, and three were missed.
One fix covers all four: keep the model id, its prices and its supported parameters in a single object that call sites spread, so changing the model means changing one thing. And check the provider’s parameter list on every swap rather than assuming.
Then choose on the cases that separate models
When you do benchmark, the obvious cases tell you little, because every capable model passes them. Score the edges: requests that should not call a tool, honesty when a tool fails, retries after a user declines, instructions hidden in data. The last two are covered in testing an agent without real data.
Common questions
Why is saying hi to an agent expensive?
Because the full system prompt and every tool definition are sent before the model reads the message. One production agent logged a turn at over 90,000 prompt tokens, and a greeting is barely cheaper, since the part that varies is the user's two words.
Can't the prompt just tell the model to keep greetings short?
It can, and it will, but the cost was already paid before the model read that instruction. Prompting changes what the model does. It cannot change what you sent it. Only routing does that.
What breaks when you switch models?
Quietly, a lot. Sampling parameters the new model does not support get ignored while the code still reads as though they apply. A price table left on the old model makes cost reporting wrong. A provider pin can outlive the model it was chosen for. Keep the model id, its prices and its parameters in one object edited together.
Keep reading
- How to test an AI agent that takes actions, without touching real data
Run it for real against development data, the five cases worth deliberately breaking, and why a timer on real execution backfires.
- Why your embedded script silently does nothing under a Content Security Policy
The failure that cannot report itself, the second directive everybody forgets, and why strict-dynamic makes your allowlist irrelevant.
- The Next.js build, not the app, is what runs the server out of memory
Why a build takes down a small server when the app it produces runs fine, and the one-line cgroup fix.
Verb is this, built. An AI assistant you embed in your SaaS with one script tag: it calls your own API as the signed-in user, confirms before it changes anything, and logs every action. Free to build and test.