How to pick the right AI model for each task
Newer and cheaper per token doesn't mean cheaper per task. How to match each kind of work to the right model, and how to prove it.
In July, a developer relations team at Microsoft tested what happens when you switch an agent to a newer, cheaper model.
They ran 150 agent tasks in GitHub Copilot Chat on two models: Claude Sonnet 4.6 and the newer Sonnet 5, which costs 33% less per token.
On code upgrade tasks, the newer model cost $2.01 per run. The older one cost $0.55.
The "cheaper" model was 3.7x more expensive, because it used about 10x the tokens to do the same job. On architecture tasks it came out 12% cheaper, but the older model wrote better output, scoring 90% against 78% on their check for following established patterns. Then on the code upgrades, the newer model followed the instructions every time, while the older one got there in 60% of runs.
Same two models. Different winner depending on the task, and a different winner depending on whether you were counting cost or quality.
The right model for a task is the cheapest one that clears your bar on that task, and the only way to find it is to test.
"Just give everyone the best model"
This is the most common default, and it has a real argument behind it:
Engineers cost far more than tokens. If the top model saves anyone ten minutes, it's paid for itself. Why make people think about it?
For a person working through a hard problem, that's often right. But most AI spend doesn't come from one person's hard problem. It comes from the same tasks running thousands of times: agents in CI, scheduled jobs, ticket triage, summaries, extraction. At that volume, a model that costs 5x more per task is 5x more spend for the same result.
And "best" moves around. In the Microsoft test, the newer model was cheaper on architecture work and wrote worse output there.
The opposite default, "use the cheapest model", fails the other way. A cheap model that needs three tries, or a person to fix its output, isn't cheap.
Sort the work by what a mistake costs
Every major provider sells models in tiers. Anthropic has Opus, Sonnet, and Haiku. OpenAI and Google have their own large, mid-sized, and small models. The names change every few months, but the split stays about the same.
Start by sorting your work by how much it costs you when the output is wrong.
| Tier | Good for | A mistake costs |
|---|---|---|
| Top (Opus class) | Multi-step planning, system design, deep research, reviewing business rules, anything a customer reads | A bad plan, a wrong decision, or your brand on a generic answer |
| Middle (Sonnet class) | Everyday coding, writing, analysis, most production workflows | A rerun or a quick edit |
| Small (Haiku class) | High-volume checks, pulling fields out of documents, classification, quick lookups, prototypes | Very little, and it's usually easy to catch |
The middle tier is the safest default for most work. Move up when mistakes are expensive. Move down when the task is narrow, runs often, and is easy to check.
This table gets you a starting point. It doesn't tell you what works for your tasks. For that, you test.
Start high, then step down
Testing a model doesn't need a research team. It needs a handful of real examples and a clear idea of what a good result looks like.
- Set the bar with the top model. Run 10 to 20 real examples of the task through the most capable model you have. Its output becomes your reference for what good looks like.
- Write down what "good" means. Did it follow the instructions? Is the answer correct? Does it match your format? A few concrete checks beat a gut feeling.
- Step down a tier and run the same examples. Then step down again.
- Run each example more than once. The same model can use very different amounts of tokens on the same task. Microsoft ran every scenario 5 times per model.
- Compare cost per finished task. Price per token tells you little. Total the cost of the runs that met your bar, including retries.
- Pick the cheapest model that passes. Then make it the default for that task.
Repeat it when a new model comes out. Microsoft's team put it well:
A model upgrade is a hypothesis, that newer means better for your specific tasks.
Make the good choice the default
Testing tells you which model to use. The savings only show up when people and systems actually use it without having to remember.
Split big jobs between models
For long, complex work, one model doesn't have to do everything. A capable model can plan the work and hand smaller pieces to cheaper models.
Anthropic built its research feature this way. A lead agent on Claude Opus 4 directing subagents on Claude Sonnet 4 beat Opus 4 working alone by 90.2% on their internal research eval. That setup also used about 15x the tokens of a normal chat, so Anthropic only recommends it for valuable tasks that can run in parallel. They note that most coding tasks don't fit.
Splitting work saves money when the subtasks are simple and numerous. It costs money when they aren't.
Route requests automatically
When a system sends thousands of requests a day, you can route each one by how hard it looks. Easy requests go to a small model, hard ones go to a large one.
In one open research test, routing between a large model and a much smaller one cut costs by over 85% on a standard benchmark, while keeping 95% of the large model's quality. The routing can sit in a gateway between your systems and the model providers, so your application code doesn't have to change.
Give the model less to figure out
A lot of tokens go to the model working out what you mean: which customer, which table, which version of a document. Clear definitions, clean data, and good documentation shrink that work on every model you use.
The Microsoft test showed the limit of model choice here. On the code upgrades, neither model got the configuration right, because the steps weren't written down anywhere.
Spending more on a newer model doesn't fix a content gap.
Set defaults for people, too
Pipelines are easy to configure once. People are harder. Most AI tools start everyone on whatever model the tool picks.
With Tokenize model control, you choose which models each team can use and which one new Claude Code and Codex sessions start on. Research can get the top model while routine work starts on a cheaper one. For production agents, Tokenize watches spend by API key and flags agents running on a pricier model than they need.
Questions to ask before you pick
How much does a wrong answer cost?
If the output goes to a customer or into a decision, start at the top tier and only step down with evidence. If a person reviews it right away, the middle tier is usually enough.
How often does it run?
A task one person does a few times a week isn't worth testing. Give them the best model and move on. A task that runs thousands of times a month is where an hour of testing pays back.
Can you check the output automatically?
If you can tell a good result from a bad one with a simple check, a small model is worth trying, because its misses are cheap to catch and rerun.
When did you last test it?
If the answer is "when we set it up", run your examples again on the newest models. The cheapest model that passes may have changed, in either direction.
Spend where it counts
Picking models this way takes some effort up front. You find a handful of tasks where the top model earns its price, a lot of work where the middle tier is plenty, and some high-volume jobs a small model handles fine. Then you set it as the default and stop thinking about it until the next model ships.
If you'd like to see which models your teams use today and what they're spending on each, book a demo and we'll show you on your own data.
Sources: Waldek Mastykarz, Microsoft, Not all model upgrades are upgrades, July 6, 2026. Anthropic, How we built our multi-agent research system, June 13, 2025. LMSYS, RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing, July 1, 2024.