One model for everything is the expensive default
Sound familiar?
Every request goes to the same top-tier model, including "thanks!", "what are your opening hours?" and "reset my password". The hard questions get great answers. So do the easy ones, at the same price.
The price spread between AI model tiers is now enormous. OpenAI's GPT-6 Astra lists at $10 per million input tokens, and GPT-6 Luna at $0.10: a 100x difference within one provider. Most real traffic is a mix of easy and hard requests, so running everything through one model means overpaying for the easy ones or under-serving the hard ones.
Model routing fixes this by choosing a model per request, based on how hard the request is, how fast the answer is needed, and what it is worth. In 2026 it has moved from research idea to standard infrastructure: AI gateways such as OpenRouter, LiteLLM and Portkey offer routing as configuration rather than custom code.
Four ways to route
| Strategy | How it works | Good for |
|---|---|---|
| Rules | Fixed logic by feature or request type, e.g. summaries go to a small model and contract review to a large one | Products whose AI features differ clearly in difficulty |
| Classifier | A small, cheap model reads each request and labels it easy, medium or hard before it is sent on | Mixed traffic in one feature, such as a support chat |
| Cascade | Try the cheap model first and escalate to a bigger one when its answer fails a check | Tasks with an automatic way to check the answer, such as structured data extraction |
| Effort levels | Keep one model but ask it to think less or more per request | Teams that want one provider and one set of prompts |
Effort levels are now built into the major models. OpenAI's GPT-6 Astra accepts reasoning effort settings from low to max, Anthropic's Claude Fable 5.1 takes an effort parameter, and DeepSeek offers low, high and max reasoning effort on V4. Lower effort means fewer output tokens, and output tokens are the expensive ones.
What routing saves: a worked example
Take the support chat from our AI API pricing guide: 5,000 conversations a day, 2.7 billion input and 225 million output tokens a month. Suppose a classifier finds that 70% of turns are simple, 25% are standard and 5% are genuinely hard. Each class goes to a matching OpenAI tier. The classifier itself runs on GPT-6 Luna and adds about $47 a month.
| Everything on GPT-6 Astra | 38250 USD |
|---|---|
| Everything on GPT-6 Sol | 7650 USD |
| Routed: 70% Luna, 25% Sol, 5% Astra | 4140 USD |
Source: Tricolens calculation from OpenAI list prices, Sep 28, 2026. No caching or batch discounts; routed figure includes the classifier.
The routed setup costs 46% less than running everything on the mid-tier model, and it still sends the hardest 5% of questions to the most capable model. Compared with running everything on the top tier, it costs about a ninth as much.
Want to know what routing would save on your traffic?
Get a free consultationHow to build a router
Put one interface in front of every model
All AI calls go through one layer in your code, so adding or swapping a model is a configuration change, not a rewrite.
Collect and label real requests
Take a few hundred real requests and mark which ones a small model handles well. This becomes both your routing data and your test set.
Start with rules, add a classifier later
Route whole features by rule first. Add a classifier only where one feature has clearly mixed difficulty.
Log every decision
Record which model handled each request, its cost and its response time, so you can see savings and trace any quality complaint to its route.
Add fallbacks
If a provider times out or has an outage, retry on another model instead of showing an error.
Pitfalls to avoid
- Routing without a test set: you will not notice when a cheap route starts giving worse answers
- Cascades on tasks you cannot check automatically: you pay for the cheap attempt and still cannot tell if it failed
- Assuming prompts are portable: the same instructions can behave differently on another provider's model
- Ignoring data terms: every provider in the route must meet your privacy and residency rules
- Over-engineering early: rules plus one fallback cover most products at first
The bigger benefit: freedom to switch
A router makes cost savings visible, but its long-term value is flexibility. New models arrive every few months. See our September 2026 model comparison and the open-weight options. With routing in place, trying a new model means adding a route and running your test set, not rebuilding the feature.
If you want a router designed and built for your product, see our AI integration services.
Key takeaway
Match the model to the request, not the product. Start with simple rules, keep a test set, log every route, and you get lower bills and the freedom to adopt the next model the week it ships.
Frequently Asked Questions
Written by
Hiren Patel
AI & Engineering
