Skip to main content
Two requests. Same app, same code, same API key. Watch where they land.
Nothing in the app changed between those two calls. No flag, no config, no second API key. One prompt was cheap to answer and one was not, and the router noticed. That is the whole product. Everything below is just how it works.

What happened in those few milliseconds

1. It read the prompt. A scorer looks at 16 signals: length, whether code is present, reasoning markers, imperatives, formatting requests, and so on. This is pattern matching, not a model. It runs inside the process in about 3ms and never leaves your server. 2. It graded the difficulty. The score maps onto five tiers, and the scorer reports how sure it is.
3. If it was unsure, it got a second opinion. Below 0.60 confidence the decision moves to an embedding model that compares your prompt by meaning against 5,090 pre-labeled examples. The closest matches vote. That is what happened to the planning request above: the rules landed on STANDARD but only at 0.50 confidence, so the embedding path confirmed it. 4. It picked the model. Tier plus topic points at a model in the catalog. Your routing profile can then override: never below this tier, never above that price, must support vision. 5. It called the model, with a net underneath. If that provider fails, the request falls to the next model instead of to your user. 6. It told you everything. Every response carries the decision in its headers.
Routing never calls a language model to decide which language model to use. The whole decision is local, which is why it costs milliseconds instead of a round trip, and why the same prompt always produces the same answer.

The other way to build this

Some routers decide by watching the market. They sort your prompt into a task type, look up which models the community spends the most on for that type, and send it there. It is a reasonable idea, and it has one genuine advantage: the list updates itself as the market moves. But it answers a different question than the one you asked. Three of those deserve a sentence. Spending is not the same as suitability. Ranking by spend tilts toward expensive models, because expensive models produce more spend per request. Popular is not the same as right for your prompt. Topic is not difficulty. A one-line typo fix and a week-long architectural rewrite are both “coding”. Sorting by topic cannot separate them. Grading by difficulty is the entire point. Reproducibility is a feature. If you run evals against your own app, a router that quietly changes its mind between runs makes your results unreadable. Ours is deterministic by design.

Where market routers are better

Being straight about this matters more than winning the comparison.
  • They never go stale. A new model gets popular and their ranking picks it up within days. Our catalog is curated, which means a human reviews it.
  • They carry more models. Hundreds against our 17.
  • They add no routing fee. Routor adds 10% on top of provider cost.
We think a reviewed catalog is worth more than an exhaustive one, because someone has actually checked what each model is good at and written down where it falls short. But that is a tradeoff, not a free win, and you should pick the side that fits your work.
Want to see a decision without spending anything? Send a prompt to POST /v1/routing/debug. It runs the real pipeline, returns the full decision, and never calls a model or bills you.