> ## Documentation Index
> Fetch the complete documentation index at: https://docs.routor.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Why Not Just Auto?

> How Routor decides which model answers, and why that differs from routers that follow the market.

Two requests. Same app, same code, same API key. Watch where they land.

```text theme={null}
"Summarize my todo list into three cards"
   tier SIMPLE      decided by rules        3.6ms     google/gemma-4-31b

"I have 14 tasks across 3 projects with hard deadlines and
 dependencies. Two need legal review, one is blocked on a
 vendor. Work out the critical path and tell me what to drop."
   tier STANDARD    decided by embedding    ~9ms      zai/glm-5.2
```

Nothing in the app changed between those two calls. No flag, no config, no second API key. One prompt was cheap to answer and one was not, and the router noticed.

That is the whole product. Everything below is just how it works.

***

## What happened in those few milliseconds

**1. It read the prompt.** A scorer looks at 16 signals: length, whether code is present, reasoning markers, imperatives, formatting requests, and so on. This is pattern matching, not a model. It runs inside the process in about 3ms and never leaves your server.

**2. It graded the difficulty.** The score maps onto five tiers, and the scorer reports how sure it is.

```text theme={null}
NANO  ->  SIMPLE  ->  LIGHT  ->  STANDARD  ->  COMPLEX
```

**3. If it was unsure, it got a second opinion.** Below 0.60 confidence the decision moves to an embedding model that compares your prompt *by meaning* against 5,090 pre-labeled examples. The closest matches vote. That is what happened to the planning request above: the rules landed on STANDARD but only at 0.50 confidence, so the embedding path confirmed it.

**4. It picked the model.** Tier plus topic points at a model in the catalog. Your routing profile can then override: never below this tier, never above that price, must support vision.

**5. It called the model, with a net underneath.** If that provider fails, the request falls to the next model instead of to your user.

**6. It told you everything.** Every response carries the decision in its headers.

```text theme={null}
X-Routor-Model       google/gemma-4-31b
X-Routor-Tier        SIMPLE
X-Routor-Method      rules
X-Routor-Confidence  0.65
X-Routor-Savings     97%
```

<Note>
  Routing never calls a language model to decide which language model to use. The whole decision is local, which is why it costs milliseconds instead of a round trip, and why the same prompt always produces the same answer.
</Note>

```mermaid theme={null}
flowchart LR
    P(["Prompt"]):::step --> R["Score 16 signals\nlocal, ~3ms"]:::step
    R --> C{"Confident?"}:::step
    C -->|"Yes"| T["Tier + category"]:::positive
    C -->|"No"| E["Embedding vote\n5,090 examples"]:::warning
    E --> T
    T --> M(["Cheapest model that\nclears the bar"]):::positive

    classDef step fill:#0d0d0d,stroke:#9ca3af,color:#d1d5db
    classDef positive fill:#0d0d0d,stroke:#f3f4f6,color:#ffffff,stroke-width:2px
    classDef warning fill:#0d0d0d,stroke:#6b7280,color:#9ca3af
```

***

## The other way to build this

Some routers decide by watching the market. They sort your prompt into a task type, look up which models the community spends the most on for that type, and send it there. It is a reasonable idea, and it has one genuine advantage: the list updates itself as the market moves.

But it answers a different question than the one you asked.

| | Market routers | Routor |
| - | - | - |
| What it measures | What people spend on this kind of task | How hard *this* request is |
| Granularity | About 30 task types | 5 difficulty tiers across 17 categories |
| Same prompt next week | May route elsewhere as spend shifts | Same answer, every time |
| Cost control | A band you set once for everything | Judged per request |
| What you can see | Which model answered | Model, tier, confidence, method, savings |

Three of those deserve a sentence.

**Spending is not the same as suitability.** Ranking by spend tilts toward expensive models, because expensive models produce more spend per request. Popular is not the same as right for your prompt.

**Topic is not difficulty.** A one-line typo fix and a week-long architectural rewrite are both "coding". Sorting by topic cannot separate them. Grading by difficulty is the entire point.

**Reproducibility is a feature.** If you run evals against your own app, a router that quietly changes its mind between runs makes your results unreadable. Ours is deterministic by design.

***

## Where market routers are better

Being straight about this matters more than winning the comparison.

* **They never go stale.** A new model gets popular and their ranking picks it up within days. Our catalog is curated, which means a human reviews it.
* **They carry more models.** Hundreds against our 17.
* **They add no routing fee.** Routor adds 10% on top of provider cost.

We think a reviewed catalog is worth more than an exhaustive one, because someone has actually checked what each model is good at and written down where it falls short. But that is a tradeoff, not a free win, and you should pick the side that fits your work.

<Note>
  Want to see a decision without spending anything? Send a prompt to `POST /v1/routing/debug`. It runs the real pipeline, returns the full decision, and never calls a model or bills you.
</Note>
