Here is a problem every per prompt router has, and most of them lose to.
You are twenty messages into debugging a race condition. You have pasted stack traces, the model has written three patches, and you type:
Scored on its own, that prompt is four trivial words. Scored honestly, it inherits everything above it. A router that only looks at the current message sends it to the cheapest model on the list, and you get a useless answer at the worst possible moment.
Routor reads the conversation, not just the message.
The history floor
Before choosing a model, Routor scans the recent conversation and may set a floor: a tier the request cannot fall below, no matter how simple this particular message looks.
The floor is applied as max(scored tier, floor). It can only raise a request, never lower one.
What raises it
Error detection is not keyword spotting on the word “error”. It matches real failure shapes across languages: tracebacks, panics, segfaults, borrow checker complaints, npm ERR!, ModuleNotFoundError, NullPointerException, cannot read propert..., index out of bounds.
What it reads
The last 6 conversational messages, capped at 4,000 characters each, system and tool messages excluded. The current prompt is deliberately left out, since it has already been scored on its own merits. Threads shorter than two turns are skipped entirely.
History alone never forces COMPLEX. The floor is capped at STANDARD, so a long conversation can keep you in capable territory but cannot by itself escalate you to the most expensive tier. Reaching COMPLEX requires the current prompt to earn it.
Saying “thanks” does not cost you money
The obvious failure of a floor is that it makes everything expensive forever. Type “hi” in a coding session and you would pay STANDARD rates to be told hello.
So trivial messages are exempt. Anything at or under 60 characters that is recognisably an acknowledgement skips the floor and routes on its own merits:
A greeting in the middle of a long debugging session still goes to a cheap model. The floor comes back the moment you say something real.
Category carry forward
The floor decides how capable the model must be. A second, quieter mechanism decides what kind of model it should be.
When history shows a coding session, Routor remembers the task category and carries it forward through thin follow ups. A vague “now the tests too” stays in the coding category instead of drifting into general chat, because the session already established what you are doing. An error trace marks the session as debugging specifically.
Seeing it happen
When the floor is applied, the response says so:
The reasoning string names each signal and the turn it came from, so you can tell whether it was the stack trace you pasted or the code block the model wrote that held the floor up. When no floor applies, the header is absent.
See Router Metadata for the full header list.
Why this matters more for agents
Agent loops are where this pays for itself. An agent sends dozens of short, context free instructions: “continue”, “fix it”, “now the tests”. Every one of those, judged alone, looks trivial. Judged in context, most of them are the hard part of a coding task.
Without conversation memory, a per prompt router quietly downgrades an agent at exactly the wrong moments, and the loop produces worse work while appearing to save money.