Mira Solène 7 min readChoosing one model for an entire AI product is simple, but rarely durable. Extraction, drafting, classification, code generation, and high-stakes reasoning do not impose the same demands. A model that excels at one may be unnecessarily slow, expensive, or unreliable at another.
A model router turns that mismatch into an architectural advantage. It examines each request, chooses an appropriate model or workflow, and records whether the choice worked. The concrete outcome of this guide is a first production-ready routing policy: narrow enough to evaluate, explicit enough to audit, and instrumented enough to improve.
Step 1: Define the Routing Unit
Begin by deciding what the router will classify. Do not route an entire conversation merely because that is the object your application receives. One conversation may contain several distinct jobs.
The useful routing unit is usually an atomic task: summarize a document, extract invoice fields, answer a policy question, produce code, or rewrite text under constraints. Give each task a stable name and define its expected output.
| Task | Expected output | Primary risk | Useful signal |
|---|---|---|---|
| Invoice extraction | Validated fields | Incorrect values | Schema and document type |
| Support reply | Grounded response | Invented policy | Topic and retrieval confidence |
| Code repair | Passing patch | Regression | Language, repository context, test result |
| Marketing rewrite | On-brand copy | Tone mismatch | Length, audience, style constraints |
Common mistake: creating categories such as “easy” and “hard.” These labels are subjective and provide no operational guidance. A short compliance question may be riskier than a long rewrite. Name the work, the required output, and the failure that matters.
Step 2: Establish a Baseline Before Routing
Select one capable model as the baseline and run it across a representative evaluation set. The set should contain real task shapes: short and long inputs, ordinary and unusual requests, malformed data, ambiguous instructions, and adversarial phrasing where relevant.
Record four dimensions for every example:
- Quality: Did the output satisfy the task-specific rubric?
- Latency: How long did the user or downstream process wait?
- Resource use: How much input, output, retrieval, and tool activity occurred?
- Failure severity: Was the error cosmetic, recoverable, or consequential?
For extraction, quality may mean exact field agreement plus schema validity. For support, it may require policy grounding and correct escalation. For code, executable tests are stronger than a reviewer’s impression.
Common mistake: comparing models with a universal score. Routing requires conditional evidence. You need to know which model succeeds on which task under which constraints, not which model wins an abstract contest.
Step 3: Build a Small Candidate Portfolio
A useful first portfolio has distinct roles rather than many nearly interchangeable models. Consider a fast default model, a stronger reasoning model, and a specialist model or deterministic workflow for a narrow task. The specialist might be an OCR pipeline, a classifier, or code execution paired with a model.
Run every candidate against the same evaluation cases. Then construct a capability matrix. Avoid forcing a single ranking; mark where each candidate is acceptable, unacceptable, or only acceptable with verification.
Suppose the fast model handles routine rewrites and simple classification reliably, while the stronger model performs better on ambiguous policy questions. Invoice extraction may work best through OCR, schema-constrained generation, and validation rather than through either model alone. The router is therefore selecting systems, not merely model names.
Common mistake: adding candidates because they are newly released. Every additional route creates evaluation, integration, fallback, and monitoring work. Add a candidate only when it occupies a meaningfully different point in the quality, latency, risk, or deployment landscape.
Step 4: Convert Evidence Into Explicit Rules
Start with deterministic routing rules. They are inspectable, reproducible, and easier to debug than a learned router. Useful inputs include task type, input length, language, required tools, data sensitivity, customer tier, time budget, and failure severity.
A first policy might read:
- If the task is invoice extraction, use the extraction workflow.
- If the task is a routine rewrite below the context threshold, use the fast model.
- If the request concerns regulated policy, use retrieval plus the stronger model.
- If required context is missing, ask a clarifying question rather than selecting a model.
- If no rule matches, use the baseline model and mark the event for review.
Notice that abstention is a route. So are human review and deterministic software. A router should choose the safest processing path, not ensure that a language model answers everything.
Common mistake: routing primarily by prompt length. Length affects context capacity and cost, but it says little about ambiguity or consequence. A compact request can demand substantial reasoning; a long document can require straightforward extraction.
Step 5: Add Confidence Without Trusting Self-Confidence
Rules eventually encounter uncertain boundaries. A lightweight classifier can identify task type, sensitivity, or required capability, but its confidence must be calibrated against observed outcomes.
Use external signals where possible: schema validation, retrieval coverage, test execution, citation presence, agreement between independent methods, or a task-specific verifier. Model declarations such as “I am highly confident” are not sufficient control signals.
Define three bands for each route. In the accepted band, proceed normally. In the review band, invoke verification or a stronger system. In the rejection band, ask for missing information or escalate. Set these boundaries from evaluation errors and their consequences rather than intuition.
Common mistake: treating a confidence score as a universal probability of correctness. A classifier score concerns its own label space. It does not automatically measure whether the final answer is factual, safe, or useful.
Step 6: Design Fallbacks as a State Machine
Fallbacks should be finite and purposeful. Without limits, systems can bounce among models, repeat tool calls, and accumulate delay while producing no better answer.
Represent processing as explicit states: received, classified, executing, verifying, escalated, and completed. Permit only deliberate transitions. For example, a failed schema check may trigger one repair attempt; a second failure moves to review rather than another generation.
Preserve the original request, route decision, intermediate outputs, validation results, and fallback reason. This provenance makes failures diagnosable and prevents a fallback model from silently operating on corrupted intermediate data.
Common mistake: defining fallback as “send it to the largest model.” Some failures arise from absent data, broken tools, unsupported file formats, or contradictory instructions. More model capacity cannot repair a missing prerequisite.
Step 7: Deploy in Shadow Mode
Before allowing the router to control production responses, let it make decisions silently beside the existing baseline. For each live request, log the proposed route while the baseline continues serving users. On a sampled subset, execute the proposed route and compare outcomes offline.
Examine disagreements rather than averages alone. Where did the router choose a weaker path? Where did it invoke an expensive path without improving the result? Which requests failed to match any category? These cases reveal flaws in the taxonomy and policy.
Then introduce routing gradually. Begin with a low-risk, easily verified task such as formatting or structured classification. Keep consequential routes behind review until their failure modes are understood.
Common mistake: evaluating only successful requests. Timeouts, refusals, malformed outputs, empty retrieval, and tool failures are part of route quality. Excluding them makes fragile paths appear efficient.
Step 8: Create the Learning Loop
Log a compact decision record for every routed task: task label, relevant features, chosen path, policy version, latency, verification result, fallback, user correction, and final disposition. Avoid storing sensitive raw content when derived features or redacted traces will suffice.
Review records by route and failure type. A rising fallback rate may indicate model drift, a changed input distribution, or a broken integration. Frequent user edits may expose quality problems that automated validators miss. Unmatched requests may justify a new task category, but only after a repeated pattern appears.
Version the policy so an outcome can be traced to the rules active at that moment. When changing a rule, replay historical cases before release. This converts routing from ad hoc orchestration into an evidence-bearing control layer.
Common mistake: optimizing cost before defining acceptable quality. The correct sequence is to establish a quality floor, identify every route that clears it, then choose among those routes using latency, resource use, privacy, or resilience.
The Concrete Outcome: A Minimum Viable Router
Your first router does not need machine-learned optimization. It needs a task taxonomy, a shared evaluation set, a small candidate portfolio, explicit selection rules, bounded fallbacks, and decision logging.
Its governing principle is precise: use the least burdensome system that reliably satisfies the task’s quality and risk requirements. That principle avoids both extremes—sending trivial work to the most capable model and sacrificing consequential work for superficial efficiency.
Once the router has accumulated trustworthy outcome data, more adaptive methods become possible. Until then, transparent rules are not a primitive substitute. They are the instrument that reveals what a future learned router should actually learn.
This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.
From our own rounds
Measured on The Curator, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 163
- Questions per round
- 1.7
Rate this article
Discussion
Comments are moderated. Read our editorial policy.