MLGuerrillaStart with M1 →
Free · in beta·intermediate·M14·25 min read·Prereq: Classification (M13), whose taxonomy, thresholds and calibration carry over unchanged. LLMOps & Cost (M12) supplies the price table every number here is computed from.

Routing

The capability

What a router is

A router reads a request and picks which handler does the work, before the work happens. The handler might be a model, a piece of plain code, a retrieval pipeline, an agent or a person.

The word "picks" is carrying a lot of weight in that sentence, so it is worth being precise about what the router is picking from. It can only return a handler you already wrote into your code, which means there is no sixth option until somebody edits the code and ships it. That sounds like a restriction, and the restriction is the point, because a closed set of options is what lets you measure every one of them ahead of time. By the time you sit down to write the routing rule you already know what the database lookup costs, what the cheap model costs, what the frontier model costs, and roughly how long each of them takes to answer.

The second thing worth knowing is where the router sits. It runs before the expensive work, so every millisecond it spends gets added to every request, including the ones it ends up sending to the cheapest handler. A router that takes four seconds to make up its mind has spent longer deciding than the handler will spend answering.

And it can be wrong in two directions, which cost quite different amounts. If it sends hard work to a cheap handler, the customer gets a bad answer. If it sends easy work to an expensive one, the customer gets a perfectly good answer and you get a larger bill. Section 7 is mostly about that gap, because tuning a router as though those two mistakes were equally bad is how teams end up worse off than when they started.

One customer request enters a router on the left. The router chooses between five handlers, each with a cost and a latency written beside it: an order lookup in the database at no cost and 40 milliseconds, a small model at $0.0080 a conversation and 1.9 seconds, a frontier model at $0.1443 and 4.2 seconds, an agent with tools at a higher cost and over 10 seconds, and a person at $2.50 and minutes. The order lookup is marked chosen. A note underneath says the router can return one of these five and nothing else, that the cheapest model call and a person are about 300 times apart, and that the router cannot improve any handler.
The router can return one of these five and nothing else. You measured what each costs before writing the rule that chooses.

One thing a router cannot do is improve any of the handlers it calls. All it decides is which one gets the work, so if every option on the list is mediocre, what you get back is mediocre answers slightly more cheaply.

Three different things get called routing

People use the word routing for three different problems, and they tend to get solved by different people in different parts of a system, so it helps to pull them apart before going further.

Model routing is what most of this module is about. Something reads the request and decides whether it goes to the cheap model or the frontier one, the fast one or the thorough one.

Capability routing is a step up from that, because it decides what kind of handler answers at all. That might be a database lookup, a retrieval pipeline, an agent, or a person. Section 8 comes back to it, and at the flower company it turns out to be worth more money than model routing is.

Infrastructure routing decides which provider, region or deployment serves the call. M11-1 built that one already, as the fallback chain and the load shedding that kick in when something breaks.

The first two are choices you make on purpose, before anything has gone wrong. The third is mostly a reaction to a failure, which is why it sits in the rails and this module leaves it alone.

Why anyone bothers

Cost is the reason with the most evidence behind it, and the one section 5 puts numbers on. Most production traffic never needed the frontier model in the first place, so the money spent sending it there was never buying anything.

Latency is the second reason, and it is easier to feel than to measure. A small model answers faster, and somebody watching a cursor blink notices the difference.

Quality is the reason people claim most often and measure least. Some models genuinely are better at particular kinds of work, so a mix can beat any single model, though that argument only holds if somebody ran the comparison.

Availability is the fourth, and it overlaps with what M11-1 already covered. A router that knows two paths can take the second one when the first is down, which is the same fallback with the decision made in advance.

Where it shows up

A few examples. None of them is this module's running example, which starts in section 2.

  • A chat product deciding whether a message needs a reasoning model or a fast one. ChatGPT does this, and section 6 looks at what happens when the provider makes that decision for you.
  • A voice assistant deciding whether "open the calendar" is a command it can run on the device, or whether "what should I plant in November" needs a model. The first is matched against a list of known commands in milliseconds and never leaves the phone. The second is a network call to something large, and the answer comes back as speech.
  • A support agent sending order lookups to a database and policy questions to a model.
  • A coding assistant choosing between a small model for autocomplete and a large one for a refactor.
  • A triage desk of any kind, including one with people on the other end, which is the version that predates all of this by about a century.

Where this starts

Every turn costs the same at the flower company

The support agent we have talked about before runs every turn through GPT-5.6 Sol, the frontier model. It has done this since the first version, because that was the simplest thing to build and nobody has had a reason to change it. The company is made up, and so are its numbers, and the prices are real ones read in September 2026.

Module 12 priced the year. About 60,000 conversations in February, a mean of $0.1443 of model spend each, so about $8,655 of model spend for the month.

Module 7 already knows the shape of that traffic. Someone read 500 tickets and found that 45% are "where's my order" and 15% are delivery changes. M13 then built a classifier that puts a label on every ticket as it arrives, in a few hundred milliseconds, with a score attached.

So 60% of the agent's turns are a tracking link or a change form, the label saying so is already computed, and all 60% go to the frontier model anyway.

The question this module answers is what happens if they stop.

The overlap with M13

A router is a classifier whose labels cost money

Most of M14 is M13 with money attached. Saying exactly what carries over keeps this section short.

What M13 gives M14 for free
M13 builtM14 uses it as
A closed label set with written definitionsThe handlers you wrote into the code
A score per labelThe confidence that decides whether to route or ask for help
An abstain band that hands the middle to a personA band that sends uncertain requests to the safe expensive handler
Per-class precision and recallPer-handler precision and recall
The calibration checkThe same check, and it carries more weight here, since the score now moves money
Drift monitoring on the inputThe same, plus drift in the handlers themselves, which section 9 covers

What is new in M14 comes from the prices, and there are two of those changes worth separating.

The first is that the labels cost money now. A classifier's labels are just names, and getting one wrong is wrong in a fairly uniform way. A router's labels each carry a price, and one handler can cost eighteen times another, so being wrong in one direction is a completely different event from being wrong in the other. Section 7 works out what each direction costs.

The second is bigger, which is that some wrong answers can be undone. That gets its own section next, because it changes the threshold arithmetic M13 spent a whole section on.

The code looks much like M13's, with the handler attached to the answer:

route.py — the seam in front of every handler
HANDLERS = {
    "lookup":  dict(fn=order_lookup,   cost=0.0,    p50_ms=40),
    "cheap":   dict(fn=answer_on_luna, cost=0.0080, p50_ms=1900),
    "frontier":dict(fn=answer_on_sol,  cost=0.1443, p50_ms=4200),
    "person":  dict(fn=hand_off,       cost=2.50,   p50_ms=None),
}

def route(request) -> str:
    """Returns a key of HANDLERS. Runs before any handler does."""
    ...

choice = route(request)
HANDLERS[choice]["fn"](request)

M4 put every model call behind one complete() function, and M13 put every classifier behind one classify(). This is that same move a third time, and it is what makes section 6's ladder practical, since you can swap one router for another without touching any of the handlers underneath it.

Recoverability

Wrong routes you can undo, and wrong routes you can't

A misfiled support ticket stays misfiled until somebody notices, which is why M13 spent so long on getting the label right the first time. A misrouted request is often different, because the answer comes back and you can look at it before anybody else does.

That difference splits routing decisions into two kinds, and which kind you are dealing with changes almost everything downstream.

In the first kind, being wrong is recoverable. You send the work somewhere cheap, check what comes back, and send it somewhere better if the check fails. The whole cost of being wrong is one wasted cheap call and some extra waiting.

In the second kind, the route commits you. The money has moved, the email has gone out, the customer has already read the reply. There is nothing to check after the fact because the consequence has happened, so being wrong costs whatever M7's failure ladder says it costs.

What decides which kind you have is whether a check exists that can separate a good answer from a bad one before anything acts on it. M9 built three of those and they all apply here: the business rules that compare an answer against what your code already knows, the citation check that catches a reference search never returned, and the coverage check that declines when no source covers the question.

A cascade is routing with no router in it

Once the answer is checkable, you can skip the router altogether. Send everything to the cheap handler, run the check on what comes back, and pass the failures up to the expensive one.

cascade.py — try cheap, check, escalate
def answer(request):
    draft = answer_on_luna(request)
    if checks_pass(draft):          # M9's business rules and citation check
        return draft
    return answer_on_sol(request)   # the same request, one tier up

Notice what is missing from that. There is no taxonomy to maintain, no training run, no labelled dataset, and nothing that has to predict anything. The share of requests that escalate falls out of the traffic on its own, and when the traffic gets harder that share goes up on its own, which is exactly the kind of drift section 9 shows a router failing to handle.

The arithmetic is one line too. A cascade costs you the cheap call every time, plus the expensive call on however many requests escalate, so it beats always paying for the expensive model whenever:

“escalation rate < (expensive − cheap) ÷ expensive”

For the flower company's two models that threshold works out at 94.5%, because Luna at $0.0080 a conversation is about eighteen times cheaper than Sol at $0.1443. Which is to say a cascade there wins at almost any escalation rate you are likely to see, since you would need nearly every request to fail its check before the cheap call stopped paying for itself. On a closer pair the margin is much tighter. Claude Haiku 4.5 against Sonnet 5 only breaks even at 50%.

A cascade on 60,000 conversations
PairEscalation rateModel spend a monthAgainst always-expensive
Luna to Sol20%$2,210$6,445 cheaper
Luna to Sol40%$3,941$4,714 cheaper
Luna to Sol94.5%$8,655break-even
Haiku to Sonnet20%$3,029cheaper
Haiku to Sonnet50%$4,328break-even

What you give up in exchange is latency, and only on the requests that escalate, since those end up waiting for two calls one after the other. That trade is usually fine behind a queue where nobody is watching, and usually wrong in front of somebody waiting for a reply.

The numbers

The flat mix, priced

Routing the flower company's traffic by the M13 label sends 60% to Luna and keeps the rest on Sol. With a router that never makes a mistake:

February, three ways
ArrangementModel spend a monthCost per resolved ticketPass rate
Everything on Sol$8,655$0.95197%
Everything on Luna$479$1.00386%
Routed, 60% to Luna$3,750$0.86797% on easy, 97% overall

The middle row is the one worth sitting with for a moment. Moving everything to Luna makes the model line eighteen times cheaper and still leaves you paying more per resolved ticket than you were before. What happens is that the pass rate falls from 97% to 86%, and every conversation the agent fails to resolve turns into a handoff to a person at $2.50, against a model call that costs a fraction of a cent. So the $8,176 you saved on tokens gets spent several times over on support staff. A cheap model is only cheap when it can do the work in front of it.

The third row is where routing earns its place, and it produces two numbers that sound like they contradict each other. Routing saves 57% of model spend and 9% of cost per resolved ticket. Both of those are true, and they are measuring different things. M12 already established why, which is that the model is only about 17% of what a resolved ticket costs once you count the people, so cutting more than half of a small share still leaves you with a small number.

If you take this to a finance team they will want to know which of the two is the real one, and the honest answer is both, depending on what they are asking. The number the industry quotes is always the model-spend one. RouteLLM's authors reported "cost reductions of over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K as compared to using only GPT-4, while still achieving 95% of GPT-4's performance" (LMSYS, 2024). Those are real measurements on public benchmarks, and what they measure is token spend. How much of that reaches your cost per outcome depends entirely on how large a share of it the tokens were in the first place, which is a number nobody else can tell you.

That 85% is worth one more sentence, because it gets quoted constantly and it does not come from where people think. It is from the LMSYS blog post about the paper. The paper's own abstract claims something a good deal more modest, that the approach "significantly reduces costs-by over 2 times in certain cases-without compromising the quality of responses" (Ong et al., 2024). Anyone who opens the paper looking for the 85% will not find it and will reasonably wonder what else is off, so cite whichever of the two you mean.

The ladder

Pick the router your traffic can support

Seven ways to make this decision are worth knowing about, and choosing between them starts with the same question M13 asked about classifiers. What does your data support, and how much time does the request have to spare?

Seven rungs compared on added latency and what each one needs. Rules add under one millisecond and need nothing. Embeddings plus kNN add about ten milliseconds and need a few hundred labelled requests. A trained router adds five to fifty milliseconds and needs thousands of graded pairs. A typed-decision model adds a claimed seventy to five hundred milliseconds and needs nothing. An LLM router adds five hundred to four thousand milliseconds and needs nothing. A provider-native router is included in the call and gives up control. A cascade adds no routing latency at all, and pays for a second call on the share that escalates. A band underneath gives the routing cost per sixty thousand conversations: zero for rules and embeddings, fifty cents for the typed-decision model and $3.12 for the LLM router.
The routing decision itself costs almost nothing in money. What it costs is latency, and the top two rungs spend a real share of the turn's deadline.

Rules

A router that reads the M13 ticket class costs nothing and answers in microseconds, because that label was going to be computed anyway. The same is true of most of the other signals sitting in a request already: which tier the customer is on, how long the message is, whether there is an order id in it, whether an image is attached. None of that needs a model to read.

Start here, and stay honest about the fact that a good deal of the reported gains in this module's own sources came from something about this simple.

Embeddings plus kNN

Embed the request once, compare that vector against a few hundred labelled examples you have already embedded, and take whichever neighbours come out closest. Almost all of the ten milliseconds this costs is the embedding call itself, and the comparison afterwards is arithmetic over a few hundred vectors, which barely registers.

This rung has better evidence behind it than its reputation suggests. Yang Li tested it against the learned routers that get the attention and found that "a well-tuned k-Nearest Neighbors (kNN) approach not only matches but often outperforms state-of-the-art learned routers across diverse tasks", concluding that the result "challenges the prevailing trend toward sophisticated architectures and highlights the importance of thoroughly evaluating simple baselines before investing in complex solutions" (Li, 2025).

A trained router

Here you train a small classifier on which model won on which request, which is what RouteLLM is and where the 85% came from. The catch is the dataset. You need thousands of graded pairs, and producing them means running both models over the same traffic and then scoring both sets of answers, so the training data costs real money before the router exists at all. M21 covers the training run itself.

A typed-decision model

This rung did not exist two years ago, and it is where a lot of routing work has been moving. Jev, from TypeSafe, takes a question whose options you list in advance and hands back a probability for each one, which is close to exactly the shape a router wants. TypeSafe prices it at "$0.042 / MTok ($42 per billion tokens)" with output free, and claims "End-to-end response time is 70ms-500ms" (TypeSafe).

You can watch the substitution happening in real projects. The routing engine Bifrost filed a proposal to adopt it, and the useful part of that proposal is how plainly they describe what they were trying to get away from. Their existing semantic classifier is "limited to three fixed labels (SIMPLE/MEDIUM/COMPLEX) based on embedding similarity", and their LLM fallback is "slow (4-second budget), expensive per request, generates unparseable text". Jev, they wrote, "returns typed answers with a probability distribution and a calibrated confidence", with no text to generate and none to parse, at "roughly 1/50th to 1/500th of an LLM-classifier call" (maximhq/bifrost#7278). LangChain documents the same pattern, using it to pick "the least costly model that can complete the task" (LangChain).

One independent measurement exists, and it is the only figure here that TypeSafe did not produce. An engineer at Classmethod built a four-tier router on TypeSafe's API and measured a median latency of 0.643 to 0.674 seconds, a cost of $0.000025 to $0.000027 a call, and ten correct classifications out of ten in each tier (DevelopersIO). Before leaning on that, keep in mind it is forty calls on one person's taxonomy, so it is a sighting, and some way short of a benchmark. Their measured latency also came out above the 70 to 500 ms TypeSafe claims, and a gap like that is worth reproducing on your own traffic before you design a deadline around either number.

The confidence figures from that test are the part worth paying attention to. Jev came back at 1.0 on the simple, complex and reasoning tiers, and at 0.57 to 0.67 on the medium one, which happens to be the tier a person would also struggle to call. That behaviour is what you want, because a router that tells you which requests it found hard is a router where M13's abstain band has something to work with.

An LLM router

Here you write a prompt describing the options and ask a model to pick one. This was the default approach for about two years, and it has quietly become the expensive rung, costing seconds of latency and handing back an answer your code then has to parse. It still earns its keep when the routing criteria are complicated enough that writing them out in a prompt is genuinely easier than any alternative, and when the requests are rare enough that a few seconds of deciding does not hurt.

A provider-native router

On this rung you do not make the decision at all, the provider does. GPT-5 ships with "a real-time router that quickly decides which to use based on conversation type, complexity, tool needs, and your explicit intent", which "is continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness" (OpenAI).

The sentence after that one is worth reading slowly, because it describes a behaviour your capacity planning needs to know about. "Once usage limits are reached, a mini version of each model handles remaining queries." The model answering your request changes based on your consumption, without a deploy, which is the class of silent change M11-2 built a release manifest to catch.

A cascade

The seventh rung has no router in it at all, which is why section 4 got to it first. It adds no routing latency, and pays for a second call on whatever share of requests escalate.

What the decision itself costs

Whichever rung you land on, the decision itself is close to free in money terms:

The router's own bill, 60,000 conversations
RungPer decisionPer monthShare of the saving
Rules on the M13 label$0$00%
Embeddings plus kNN on your own box$0$00%
A typed-decision model$0.0000084$0.500.01%
An LLM router on Luna$0.000052$3.120.06%

Which means the argument against any of these is never what it costs to run. It is the latency, and the ongoing work of keeping the thing correct. Measured against M11-1's 20-second turn deadline, rules spend 0.01% of the budget on deciding, embeddings 0.05%, a typed-decision model somewhere between 2.5% and 3.4%, and an LLM router a full 20%.

Metrics

What a wrong route costs in each direction

M13's per-class precision and recall come across unchanged, so there is nothing new to learn about measuring a router's accuracy. What does change is that the two kinds of error now carry prices, and those prices are nowhere near each other.

A two-by-two grid. The columns are what the request needed, easy or hard. The rows are where the router sent it, cheap or expensive. Top left, easy work on the cheap model, is correct and costs $0.0080. Top right, hard work sent to the cheap model, is a bad answer that costs $0.0080 plus a $2.50 handoff. Bottom left, easy work sent to the expensive model, is a fine answer costing $0.1443, about eighteen times more than it needed. Bottom right, hard work on the expensive model, is correct at $0.1443. A note says the two error cells differ by a factor of about twenty.
Routing up wastes about 14 cents. Routing down risks $2.50. A router tuned as though both errors were equal is tuned wrong.

Routing up costs you money and nothing else. Routing down costs you a bad answer, and at the flower company a bad answer means the conversation ends in a handoff at $2.50, which is roughly twenty times what the wasted frontier call would have cost you.

So the threshold ends up where M13 said thresholds always end up, which is wherever the cost of each kind of error puts it. Given a twenty-to-one gap, that means routing down only when the router is confident and routing up whenever it is unsure, which turns M13's abstain band into a band that quietly sends everything in the middle to the expensive model.

The accuracy of the router moves the result less than you would expect once that is in place:

The same mix, as the router gets worse
Router accuracyModel spend a monthCost per resolved ticketSaved against all-Sol
100%$3,750$0.8679%
95%$3,831$0.8768%
90%$3,913$0.8867%
80%$4,077$0.9055%

Dropping twenty points of router accuracy costs you four points of the saving. That is a much smaller effect than people expect, and it comes from the same arithmetic as everything else in this module, since the handoff dominates the cost of a ticket and the model is a minority of it either way. Practically, it means that when you have a mediocre router the right move is usually to ship it and watch it for a month before deciding whether improving it is worth anyone's time.

One warning about which number you report. Publish cost per resolved ticket, because model spend on its own will make a bad router look excellent. A router that sends absolutely everything to the cheapest model minimises model spend perfectly, and you already saw in section 5 what that does to the bill that counts.

Capability routing

Routing to handlers that are not models

The larger saving at the flower company turns out to have nothing to do with which model answers.

M7 measured that 60% of tickets need no model at all. A tracking link answers "where's my order" and a change form handles a delivery change. Both of those handlers cost nothing per request and answer in tens of milliseconds, against $0.0080 and about two seconds for the cheapest model.

The same 60% of traffic, two ways
HandlerCost per requestLatencyAnswers correctly
A frontier model$0.1443about 4.2 s97%
A small model$0.0080about 1.9 s97% on this class
A database lookup and a template$0about 40 ms100%, since it reads the order

The lookup wins on every column for this class of request, and it is still the handler teams skip, usually because the model is already wired up and already works. M7's framing page made the same point a year earlier, in the module where that argument belongs.

Capability routing does come with one rule that model routing does not have. Here the handlers behave differently, and the difference goes well past cost, so the router has to be right about what kind of request it is looking at, and being right about how hard it is does not help. "Where is order 44718" belongs at the lookup whether the question is easy or difficult. Getting that wrong is also usually unrecoverable in the cascade sense, because a database lookup cannot answer a policy question badly. It answers a different question, or it returns nothing, and neither of those is something a quality check will catch and escalate.

A person belongs on the same list of handlers as everything else. M9 built the handoff and M7 priced it, and it is the one to pick when the router is unsure and the decision moves money.

In production

Routers drift, and so do the handlers they choose between

A classifier drifts when the traffic coming into it changes. A router has that problem too, and then a second one on top of it, because several of the handlers it chooses between are operated by other companies who change them on their own schedule.

The traffic problem is the familiar one. February runs at eight times a normal week and skews toward delivery changes, so a router tuned against November's mix is routing a different distribution two months later. That is M13's drift and it behaves the same way here.

Prices are the second thing that moves, and this one has no equivalent in M13 at all. Every number in this module is a list price read in September 2026, and M12 already noted that Sol's $4 and $20 are promotional, "available at least through November 21, 2026". A routing rule written against a price is a rule with an expiry date on it, whether or not anyone wrote that date down.

The handlers themselves change as well. Models get retired, released and renamed, so a router with a model id hard-coded into it is a release artifact in the sense M11-2 means, and it belongs in the same manifest as the prompt version, where somebody will see it change.

More awkwardly, the model behind a name can change while the name stays put. M4 covered the general case and GPT-5's own documentation gives a specific one, where hitting a usage limit swaps in a mini version. Your router picked a name, and the provider picks the weights.

If you run a cascade, the check drifts too, though at least it tells you. The escalation rate is a live signal, and when it starts climbing it means either the traffic got harder or the cheap model got worse. Separating those two takes M1's eval dataset run against both models, because from the outside they look identical.

Beyond what M13 already asked you to watch, two numbers are worth putting on a dashboard. The share of traffic going to each handler, because that is how a handler changing without anyone telling you first shows up. And, if you run a cascade, the escalation rate, since it is simultaneously your biggest cost driver and your earliest warning.

Shipping a router follows M8 and M11-2 without modification. Run it in shadow first, comparing the choice it would have made against what the live system did, with no effect on any customer. Then canary it, with cost per resolved ticket sitting on the guardrail list next to the quality numbers.

Putting it together

Putting it together

The flower company comes out of all this with two routers and a cascade, which is more separate arrangements than the word routing tends to suggest.

The routing spec
LineThe flower company's answer
Capability routeM13's label sends "where's my order" to the order lookup and delivery changes to the change form, which is 60% of tickets at no model cost
Model routeOf what is left, the router picks Luna or Sol by M13's confidence, routing up whenever unsure
CascadeBehind the queue, where nobody is watching a cursor, Luna answers first and M9's checks escalate the failures
Router chosenRules on the M13 label, because the label is already computed and adds no latency
Quality barCost per resolved ticket, alongside the M1 pass rate, since model spend alone rewards the wrong behaviour
ThresholdsRoute down only above the confident band, route up in the middle, hand to a person at the bottom
MonitoringShare of traffic per handler, escalation rate, cost per resolved ticket, and the prices the rules assumed
RolloutShadow, then a canary, with cost per resolved ticket as a guardrail

M15 takes the next capability, which is checking whether an answer is any good. That bears on this module more than it might look, because a cascade is worth exactly as much as its check is, so the two modules run into each other straight away. M19 covers retrieval, which is one of the handlers here, and M30 covers agents, which is the most expensive one.

Checkpoint · recall · 6 questions

What the module said

  1. 01

    What does a router do?

  2. 02

    Why is a router described as a classifier with a price list?

  3. 03

    What is a cascade?

  4. 04

    What is the break-even escalation rate for a cascade?

  5. 05

    What did Yang Li find about learned routers?

  6. 06

    What does GPT-5's documentation say happens once usage limits are reached?

0 / 6 answered

Checkpoint · understanding · 6 questions

Reason it through

  1. 01

    Routing the flower company's traffic saves 57% of model spend and 9% of cost per resolved ticket. Why the gap?

  2. 02

    Moving every conversation to Luna cuts model spend eighteenfold and raises cost per ticket. What happened?

  3. 03

    Your router is 80% accurate and a teammate wants to spend a month improving it. What does the table say?

  4. 04

    A product has a person watching a cursor, a 3-second budget, and two models 18 times apart in price. Router or cascade?

  5. 05

    Which routing signal would you reach for first at the flower company, and why?

  6. 06

    Your router sends "where is order 44718" to a database lookup and it returns nothing, because the order id was mistyped. Is that a recoverable misroute?

0 / 6 answered

Checkpoint · debugging · 5 questions

Debug it

  1. 01

    Model spend dropped 60% after the router shipped and cost per resolved ticket went up. What happened?

  2. 02

    Your cascade's escalation rate went from 22% to 48% over three weeks with no deploy. Where do you look?

  3. 03

    A rule routes anything under 50 tokens to the cheap model. It worked for months and now misroutes constantly. What is the likely cause?

  4. 04

    Your routing rules assume Sol costs $4 per million input tokens. What breaks in December?

  5. 05

    Traffic to your expensive handler doubled overnight. The router code is unchanged and the mix looks normal. What do you check?

0 / 5 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.