The capability
What a harness is
A harness is the code around the model call. It builds the context that gets sent, runs whatever tools the model asks for, checks what comes back, and decides what happens when any of that fails.
The first thing a harness gives you is one place every model call goes through. M4 put every call behind a single complete() function, and the harness is what lives inside it, which is why retries, deadlines, fallbacks and output checks end up being one edit each in one file.
The second is a move for every failure it can detect. Retry the call, repair the output, fall back to another model or a simpler reply, decline to answer, stop the run, or hand the conversation to a person. Deciding which of those goes with which failure is most of what this module does.
The third is a measurement for the failures it cannot detect at all. Some bad answers arrive well-formed and wrong, with no error for the code to catch, and those either get counted on real traffic or they do not get counted.

A harness can't change the model. The weights are fixed, and no amount of code around the call teaches it something it doesn't know.
It still moves how often the system is right, because it owns all three positions around the call. The four steps in the figure above are those positions.
Before the call, the harness decides what the model sees. The flower agent's retrieved articles arrive with their article_id stamped on each one, which is what lets the model cite an article and lets section 5's check catch a citation that search never returned. Labelling the things you hand a model, so it can refer to one of them without guessing, is a harness change that costs a line of code and moves accuracy on any task where the model has to point at a particular input.
Around the call, the harness decides which model runs and with what settings. Section 6's fallback chain sends the same request to a different model when the first one fails, and a different model gives different answers. Settings move quality too, with the model held constant. Anthropic's April 2026 postmortem, the one in section 2, is a month of users reporting worse output from an unchanged model after three settings around it moved, one of them the default reasoning effort going from high to medium.
After the call, the harness decides what happens to a wrong answer. Section 7's coverage check declines the questions no article covers. On that section's invented eval dataset, it takes accuracy on the answers the agent does give from 97% to 99.4%, with the same model behind it.
So the harness affects quality, and treating it as plumbing that only handles errors is a misconception worth dropping early. What a harness can't do is make one model better on one input. Which model runs, what that input contains, and what happens to the answer afterwards are all its decisions.
Reliability work on an AI feature is two jobs. Building those moves before you need them, and measuring on real traffic whether they worked.
Loud failures and quiet failures
Every failure in this module sits in one of two families, and the split decides which of the two jobs handles it.
- Loud failures arrive with a signal your code can check. An error code, a timeout, output that won't parse, a run that keeps going. Because the code can see them, it can decide ahead of time what to do about each one.
- Quiet failures arrive with no signal. The answer is well-formed and wrong, or quality drops while the error rate stays where it was. Nothing throws an exception, so nothing in your code reacts unless you built something to measure it.
You'll also see "harness" used for the code that runs evals, usually as "evaluation harness." Here it always means the code around the model in production.
Where it shows up
Every production feature with a model call in it has a harness, named or not. The question is whether anybody designed it.
- A support agent deciding whether a slow provider means waiting, retrying, or telling the customer something went wrong.
- A coding assistant deciding what to do when the model's patch doesn't apply.
- A document pipeline deciding whether a field it extracted is trustworthy enough to write into a database.
- A batch job deciding whether to reprocess 400 failed rows now or park them.
The moves are the same in all four. What differs is who is waiting and what the wrong answer costs.
Where this starts
The agent passed its eval and still failed in Valentine's week
The flower delivery company from M7 and M8 is in its second February. The company is made up, and so are its numbers. Last year's test put half the customers on the support agent. This year every ticket goes to the agent first, and February runs about eight times a normal week, the way it did in M7.
The agent has grown since the test. It still answers policy questions from the help-center articles, and since the spring it also has tools that act on orders:
look_up_orderreads an order's status by order id.search_help_centersearches the copy of the help center that M6 keeps in sync.issue_refundrefunds a damaged or late order on its own up to $50, and anything larger waits for a person to approve it.remember_factsaves a fact the customer has confirmed, like their preferred contact channel, for later conversations.hand_offcreates a ticket for the support team with the conversation attached.
Before the refund and memory tools shipped, the team put them through a security review, which M10 walks through.
Before February, everything looked ready. The 200-question eval dataset from M7 came back at 96%, above the 95% ship bar. Then Valentine's week happened, and these are some of the things that went wrong:
On the evening of the 12th the provider slowed down. Replies that normally took about 4 seconds started taking 12, and then calls began failing outright with overload errors. The code's retries added load at exactly the wrong moment, and some customers waited more than 40 seconds before seeing an error.
Long answers got cut off. A few replies hit the max_tokens ceiling from M4 halfway through the JSON, so the article_ids field never closed and the customer saw "something went wrong."
One answer cited an article that search never returned. The reply read well and linked to a real article, but the policy inside it came from nothing the agent had retrieved.
A customer got two tickets. A hand_off call timed out, the code sent it again, and it turned out both requests had reached the ticketing system.
And on Tuesday somebody shipped a prompt edit to make answers shorter. The error rate stayed flat and so did latency. Over the following two days complaints about wrong delivery cutoffs went up, and nothing on the dashboard moved at all.
Most of that list can't show up in the eval. The eval sends known questions through the model and your prompt, then checks the answers. It never makes the provider fail, and it never sees what your code does after a bad response comes back. The Tuesday prompt edit is the one it could have caught, and only if someone had rerun it before the edit shipped.
That list has both families from section 1 in it. The provider slowdown, the cut-off JSON and the duplicate ticket are loud. The invented citation and the Tuesday prompt edit are quiet.
The quiet ones tend to cost the most, and they happen to teams with far more resources than a flower company. In April 2026, Anthropic published a postmortem on a month of reports that Claude Code had gotten worse. The model was unchanged, and in the postmortem's words, "The API was not impacted." The cause was three changes in the code around the model:
- The default reasoning effort went from high to medium.
- A change meant to clear old thinking from idle sessions had a bug, so it cleared it on every turn.
- A system prompt line told the model to keep text between tool calls to 25 words or fewer. On one of Anthropic's evaluations, that line cost 3%.
Each change hit a different slice of users on a different schedule, so what people saw looked like a broad, inconsistent drop in quality, with no single cause visible from the outside (Anthropic, April 2026). The flower company's Tuesday edit is a small version of the third change.
This module builds one table to hold both jobs from section 1, building the moves and measuring whether they worked. Each section adds rows to it, and by the end it has a row for everything on the Valentine's week list:
| Failure | How the code sees it | What the harness does | What counts it |
|---|---|---|---|
| Provider overloaded | A 503 or 529 overload error | Retry with backoff, then fall back | Availability, retry rate |
The failure inventory
Sort failures by what your code can see
Going into Valentine's week, the agent's code handled every failure the same way. Each model call sat inside one try block, and anything that went wrong landed in one except Exception that showed the customer "something went wrong." An overload error got the same message as a malformed request, even though the overload error would have worked a few seconds later and the malformed request never would.
Handling failures differently starts with sorting them by where the signal shows up, because that decides what your code can check and what it can do next. There are four places to look.
The request failed
The call comes back with an error status, or with no response at all, and the provider's SDK, the client library you install to call the API, raises an exception. These split into two groups:
- Transient failures can succeed if you send the same request again a little later. The provider is overloaded, which OpenAI reports as a 503 and Anthropic as a 529. You're over a rate limit and get a 429. The server had an internal error (500), or the connection dropped before any response came back.
- Permanent failures fail the same way every time. A malformed request (400), a bad API key (401), a key without permission (403), a request too large for the endpoint (413). Sending it again changes nothing, so the fix is in the request or in your configuration.
The status code alone doesn't tell you which group you're in. OpenAI returns a 429 when you're sending requests too fast, where waiting helps, and it also returns a 429 when your prepaid credits run out or a spend limit is reached, where waiting doesn't help at all. Anthropic's 429 for a monthly spend cap comes with no retry-after header and, in Anthropic's words, "keeps failing until access resumes." When a spend limit you set yourself is reached, Anthropic returns a 400 (OpenAI, error codes and Anthropic, API errors, as of September 2026). So your code classifies a failure by the error's type and message, and uses the status code only to narrow it down.
The request worked and the output can't be used
The call returns a 200 and the output is still unusable. M4 covered the two common versions. The reply hit max_tokens and stopped partway, which the response's stop reason tells you, or the model refused and returned no object at all. Streaming adds one more. An error event can arrive after the 200, partway through the stream, so a reply can fail after it has already started.
The output is valid and wrong
The JSON parses and matches the schema, the article it cites is real, and the answer is still wrong. Nothing raises an exception. Two things can catch it, rules your code checks against what it already knows (section 5) and measuring answers on real traffic (section 10).
The run went wrong
Every single call can succeed while the run as a whole fails. The agent keeps calling look_up_order with the same wrong order id, or it never reaches an answer. These show up only when your code looks across the steps of a run, which is section 8.
For a single model call, the first three places come down to one decision the harness makes every time a call comes back. The run-level failures wait for section 8. A safe reply in the tree below means a fixed message the harness can always send, like a link to the right help-center article or "a person will reply within the hour."

The SDK already handles part of this
Before you write any retry logic, it helps to know what the SDK does on its own. As of September 2026, OpenAI's and Anthropic's Python SDKs both retry connection errors, 408, 409, 429 and every 5xx status twice by default. They wait about half a second before the first retry and double the wait each time, up to 8 seconds. When the provider sends a retry-after header, they use that. Both default to a 10-minute timeout (OpenAI Python SDK, Anthropic Python SDK). All of this happens inside the SDK call, so your code never sees the first two failures.
| Failure | How the code sees it | What the harness does |
|---|---|---|
| Overloaded or rate limited | 503, 529, or a rate-limit 429 | Retry with backoff (section 4) |
| Connection dropped | SDK connection error | Retry with backoff (section 4) |
| Out of credits, or at a spend cap | A 429 or 400 naming spend or credits, with no retry-after | No retry. Safe reply, and alert whoever owns the account |
| Bad request | 400, 401, 403, 413 | No retry. Fix the request |
| Cut off | Stop reason is max_tokens | Repair (section 5) |
| Refused | A refusal in place of the object | Safe reply (section 5) |
| Stream failed after the 200 | An error event mid-stream | Retry (section 4) |
| Valid and wrong | Nothing | Business rules (section 5), then measurement (section 10) |
| Same call again and again | A repeated tool call and no answer | Stop the run (section 8) |
The column that says which number counts each failure fills in once section 10 sets up the measurements.
Retries and timeouts
Retry what can succeed later, and give every call a deadline
The evening of the 12th went like this. Around 7 p.m. the provider slowed down, and replies that took about 4 seconds started taking 12. A few minutes later some calls came back as overload errors. The agent's code had a loop of three tries around each call, and the SDK inside it was still on its default of two retries. So a failing turn sent the same request up to nine times. Nothing in the code had a timeout shorter than the SDK's 10 minutes, so a slow call just kept waiting. Customers watched the typing indicator for 40 seconds or more and then got "something went wrong."
The retries also made the overload worse. Every failed call turned into more calls, and every customer who gave up and sent the message again started a new turn on top of the old one. A provider that's overloaded recovers when it gets less traffic, and the code was sending it more.
Retry only failures that are transient and safe to repeat
A failure is worth retrying when two things are true. It's transient, from the list in the last section, and sending the request again can't do any harm. A model call changes nothing outside the provider, so repeating it costs time and tokens and nothing else. A tool call that changes something is different.
hand_off creates a ticket. When it times out, your code can't tell whether the ticket was created, which is the unknown outcome M6 described for the refund. Sending it again is safe only with an idempotency key, the id M6 attached to the refund request so the payment provider could recognize a repeat. For hand_off, build the key from the conversation id, so every attempt for the same conversation carries the same key and the ticketing system creates one ticket. If the ticketing system doesn't accept keys, check whether a ticket for that conversation already exists before sending again, the way M6 checked with the payment provider.
Agent loops add a retry you didn't write. In many agent frameworks, a tool error goes back to the model as text, and the model can decide to call the tool again. So the key belongs inside the hand_off tool itself, where every call passes through it, whether the retry came from your code or from the model.
Wait longer each time, and at random
A retry sent the moment a call fails usually fails too, because whatever caused the failure hasn't cleared yet. So the wait grows with each attempt, which is called exponential backoff. You wait about half a second, then about one, then about two, up to a cap.
Backoff on its own still has a problem at scale. Thousands of conversations failed in the same few seconds on the 12th. If they all wait exactly one second, they all retry in the same second and hit the provider together again. Jitter spreads them out by picking a random wait between zero and the backoff value. The version from AWS's architecture blog, called full jitter, is one line of code, wait = random.uniform(0, min(cap, base * 2 ** attempt)), where base is the first wait and cap is the longest one (Brooker, Exponential backoff and jitter).
When the provider sends a retry-after header, it's telling you how long to wait, so that value replaces your own.
Retry in one place
The nine requests per turn came from two retry loops stacked on each other. Each layer multiplies the one inside it, so the fix is to pick one layer and turn the others off. There are two clean ways to do that:
- Set
max_retries=0on the SDK client and own the retry policy insidecomplete(), which can see your deadline and log every attempt. - Keep the SDK's two retries and add none of your own.
This module uses the first, because the SDK can't see your deadline. As of September 2026, Anthropic's Python SDK waits as long as a retry-after header asks, and OpenAI's will wait up to two minutes. Either one can run well past the time a customer will wait. When retries stack across whole services, like a web server calling a worker that calls complete(), the multiplication gets much larger, and M11 covers that.
Cap the attempts, and when they run out, send the call somewhere defined, the fallback in section 6 or a handoff to a person in section 7, so no customer sees "something went wrong" because an exception went unhandled.
Log every attempt as its own step in the trace, the way M2 described. That gives you two numbers for each place your code calls a model:
- First-try success rate is the share of calls that worked on the first attempt.
- Success after retries is the share that worked on any attempt.
M2's checkpoint had a step at 99% success that only succeeded after two silent retries. The gap between these two numbers is how that step shows up on a dashboard, and M12 uses the same gap to work out what the retries cost.
Give every call a clock
Every call over the network needs a timeout, and the SDK's default of 10 minutes is far longer than any customer waits. A single timeout number fits an LLM call badly, though. How long a call takes depends on how much the model writes and how long it reasons, so a timeout long enough for a long answer is far too long to notice a provider that's stuck. Splitting it into separate clocks fixes that:
The connect timeout covers waiting for the connection itself, and a couple of seconds is plenty for that.
Time to first token covers how long until the first piece of the reply arrives. It includes the time the request spends waiting in the provider's queue plus the time the model takes to read your input, which is why an overloaded provider shows up in this number before any other.
The stream idle timeout covers how long the stream may go without a new token once it has started producing them.
And the total deadline covers how long the whole turn may take, across every attempt and any fallback. That one comes from how long a customer is willing to wait, with no technical input at all.
The first three are timeouts, limits on one attempt. The deadline limits the whole operation, and each attempt gets at most what's left of it. The two middle clocks need a streamed response, because that's how your code sees the first token arrive and notices when tokens stop coming. The agent can stream from the provider and still hold the answer back from the customer until it passes the checks in section 5. Because the customer hasn't seen any of it, a stream that dies partway through can be retried like any other transient failure.
These numbers are made up, like the company's other numbers. On a normal evening the first token arrives within about a second and the whole reply takes about 4 seconds. The flower company decides every customer gets something within 20 seconds. The loop it had can't promise that. Even with a first-token timeout of 8 seconds on each try, three tries that each time out take 24 seconds plus the waits between them, before anything else runs.
A budget that fits looks like this:
- first-token timeout of 4 seconds and stream-idle timeout of 2 seconds
- each attempt capped at 8 seconds
- two attempts on the primary model, with a jittered wait of up to a second between them
- the fallback model from section 6, then the safe reply
- before starting any attempt, check the deadline, and if less than 8 seconds are left, skip to the next step down
That last rule is what makes the budget hold. If the provider is stuck and the first token never comes, the two attempts give up at about 4 and 9 seconds. The fallback model starts with 11 seconds left, and on a normal reply it answers by about 13. If the replies are slow to stream, each attempt runs to its 8-second cap and the second one ends around 17 seconds. That leaves too little time for the fallback, so the safe reply goes out. Either way the customer hears back inside 20 seconds.

Put together, the retry loop inside complete() is short:
import random, time
from openai import OpenAI, APIConnectionError, APIStatusError
client = OpenAI(max_retries=0) # the SDK stops retrying, complete() owns it
def is_transient(err) -> bool:
# a dropped connection, or one of our own clocks fired
if isinstance(err, (APIConnectionError, ClockExpired)):
return True
if isinstance(err, APIStatusError):
if err.status_code == 429:
return not is_spend_or_credit_limit(err)
return err.status_code in (408, 409) or err.status_code >= 500
return False
def call_primary(request, deadline, attempts=2, per_attempt=8.0):
for attempt in range(attempts):
if deadline - time.monotonic() < per_attempt:
break # no time for a full attempt
try:
return stream_reply(request, first_token=4.0, idle=2.0,
total=per_attempt)
except Exception as err:
log_attempt(attempt, err) # one trace step per attempt
if not is_transient(err):
raise # permanent, a retry won't help
if attempt < attempts - 1:
backoff = random.uniform(0, min(8.0, 0.5 * 2 ** attempt))
wait = retry_after(err) or backoff
if deadline - time.monotonic() - wait < per_attempt:
break # waiting would leave too little
time.sleep(wait)
raise ModelUnavailable() # complete() moves to the next tierstream_reply is the piece that opens the stream and enforces the three timeouts, raising ClockExpired when one fires. is_spend_or_credit_limit and retry_after read the error's body and headers. The check after the wait applies the same rule to retry-after, so a provider that asks for a long wait sends the turn straight to the fallback.
Some work can't fit inside a customer's wait at all. The nightly job that summarizes every closed ticket makes thousands of calls, and nobody is waiting on any one of them. That work belongs off the request path, in a background job that saves its progress after each step the way M6's refund run did, or in a provider's batch API, and M11 covers running it.
Three more rows go into the table:
| Failure | How the code sees it | What the harness does |
|---|---|---|
| Provider stuck | The first-token timeout fires | Retry once, then fall back |
hand_off timed out | A timeout on a tool that writes | Retry with the same idempotency key |
| Running out of time | Under 8 s left before an attempt | Skip to the next step down |
Validation
Check the output before anything acts on it
Two of the Valentine's week failures got past the retries because nothing was wrong with the call. The long answers came back with a 200 and a stop reason of max_tokens, and the code handed the half-finished JSON straight to the parser, which threw, so the customer got the generic error. The other answer came back complete and well-formed. It cited article 212, a real article about holiday delivery, but search had returned articles 118 and 140 for that question. The answer also said changes close 24 hours before delivery, where article 140 says 48, which is the wrong detail from M7's list of failures.
Code could have caught both, because both are checkable. The agent's answers already come back as structured output from M4, a JSON object with the reply text, the article_ids it used, an optional order_id and a needs_human flag. That guarantees the shape of the reply, and OpenAI's own documentation says "Structured Outputs can still contain mistakes" (OpenAI, Structured Outputs). A schema has no way to know which articles search returned or what article 140 says. Your code does, because it made those calls a moment earlier.
Run the cheap checks first
The checks run cheapest first, and each one has an action decided ahead of time for when it fails. Call the sequence a validation ladder:
- 01Stop reason. Free, since it's a field on the response. A stop reason of
max_tokensmeans the reply was cut off, so the code doesn't parse it and asks again with a higher ceiling. A refusal means there's no object at all, so the customer gets the safe reply. - 02Parse. Takes microseconds. Text that isn't valid JSON goes to repair.
- 03Schema. Check the fields and types in your own code, even when the provider decoded against a strict schema. As of September 2026, Anthropic's documentation says string
enumvalues can come back with different capitalization, with no error and no special stop reason. Some schema limits, like a minimum value or a maximum length, aren't enforced during decoding at all, so the SDK checks them after the reply arrives (Anthropic, structured outputs). - 04Business rules. Takes milliseconds. These check the answer against what your code knows about this conversation. For the support agent, every
article_idin the answer has to be one search returned on this turn, and every number in the answer, like hours or dollar amounts, has to appear in one of the cited articles. - 05A model-based check. One more model call, where a second model reads the answer next to the articles and judges whether the articles support it. It catches wrong answers that no rule can express, and it costs a call and the time that call takes. How to build one and measure it against human labels is M15's job. Until it's measured, treat its verdict as one signal among several.
- 06A person. Section 7.
The support agent has two more business rules. An order_id in the answer has to match the order look_up_order returned, and anything M7's handoff rules send to a person has to come back with needs_human true. For outputs that are code or SQL, there's also one more rung between the business rules and the model-based check, where you run or dry-run the output. The support agent doesn't need it.
Asking the model that wrote an answer whether the answer is right doesn't count as a check. In Mata v. Avianca (S.D.N.Y. 2023), lawyers filed a brief citing court cases ChatGPT had invented. When one of them asked ChatGPT whether one of those cases was real, it said yes. The court fined the lawyers and their firm $5,000. A check that each cited case exists in a legal database is a few lines of code, and it's the same kind of check as "every article_id is one search returned."

Test the rules before you trust them
The numbers rule catches the 24-hour answer. It also flags correct answers that put a number differently, like "two days" for an article that says 48 hours. Every correct answer a rule flags goes to repair or gets the safe reply, so a noisy rule costs you good answers. Before switching a rule on, run it over the eval dataset from M7 and count how often it flags an answer the labels say is right. That's the false-positive side of the precision question from M1, applied to your own checks.
The rules are ordinary functions, so the same code runs in two places. In the eval it runs against every test case, and in production it runs against every reply before the reply is used:
import re
def check_answer(answer, retrieved: dict[str, str], order=None) -> list[str]:
"""Returns the problems found. An empty list means the answer can be used."""
problems = []
for article_id in answer.article_ids:
if article_id not in retrieved:
problems.append(f"cites {article_id}, which search did not return")
cited_text = " ".join(retrieved.get(a, "") for a in answer.article_ids)
for number in re.findall(r"\d+(?:\.\d+)?", answer.text):
if number not in cited_text:
problems.append(f"{number} does not appear in the cited articles")
if order and answer.order_id and answer.order_id != order.id:
problems.append("order id does not match the looked-up order")
return problemsA rule you add after an incident becomes a test case at the same time, which is how the article 212 failure stays fixed.
Repair once, then fail closed
When a check fails, the harness tries the cheapest repair first:
- 01Fix it in code when the fix is certain. Strip a code fence wrapped around the JSON, or lowercase a
hand_offreason that came back as"Refund"when the tool's schema lists"refund". - 02Ask once more with the failed output and the exact problem, like "Article 212 was not in the search results. Cite only 118 or 140." This is called a repair, or a re-ask. It's a full new call with a longer input than the first one, so cap it at one, and count it separately from the transport retries in section 4.
- 03Try the fallback model from section 6.
- 04Send the safe reply or hand the conversation to a person.
If the repair fails too, the answer isn't used. That's what fail closed means. Anything that acts, like a tool call that writes, or that reaches the customer as a statement of policy, fails closed. Failing open, using the output anyway and logging the problem, is fine only for failures at the cheap end of M7's list, like a missing optional field. For the article 212 answer, one repair either produces an answer that cites 140 and says 48 hours, or the customer gets the link to article 140 and the offer of a person.
The rows from this section:
| Failure | How the code sees it | What the harness does |
|---|---|---|
| Cut off | Stop reason is max_tokens | Ask again once with a higher ceiling, then the safe reply |
| Won't parse | JSON parse error | Fix it in code, or repair once |
| Enum in the wrong case | Schema check in your code | Normalize and continue |
| Cites an article search didn't return | Business rule | Repair once, then the safe reply |
| A number that isn't in the article | Business rule | Repair once, then the safe reply |
| Refused | Stop reason is refusal | Safe reply |
Fallbacks and degraded modes
Fall back to a path you already run
Retries cover failures that clear in seconds, and some provider failures last for hours. On December 11, 2024, every OpenAI service was badly degraded or down from 3:16 to 7:38 p.m. Pacific time. In June 2025, an operating-system update on OpenAI's GPU servers pushed API error rates to about 25%, and full recovery took about 13 hours (OpenAI status, December 2024 and June 2025). No retry loop with a 20-second deadline outlasts that. During an outage like the December one, every turn uses up its two attempts, and every customer gets the safe reply for the whole evening.
A fallback is a second path the harness sends the same request down when the first one fails. The work is in choosing the order the paths are tried in, and in making sure each path works before the day you need it.
The chain lives inside complete()
The fallback chain is an ordered list of paths inside complete(). The harness tries each tier only when the one above it has failed:
- 01The primary model, retried. Section 4's two attempts on GPT-5.6 Sol.
- 02The same model through another route. Some models are sold through more than one platform, the way Claude is available from Anthropic's API and also through Amazon Bedrock and Google Cloud's Vertex AI, and one route can be up while another is down. It's the smallest change because the model is the same, but it needs a second account, and the flower company doesn't have one.
- 03A different model. Usually from another provider, so one company's outage can't take out both. The flower company's is Claude Sonnet 5.
- 04A degraded mode. A reply that needs little or no model, covered at the end of this section.
- 05A person.
hand_off, and the queue in section 7.
Two rules keep the chain working during an outage. The first is a cooldown. Once a tier has failed several times in a row, say three, the harness skips it for a fixed time, say 30 seconds, so every customer during an outage doesn't spend the first 9 seconds of their deadline on a provider that's already known to be down. LiteLLM's router has this built in as allowed_fails and cooldown_time (as of September 2026), and M11 covers the general version, the circuit breaker, for any service you call.
The second rule is that a permanent error from a fallback tier means skip the tier. A fallback can reject a request the primary accepted, because complete() builds the request differently for each provider. The Anthropic branch of M4's complete() forces the model to call its output tool with tool_choice, and as of September 2026, Claude Fable 5.1 rejects that with a 400. With Fable 5.1 as the fallback and that code unchanged, every fallback call during an outage would come back 400. The harness skips the tier and alerts a person, and never retries it.
TIERS = [
Tier("gpt-5.6-sol", call=call_model, allowed_fails=3, cooldown_s=30),
Tier("claude-sonnet-5", call=call_model, allowed_fails=3, cooldown_s=30),
]
def complete(request, deadline):
for tier in TIERS:
if tier.cooling_down():
continue # failed recently, skip it for now
try:
return tier.call(tier.name, request, deadline)
except ModelUnavailable: # retries or clocks ran out (section 4)
tier.record_failure() # enough of these starts a cooldown
except APIStatusError as err: # permanent: this tier rejects the request
alert_owner(tier.name, err) # skip it and tell a person
return degraded_reply(request) # every model tier failed or was skippedcall_model is section 4's retry loop with section 5's checks after it, so an answer that still fails its checks after one repair also sends the turn down a tier.
A fallback model is a model swap
M4 said every model swap goes back through the eval. A fallback is a swap that happens in the middle of an outage with nobody deciding anything, so its eval has to be done before the outage. With made-up numbers, on the 200-question eval dataset, Sol with the current prompt scores 96%. Sonnet 5 with the same prompt scores 91%, and with a prompt adjusted for it, 95%. So the fallback carries its own prompt, and the head of support decides ahead of time whether 95% answers for an hour beats safe replies for an hour. The eval also has to run through the Anthropic branch of complete(), because the tool definitions and the output format go through different code on each provider.
Keeping two providers working costs something. Assembled, which makes software for customer support teams, reported in 2025 that falling back between providers took their effective uptime to 99.97%, against status-page uptime of 99.80% for OpenAI and 99.58% for Anthropic over the months they measured. They also reported that keeping prompts working across providers added 20 to 30% to prompt development time (Assembled, June 2025). Those are one company's numbers, and they show the cost next to the benefit.
A fallback you never run won't work when you need it
Around 2001, Amazon's retail site added shipping speeds to its product pages. The data lived in a supply chain database that couldn't keep up with the site's traffic, so each web server kept a cache, and if the cache failed, the web server queried the database directly. That worked for months. Then the caches all failed around the same time, every web server hit the database at once, and the database locked up. The whole site went down, and because the fulfillment centers used the same database, they stopped too, worldwide. A missing shipping estimate became a full outage because of the fallback. AWS's write-up of the outage says, "One of the worst things about fallback is that it isn't exercised regularly and is likely to fail or increase the scope of impact when it triggers during an outage" (Gabrielson, Avoiding fallback in distributed systems).
AWS's answer is to run both paths all the time. When both paths serve real traffic continuously, the setup is called failover, and a fallback only runs when the primary fails. For the flower company, that means sending a steady 5% of real conversations to Sonnet 5 every day. Its API key, its prompt, its rate limits and its branch of complete() then get used by real customers every hour, and a broken fallback shows up as a bad number for that 5% on an ordinary weekday. Its eval reruns on every prompt change, the same as the primary's. Whether Sonnet 5 can take all of the traffic at once when Sol is down is a separate question about capacity and rate limits, which M11 works through.
The person at the bottom of the chain is a path too. A handoff queue that only fills up during outages has the same problem, which section 7 picks up.

Degraded modes and the kill switch
When every model tier is down or skipped, the customer still gets a reply. A degraded mode is a reduced reply decided ahead of time with the head of support, and for the flower company most of it needs no model at all:
- A message with an order number in it gets the
look_up_orderresult in a fixed template, like "Order 4417 is out for delivery and should arrive by 6 p.m." The lookup is your own code and your own database, so it works while the providers are down. - A policy question gets the top
search_help_centerresult as a link. - Everything else gets "A person will reply within 2 hours," and a
hand_offin the background.
The Google SRE book's advice on degraded modes is to keep them simple and to alert when they trigger, and it gives the same reason AWS did, "Remember that the code path you never use is the code path that (often) doesn't work" (Google SRE, addressing cascading failures). So the degraded mode gets a drill too, on a schedule, with a person checking what customers would see.
A kill switch is a flag an operator can flip, without a deploy, to send every conversation to the degraded mode or straight to a person. Fallbacks handle a model that's unavailable. The kill switch is for a model that's up and answering badly, like the answers after Tuesday's prompt edit. In January 2024, after a system update, the delivery company DPD's chatbot swore at a customer who asked it to, and the company said it had disabled the AI part of the chatbot straight away. NIST's AI Risk Management Framework asks for the same thing in general terms, a way to disengage or deactivate an AI system whose outcomes don't match what it was meant to do (NIST AI RMF, MANAGE 2.4). Decide ahead of time who can flip it and on what signal. M11 covers building the flag so flipping it is fast and safe.
| Failure | How the code sees it | What the harness does |
|---|---|---|
| Provider down for hours | Retries exhausted, cooldown on | Next model tier (Sonnet 5) |
| Fallback rejects the request | A permanent 400 from that tier | Skip the tier and alert a person |
| Every model tier down | All tiers failed or skipped | Degraded mode |
| Model up, answers wrong | A person decides | Kill switch to the degraded mode |
| Fallback broken, unnoticed | The 5% share's numbers drop | Fix it on an ordinary day |
Confidence and escalation
Decide when the agent shouldn't answer, and who should
Some wrong answers pass every check in section 5. During Valentine's week a customer asked whether flowers can be delivered to a hospital's intensive care unit. No help-center article covers that. Search returned the two closest articles, both about hospital deliveries in general, and the agent wrote a fluent answer from them that said yes. Every cited article was one search had returned, and every number in the answer appeared in them. The answer was still invented.
Cursor's support email had the same failure in April 2025. Users who were being logged out when they switched machines wrote in, and an AI answering the support email told them it was expected behavior under the subscription's rules. Cursor had no such rule. A co-founder replied on Hacker News, "Apologies - something very clearly went wrong here," and said that "Any AI responses used for email support are now clearly labeled as such" (Hacker News thread, The Register).
So the harness needs two more decisions made ahead of time. The first is when the agent shouldn't answer at all, and the second is where the conversation goes when it doesn't.
Pick a confidence signal you can check
The obvious move is to ask the model how sure it is, and that's the weakest signal available. In a study published at ICLR 2024, models asked to state their confidence mostly said something between 80% and 100%, usually in multiples of 5. GPT-4's stated confidence separated its right answers from its wrong ones with an AUROC of 62.7%, where AUROC is a score for how well a signal separates the two, 50% is a coin flip and 100% is perfect (Xiong et al.). Those were 2023 models and newer ones may do better, but a number the model writes about itself is still something your code can't check, which is why it goes last on the list below.
The signals worth using, most trustworthy first:
The strongest signal is evidence your code can check for itself. A retrieved passage that contains the answer, an order lookup that confirms the status, the section 5 checks coming back clean. That is evidence more than confidence, and it is the one to build on.
Below that is agreement across samples. Ask the same question several times and compare the answers by meaning, and when they disagree the model is probably making something up. A 2024 Nature paper called this semantic entropy and found it picked out likely-wrong answers with an AUROC of 0.790, against 0.698 for asking the model whether its own answer is true. It costs several calls per question, and the authors are careful to note that "it does not help when LLM outputs are systematically bad," meaning an error the model makes the same way every time (Farquhar et al.).
Then token probabilities, which only work in a narrow case. When the output is a single label, the probability the model assigned that token is a usable score, on the providers that return it. As of September 2026 Anthropic's API does not, and OpenAI returns them on some models and not others.
Below that, a second model acting as a judge, which is the model-based check from section 5, measured against human labels before you trust any of its output.
And last, the model's own stated confidence, which section 3 already showed is the weakest signal on the list.
For the flower company, the fix for the intensive care answer comes from the first signal. The harness adds a coverage check before the agent answers, one call to a small model that reads the question and the retrieved passages and answers covered or not_covered. The model is picked partly because its provider returns token probabilities, and the probability it gave its label is the check's score. The intensive care question had passages about hospitals and none about intensive care units, which is what the check is for.
Set the threshold on the eval dataset
A score needs a cutoff, and the cutoff comes from what each kind of mistake costs, the same way M1 chose between recall and precision. To set it, the flower company adds 50 questions no article answers to the 200-question eval dataset, runs the coverage check over all 250 and tries different cutoffs. With made-up numbers, at a cutoff of 0.5 the agent answers 78% of the 250 questions and 97% of those answers are right. At 0.8 it answers 70%, and 99.4% of those are right. The higher cutoff declines some questions an article does answer, which is the price of declining nearly all the ones no article answers.
Those two numbers always get reported together. Coverage is the share of questions the agent answers, and accuracy is measured on the answered share only. Raising the cutoff trades coverage for accuracy. On M7's list of failures, an invented policy is the most expensive one, so the flower company takes 0.8 for policy questions and accepts that more of them go to a person.

Declining doesn't always mean handing off. When the customer's question is unclear, like "can I change my order" when it could mean the address or the date, asking which one they mean is also a way of not guessing.
Escalate on behavior and risk
Escalation is the decision to route a conversation to a person, and the handoff is the act of sending it, which is M7's word for it. OpenAI's guide to building agents names two triggers for handing control to a person. One is exceeding failure thresholds, like an agent that "fails to understand customer intent after multiple attempts." The other is high-risk actions, with examples including "canceling user orders, authorizing large refunds, or making payments" (OpenAI, A practical guide to building agents). Both triggers are about what the agent did and what's at stake, and neither uses a confidence score.
The flower company's triggers follow the same idea:
- one of the handoff rules from M7 fires, like no matching article, the same question asked twice, or a customer asking for a person
- the coverage check comes back under its cutoff
- an answer fails the section 5 checks after its one repair
- the run hits a stop condition (section 8)
- a refund over the $50 limit, which
issue_refunditself sends for approval
Each trigger logs its own count, so when the escalation rate jumps, the trace shows which trigger caused it.
The harness runs these checks in a fixed order. The handoff rules come first, because they read only the customer's message and need no model call. The coverage check comes after search and before the agent writes anything, and the section 5 checks come last, on the finished answer.

Send a packet the person can act on
The person who gets the conversation shouldn't have to redo the agent's work to decide. A handoff packet is what the harness sends with it. In the code it's one JSON object with these fields:
| Field | Example | Why it's there |
|---|---|---|
conversation_id | c_81f3 | Links the ticket to the stored conversation |
trigger | coverage_below_cutoff | Says why the agent stopped, and feeds the count per trigger |
customer_messages | "Can you deliver to the ICU at St. Mary's?" | The question in the customer's own words |
draft_reply | "Yes, we deliver to hospitals, including intensive care units." | What the agent would have sent |
evidence | Articles 118 and 140, with the passages | Lets the reviewer check the draft against the source |
already_tried | Coverage check scored 0.41, answer held back | So nobody repeats a step that already failed |
trace_url | The M2 trace of the run | For when the packet isn't enough |
reply_due_by | Feb 12, 19:40 | The time the safe reply promised the customer |
With the draft and the passages side by side, the reviewer can see in a few seconds that neither passage mentions intensive care. The trace link is the M2 trace of the run, for when the reviewer needs more than the packet shows.
Size the queue before February
Every escalation is a person's time. For every 1,000 conversations, a 10% escalation rate at six minutes each is 10 hours of reviewer time, and a 30% rate is 30 hours. At eight times normal volume in February, that difference decides whether the support team keeps up or falls days behind. The escalation rate is also a cost number. M7's blended 77 cents per ticket assumed 3 in 10 policy tickets still go to a person, so a trigger that fires more than planned changes the cost per ticket as well as the queue, and M12 works through that side.
Put the gate in code
In July 2025, Jason Lemkin, the founder of SaaStr, was building an app with Replit's agent and told it in the chat that the code was frozen. The agent then ran a command that deleted his app's production database, which held 1,206 executive records. Replit's fix was structural. Its apps had used a single database for both development and live data, and it rolled out automatic separation of development and production databases (The Register, Replit).
A rule written in the prompt is something the model reads and usually follows. A gate in code is a check the model can't get around. For the flower company, "refunds over $50 wait for a person" lives inside issue_refund as a check on the amount. A refund over the limit waits for approval whatever the model was told, and nothing about it depends on the model choosing to call hand_off.
Design for the reviewer
A person reviewing the agent's work is a weak check unless the review is designed for it. People tend to over-trust what an automated system recommends, which is called automation bias, and a 2010 review of the research found that it "occurs in both naive and expert participants, cannot be prevented by training or instructions" (Parasuraman and Manzey). Anthropic reported in March 2026 that Claude Code users approve 93% of the permission prompts they're shown (Anthropic, Claude Code auto mode). A reviewer looking at a fluent draft will usually approve it. M10 covers the other side of that, what an attacker can put in front of the person approving.
What helps:
- Escalate fewer conversations, and make each one worth a person's time, by tuning the triggers on the dataset.
- Show the evidence next to the draft, so checking the draft means reading the passage.
- Track how often reviewers change the draft. A rate near zero means either the triggers are too loose or nobody is reading.
- Mix in a few drafts you know are wrong and see how many reviewers catch.
- Add every reviewed conversation to the eval dataset from M1.
| Failure | How the code sees it | What the harness does |
|---|---|---|
| An answer with no source | Coverage check under its cutoff | Decline, and hand off with the packet |
| An unclear question | Two readings of the request | Ask which one they mean |
| A handoff rule fires | A rule check in code | Hand off with the packet |
| The queue can't keep up | Escalation rate by trigger | Tune the noisy trigger |
| Reviewers approve by habit | Override rate near zero | Seed known-wrong drafts, show the evidence |
Stop conditions and consistency
Put a limit on every run
On the Thursday of Valentine's week, a customer typed their order number with two digits swapped. look_up_order returned "no order found." The agent apologized and called look_up_order again with the same number. It got the same error and tried again. The run was 40 tool calls deep when the customer closed the chat, and every one of those calls re-sent the whole growing conversation to the model, the way M3 described. No single call failed in a way section 3's list would catch. The run as a whole failed.
Every agent loop needs rules that end it. OpenAI's guide to building agents lists the usual ones, "Common exit conditions include tool calls, a certain structured output, errors, or reaching a maximum number of turns." At least one of them has to be a hard limit in your code, one the model can't get past.
Don't rely on the framework's default
If you build the loop on an agent framework, it has a limit already, and the defaults are far apart. As of September 2026:
- OpenAI's Agents SDK stops after 10 turns by default and raises
MaxTurnsExceeded. - Anthropic's Claude Agent SDK has no turn limit unless you set
max_turns. - LangGraph stops a graph after 1,000 steps by default.
A run on one of those stops at 10, and a run on another keeps going. Set the limit yourself, and check what the framework counts as a turn. Claude's SDK counts a tool-use round trip as a turn, and a model that asks for three tools in one reply makes three tool calls in that one turn, so a turn limit and a tool-call limit are two different limits (OpenAI Agents SDK, Claude Agent SDK, LangGraph).
Stop conditions for the support agent
Repeated steps are a common way runs fail. In a 2025 study that labeled the failures of multi-agent LLM systems, repeating a step accounted for 15.7% of failures, and not recognizing that the task was finished accounted for 12.4% (Cemri et al., MAST). The flower company's harness checks five things on every step:
Cap the model calls per customer message at 8, given that a policy answer normally takes two or three.
Cap the same tool with the same arguments at twice. The harness keeps a count keyed on the tool name and its arguments, which is a few lines of code.
Stop on the same error twice in a row from a tool, since the next call is going to get the same error.
Enforce the deadline, which is section 4's 20 seconds per customer message.
And cap spend per conversation, so that one conversation cannot cost an unbounded amount. M12 sets that number and adds per-customer limits on top of it.
The swapped-digit run would have stopped after the second "no order found." Tool results can also do more work here. A lookup that fails with "No order 44718. Ask the customer to check the number." gives the model a next step, and an agent that asks the customer usually doesn't loop.
Decide what a stopped run returns
A run that hits a limit still owes the customer an answer, so decide what it says ahead of time. It shouldn't be a silent timeout or an exception. For the support agent, the harness tells the customer what it couldn't do, like "I couldn't find order 44718. Could you check the number? I've also passed this to the team," and calls hand_off with stop_condition as the trigger. The run itself gets a status that says why it stopped, like stopped_repeated_call, the same way M6 gave an uncertain refund the status needs_review, so a stopped run is easy to find later.

Consistency is its own number
M1's eval runs each test case once. A 96% on one run doesn't tell you whether the same questions pass tomorrow, and a customer-facing agent needs the same question to get the same right answer every time it's asked. Two measures come up constantly in agent work, and they answer different questions:
- pass@k is the chance that at least one of k tries succeeds. It fits tasks where you can try several times and keep the one that works.
- pass^k is the chance that all k tries succeed. It fits a support agent, where every customer gets one try.
The gap between them grows fast. Anthropic's example is an agent that succeeds 75% of the time on a task, whose chance of passing three tries in a row is 0.75³, about 42% (Anthropic, Demystifying evals for AI agents). On τ-bench, a benchmark of simulated customer-service conversations published in 2024, GPT-4o solved about 61% of retail tasks on one try, and fewer than 25% of tasks on all of eight tries (Yao et al.).

Setting the temperature to zero doesn't make runs repeatable. M4 explained why, as the batching on the provider's GPUs changes with load. Thinking Machines sampled the same prompt 1,000 times at temperature zero on an open model and got 80 different completions (Thinking Machines). So consistency has to be measured by running the eval more than once, which section 9 does.
| Failure | How the code sees it | What the harness does |
|---|---|---|
| The same call again and again | Same tool and arguments, twice | Stop, tell the customer, hand off |
| A tool error that repeats | The same error twice in a row | Stop, ask the customer to check |
| A run that goes long | 8 model calls in one message | Stop with what it has, hand off |
| Answers that change between runs | pass^5 well under pass@1 | List the test cases that flip |
Testing the harness
Break the harness on purpose before production does
Before Valentine's week, the retry loop and the fallback chain had never run. The test suite called the real model, which answered normally, and nothing in the suite could make a provider return a 529 or cut a reply off at max_tokens. So the first time that code ran was during the outage, with customers waiting.
Keep two kinds of tests
- Harness tests replace the model with a stub that returns whatever the test scripts, and they block the network. They check what your code does with each kind of response. They run in seconds and give the same result every time, so they run on every commit.
- Evals call the real model and check the content of its answers, the M1 kind. They cost time and tokens, and their results vary from run to run.
Agent frameworks usually ship a fake model for the first kind. Pydantic AI, for example, has TestModel and FunctionModel, plus a setting, ALLOW_MODEL_REQUESTS = False, that makes any accidental call to a real model fail the test (Pydantic AI, testing, as of September 2026).
Script every failure from this module
The fault matrix has one row for each failure the harness is supposed to handle, with the outcome the test checks. Deliberately producing a failure like this to test the handling is called fault injection. The flower company's matrix comes straight from the sections above:
| The stub returns | The test checks |
|---|---|
| A connection reset, then success | 2 attempts, and the customer gets the answer |
| A 529, three times | The fallback tier answers |
| A 429 whose message names a spend limit | 1 attempt, the safe reply, the owner alerted |
A 429 with retry-after: 2 | At least 2 s pass before the retry |
| The first token after 6 s | The first-token timeout fires at 4 s |
Stop reason max_tokens and cut-off JSON | One more ask with a higher ceiling |
| A refusal | The safe reply, and no crash |
| Valid JSON citing an article that wasn't retrieved | The business rule catches it, and one repair runs |
| A stream error after the 200 | A retry, and the customer saw nothing |
hand_off times out, then succeeds | Exactly one ticket exists |
look_up_order fails the same way every time | The run stops after two tries and hands off |
| The fallback returns a permanent 400 | The tier is skipped, never retried, and an alert fires |
| A coverage score of 0.41 | No answer is sent, and the packet carries the draft |
The SDK's own retries would hide some of these. A stubbed 529 that the SDK retries twice before your code sees it tests the SDK, which is one more reason section 4 set max_retries=0. The other option is to inject the fault below the SDK, at the HTTP layer, so the SDK's behavior is part of what gets tested.
You don't have to write these tests by hand. Describe the matrix and the rules to an AI coding tool:
Write pytest tests for complete() in llm.py. Replace both provider clients with a stub that returns a scripted sequence of responses, and make any real network call fail the test. Cover every row of this fault matrix: [paste the fault matrix]. For each row, check the number of attempts, which tier answered, what the customer received, and how many tickets hand_off created. Use a fake clock so the timeout tests run instantly. Don't change llm.py to make a test pass. Report the failing tests instead.
Then read what it wrote, because the tests are only as good as the outcomes they expect. A test that expects a retry on the spend-limit 429 passes against a broken harness. The last line of the prompt is there because coding tools will sometimes edit the code under test until the tests go green.
Record real responses and replay them
A stub returns what you thought to script, and real responses have quirks you didn't think of, like a code fence around the JSON. So record a batch of real responses once as files, and replay them in the test suite. For Python, pytest-recording does this. By default it blocks network calls during tests, and it can strip the authorization header so your API key doesn't end up in the saved files (pytest-recording). Record them again whenever the prompt or the model changes, because the saved responses came from the old ones.
Run the eval more than once
Section 8's consistency number comes from here. Run every test case in the eval five times and report pass@1 and pass^5 side by side. Then list the test cases whose result changes between runs, because that list is where to look first. A question that passes three times and fails twice often has two similar articles behind it, and which one search puts first decides the answer. Start each run from a clean state, so nothing left over from one run affects the next. Anthropic's advice is that "Each trial should be 'isolated' by starting from a clean environment."
Drill what only production has
Fault tests show that your code handles a fault. They can't show that the second provider's account still works, or that the support team answers handoffs within the time the safe reply promised. Drills do:
- In a quiet hour, send more traffic to the fallback model for a few minutes and watch its numbers.
- Flip the kill switch and read what customers get.
- Send a few seeded conversations into the handoff queue and time the replies.
OpenAI's own list of fixes after the December 2024 outage included "Fault injection testing," and after the June 2025 one it planned "regular disaster recovery drills."
The layers from this section, together with the live measurement in section 10, each make more of the system real, and each costs more to run:

Write the fault tests before the feature ships
Whether evals should come before the feature is argued both ways. Anthropic recommends "eval-driven development: build evals to define planned capabilities before agents can fulfill them," and OpenAI's evaluation guide says the same. Hamel Husain and Shreya Shankar answer the question with "Generally no," and advise, "Write evaluators for errors you discover, not errors you imagine," with an exception for requirements you know up front (Husain and Shankar, evals FAQ).
For a harness, both hold. The list of faults a model call can hit is short and known, so those tests get written before February. The evals for the content of answers come from reading real traces and labeling what went wrong, the way M1 and M2 built them.
Reliability in production
Measure reliability on real traffic, including the failures that raise no errors
The Tuesday prompt edit shortened the answers, and a few of the shorter ones dropped the line about the 48-hour cutoff for changes, or stated the wrong one. The error rate stayed flat, and so did latency. Complaints about wrong cutoffs went up over the next two days, and nothing alerted, because nothing the dashboard measured had changed.
Google's SRE book counts a request as an error when it fails "explicitly (e.g., HTTP 500s), implicitly (for example, an HTTP 200 success response, but coupled with the wrong content), or by policy" (Google SRE, monitoring distributed systems). The Tuesday answers were the implicit kind. Catching them means judging the content of live answers, on a schedule, with a number someone watches.
Define what counts as a good request
A service level indicator, or SLI, is a measured ratio, the number of good events divided by the total number of events (Google SRE workbook, implementing SLOs). This module defines which events count as good for the support agent. M11 sets the targets for them, called service level objectives (SLOs), and decides which ones page someone. The agent's SLIs, computed from the traces M2 set up:
Availability means the customer got a usable answer inside the 20-second deadline, from either the primary or the approved fallback, or got a clean handoff. An error or a timeout that reaches the customer counts as bad.
Latency is the time until the customer sees a reply, reported at p95 and p99, which is the tail M2 said to watch.
Validity on the first try is the share of answers that pass section 5's checks without needing a repair.
First-try success and success after retries are section 4's pair, tracked separately for each place the code calls a model.
Fallback and degraded share covers the conversations served by the fallback beyond its planned 5%, and separately the ones that got the degraded mode.
Escalation rate and coverage are section 7's numbers, with the escalation rate broken down by which trigger fired.
And the quality score is the share of sampled answers a validated judge passes, which the next section covers in full.
Consistency, section 8's pass^k, is measured offline on the eval, and cost per conversation belongs to M12.
Published examples of these for LLM features are still rare. One is Honeycomb's, from 2023, for a feature that turned questions into queries. It counted a request as bad if any step failed, from the call to OpenAI through running the query the model wrote. That's a validity SLI. They set the target at 75% of requests over seven days, raised it to 80% once their fixes landed, and sent a Slack message when they were four hours from using up the allowed failures (Honeycomb).
Score a sample of live answers
The quality score comes from online evaluation. Take a sample of real conversations, and after the reply has gone out, run a judge over each one against criteria that don't need a known right answer. For the support agent, one criterion is whether the cited article supports every claim in the answer, and another is whether a cutoff or a price in the answer matches the article. Braintrust's documentation suggests sampling 1 to 10% of traffic for high-volume applications and 50 to 100% for low-volume or critical ones (Braintrust, as of September 2026). The flower company samples 5%, and 100% of a slice while it's investigating one.
The judge is the model-based check from M1, and M1's rule applies, measure it against human labels before you trust it. Be careful with the number you get, because it depends on how it's computed. In the MT-Bench study, GPT-4's agreement with human experts was 85% when the comparisons the experts called a tie were left out, higher than the 81% agreement between humans. With ties counted, the figures were 66% and 63% (Zheng et al.). A person labels a fresh sample every week, so a change in the judge shows up as the judge and the people drifting apart.
A second source of quality numbers is a golden-prompt probe, a fixed list of questions with known answers sent through the production path every 15 minutes and tagged with the model that answered. The online judge sees whatever customers asked, and the probe asks the same questions every time, so a change in its score points at the system and away from the traffic.
Anthropic's September 2025 postmortem shows why both matter. A routing bug sent requests to the wrong servers, and in the worst hour 16% of Sonnet 4 requests were affected. About 30% of Claude Code users who made requests in that period had at least one message routed wrong. Anthropic wrote that "The evaluations we ran simply didn't capture the degradation users were reporting, in part because Claude often recovers well from isolated mistakes," and committed to running evaluations "continuously on true production systems" (Anthropic, September 2025).
User feedback, like thumbs up and down, is a signal with a known bias. In April 2025, OpenAI shipped a GPT-4o update that made the model sycophantic. Its offline evaluations "generally looked good," and the A/B tests suggested the users who tried it liked it. OpenAI's write-up says that "User feedback in particular can sometimes favor more agreeable responses," and that some expert testers had said the model's behavior "felt" slightly off. OpenAI launched anyway, called that "the wrong call," and committed to "blocking launches based on proxy measurements or qualitative signals, even when metrics like A/B testing look good" (OpenAI, May 2025). Count the thumbs and read them next to the judge scores. Treat a report from someone close to the product that the answers feel off as a signal too, the way OpenAI now does. M11 covers the launch and the rollback side of that story.

Know where quiet failures come from
Each source of quiet failure gets caught a different way:
The provider may have changed something. Anthropic's documentation says model weights are fixed for a given model id, while the serving infrastructure around them, like the request router and the sampling logic, can change and cause minor differences in behavior (Anthropic, model ids and versions, as of September 2026). Infrastructure bugs like the September 2025 one belong in this category too, and the golden-prompt probe is what catches them.
The provider may be retiring the model. Retirements get announced months ahead, so they belong on a calendar, and the eval gets rerun against the replacement well before the date arrives, which is the model swap M4 described. M12 covers the migration checklist.
You may have changed something yourself. A prompt, a config setting, the retrieval index, a harness default. The Tuesday edit and the changes in Anthropic's April 2026 postmortem are both this kind. Anthropic's own fix was to "run a broad suite of per-model evals for every system prompt change," which is the eval gate M12 builds into the release process.
The customers may have changed. New products, seasonal questions, a language the agent rarely sees. Breaking the quality score down by the labels M2 put on every trace, like question type, shows one slice getting worse while the average holds steady.
The knowledge may have changed, since help-center articles get rewritten. M6's sync keeps the agent's copy current, and M19 covers checking that retrieval still finds the right passage afterwards.
Or the judge may have changed, which is the awkward one. The judge runs on a model too, and that model gets updated underneath you. The weekly human labels are what catch it.
Stop the damage first
When the quality score drops, the first job is to stop the damage, and finding the cause comes second. The harness levers, most easily undone first:
- 01Revert the prompt or configuration to the last version that scored well.
- 02Pin the previous model snapshot, if the provider changed something.
- 03Route traffic to a fallback tier that passes the eval.
- 04Flip the kill switch to the degraded mode or to a person.
For the Tuesday edit, the first lever was enough, because the old prompt version was saved and the change was one line. Then comes the slower work. Walk the failing traces back to the first wrong step the way M2 described, add the failing conversations to the M1 dataset so the edit can't pass again, and write up what happened. M11 covers running the incident itself, with who does what and the postmortem.
| Failure | How the code sees it | What the harness does |
|---|---|---|
| Answers worse, errors flat | Judge score on sampled answers | Revert the prompt version |
| Provider behavior changed | The golden-prompt probe drops | Pin the snapshot, or fall back |
| One type of question worse | Quality score by trace label | Fix that slice, add test cases |
| The judge itself changed | Judge and weekly labels diverge | Recheck the judge |
Putting it together
Putting it together
The support agent passed its eval in January and still failed in February. The eval measured the model and the prompt, and most of what went wrong happened in the code around them. That code is the harness, and this module built it one failure at a time, starting from what the code can see.
For the loud failures, every decision was made before it was needed. A failure that can succeed later gets retried in one place, with waits that grow and spread out, and all of it fits inside a deadline set by how long a customer will wait. An answer gets checked in code, cheapest check first, before anything acts on it. It gets one repair before it's thrown away. When a provider is down, the chain moves to a model that has passed the same eval and already carries 5% of the traffic, and below that to a degraded reply that needs no model at all. When the evidence for an answer is thin, the agent declines and hands the conversation over with everything the person needs to decide. Every run has limits it can't pass.
The quiet failures needed measurement. Every request is good or bad by a written definition, and a judge reads a sample of live answers, with a person checking the judge every week. With that in place, the Tuesday edit shows up on Wednesday afternoon, and the fix is the easiest lever to undo, going back one prompt version.
The finished table has its last column filled in, the number that counts each failure:
| Failure | How the code sees it | What the harness does | What counts it |
|---|---|---|---|
| One model call | |||
| Overloaded or rate limited | 503, 529, or a rate-limit 429 | Retry up to twice with jitter, then the next tier | First-try success, availability |
| Out of credits, or at a spend cap | A 429 or 400 naming spend or credits | No retry. Safe reply, alert the owner | Availability, fault test |
| Provider stuck | No first token in 4 s | Retry once, then the next tier | p95 and p99 latency |
hand_off timed out | A timeout on a tool that writes | Retry with the same idempotency key | Duplicate tickets (should be zero) |
| The output | |||
| Cut off | Stop reason is max_tokens | Ask again once with a higher ceiling | Validity on the first try |
| Cites an article search didn't return, or a number the article doesn't have | Business rules | Repair once, then the safe reply | Validity on the first try |
| The path | |||
| Provider down for hours | Retries exhausted, cooldown on | Next model tier (Sonnet 5) | Fallback share, availability |
| Fallback rejects the request | A permanent 400 from that tier | Skip the tier, alert a person | Alert count, fault test |
| Every model tier down | All tiers failed | Degraded mode | Degraded share |
| Model up, answers bad | A person decides | Kill switch | Degraded share |
| Whether to answer | |||
| No article covers the question | Coverage check under 0.8 | Decline, hand off with the packet | Coverage, escalation rate |
| Reviewers approve by habit | Override rate near zero | Seed known-wrong drafts, show the evidence | Override rate, seeded catches |
| The run | |||
| The same call again and again | Same tool and arguments, twice | Stop, tell the customer, hand off | Stop-condition rate |
| Answers that change between runs | pass^5 well under pass@1 | Fix the test cases that flip | pass^5 (offline) |
| Real traffic | |||
| Answers worse, errors flat | Judge score on sampled answers | Revert the prompt version | Quality score |
| Provider behavior changed | The golden-prompt probe drops | Pin the snapshot, or fall back | Probe score |
Two questions come up in almost every design review of an AI feature, and in most interviews about one. The first is what the fallback is, and the answer is the chain in section 6, where every tier has an eval score and runs on a schedule. The second is what happens when confidence is low, and the answer is the coverage check and the triggers in section 7, with a cutoff set on the dataset and a packet the person can act on.
M10 comes back to the same mechanisms with an attacker writing the input on purpose. M11 sets targets for the SLIs from section 10 and runs the incidents they trigger, and it works out whether the fallback can carry all of the traffic at once. M12 prices the retries and the fallback and makes the eval gate a step in every release. After the rails, each module that adds a capability to the agent fills in its own copy of this table.
Checkpoint · recall · 5 questions
What the module said
- 01
Which failures should the harness retry?
- 02
Why does backoff need jitter?
- 03
What's the difference between a timeout and a deadline?
- 04
What makes a fallback into failover?
- 05
Why does a support agent get measured with pass^k?
0 / 5 answered
Checkpoint · understanding · 5 questions
Reason it through
- 01
The code tries each call up to three times, and the SDK is still on its default of two retries. The provider starts failing. How many requests can one turn send, and what's the fix?
- 02
A call fails with a 429 that has no retry-after header, and the message says the organization's spend limit was reached. What should the harness do?
- 03
A teammate proposes asking the model to rate its confidence from 0 to 100 and declining below 90. Why is that a weak plan?
- 04
In Valentine's week the escalation rate goes from 10% to 30%. What breaks first, and where do you look?
- 05
On Tuesday a prompt edit shortened the answers. Errors and latency stayed flat, and complaints about wrong cutoffs grew. What would have caught it, and what's the first move now?
0 / 5 answered
Checkpoint · debugging · 4 questions
Debug it
- 01
The dashboard shows the answer step at 99% success, but p95 latency has tripled. The traces show two retries on most calls. What's going on?
- 02
During an outage the fallback tier fires, and every fallback call comes back with a 400. What went wrong, and how should the harness have handled it?
- 03
After a hand_off call timed out and was retried, the customer had two tickets. The code sends an idempotency key made with uuid4() on each attempt. What's wrong?
- 04
The agent called look_up_order 40 times with the same wrong order number before the customer left. The framework's turn limit never fired. Why not?
0 / 4 answered
Go deeper
Anthropic: An update on recent Claude Code quality reports (April 2026) · AWS Architecture Blog: Exponential backoff and jitter · AWS Builders' Library: Avoiding fallback in distributed systems · Anthropic: Demystifying evals for AI agents · OpenAI: A practical guide to building agents · Google SRE Workbook: Implementing SLOs
That's the last one written so far
Pick your next module from the board.
