MLGuerrillaStart with M1 →
Free · in beta·intermediate·M15·34 min read·Prereq: Harnessing & Reliability (M9), whose validation ladder stops exactly where this module starts. Classification (M13) supplies the thresholds and the confusion matrix, and Routing (M14) supplies the cascade this module puts a price on.

Verification

The capability

What a verifier is

A verifier reads a finished output and returns a verdict on whether that output can be used, with a number attached saying how sure it is.

The word carrying the weight in that sentence is "finished". Everything M13 and M14 built ran before the work happened, reading an incoming request and deciding something about it. A verifier runs afterwards, on an answer that already exists, and what it decides is whether your code should act on that answer. It sits in the one place in the whole system where there is something concrete in front of you, which turns out to matter a great deal.

What it reads is the output together with whatever that output was supposed to be true about. For a support reply that means the reply and the knowledge-base articles search returned on this turn. For generated SQL it means the query and the schema. For a summary it means the summary and the document. A verifier handed the output on its own, with nothing to compare it against, is doing something much weaker, and section 4 is entirely about why.

What comes back is a verdict your code branches on, which in practice is one of four things. Ship it, repair it, send it somewhere better, or give it to a person. Next to the verdict there is a score, and the score is what lets you put a threshold on the decision the way M13 did, so that clear passes go straight out and the uncertain middle goes somewhere you have decided in advance.

A finished answer flows into a verifier along with three other inputs: the source passages the answer was supposed to rest on, a rubric listing the criteria, and the request that started it. The verifier is drawn as one box with a label reading pass or fail and a score of 0.34. Four arrows leave it, going to ship it, repair it, escalate to a bigger model, and hand to a person, with the score bands that select each one marked on the arrows. A note underneath says the verifier runs after generation and before anything acts on the output.
The answer alone is not enough. What makes a check worth anything is the material next to it.

What a verifier cannot do is tell you an answer is right. It tells you the answer survived the checks you wrote, which is a smaller claim, and confusing the two is how teams end up with a green dashboard and angry customers. The gap between those two claims is the subject of this module.

Why this is a capability of its own

M9 already built checks and ran them in production, so it is fair to ask what is left. The answer is that M9's checks all had one thing in common, which is that your own code could settle them. Did the JSON parse. Is this article id one search returned. Does this order number match what the lookup came back with. Every one of those is a question your code knows the answer to already, and none of them touch whether the reply is correct.

Correctness needs something else, and that something else costs money, has its own error rate, and can be wrong in two directions that cost you completely different amounts. That combination is what makes verification a component you design and measure, the same way M13 made you design and measure a classifier.

There is also the matter of how much of the failure it accounts for. Cemri and colleagues built a taxonomy of agent failures by hand from 150 execution traces, with two annotators agreeing at κ = 0.88, and then annotated 1,642 traces across seven frameworks with it. Quality control of the final output came to 23.5% of all failures, split into premature termination at 6.20%, no or incomplete verification at 8.20%, and incorrect verification at 9.10% (Cemri et al., 2025). Close to a quarter of everything that broke in those systems broke at the step that was supposed to catch things breaking.

Where it shows up

Verification is the least visible part of a production AI system and it is running in more places than most engineers realise.

  • A test suite executing code a model just wrote, which is the strongest verifier in this whole module.
  • A grounding check on a retrieval answer, asking whether the retrieved passages support what the reply claims.
  • A moderation pass on generated copy before it reaches a publishing queue.
  • The check in the middle of a cascade, which M14 assumed was perfect and this module prices properly.
  • A reward model scoring candidate answers during training, which is M21's subject.
  • A sampling gate that pulls a percentage of live traffic into a human review queue.

The same component sits behind all of those, with a different source of truth wired into it.

Where this starts

The answer that passed every rule and was still wrong

M9 left the flower company's support agent with a validation ladder and two fixed bugs. The answer that had been cut off at max_tokens now gets one retry with a higher ceiling. The answer that cited article 212 when search had returned 118 and 140 now fails a business rule, because the code compares every cited article id against the ids search returned.

In March a different ticket arrives. A customer writes that her bouquet was delivered yesterday and half the roses are brown, and asks what can be done. Search returns article 305, the damage and replacement policy, and the agent writes back:

“I'm sorry your flowers arrived in that condition. Our policy gives you 30 days to report damage, so you're well within the window. Reply here with a photo of the arrangement and we'll get a replacement sent out.”

Every check in M9's ladder passes this answer. The stop reason is stop, so nothing was truncated. It parses. It matches the schema. Article 305 is one search returned, so the citation rule passes. The number 30 days appears in article 305, so the numbers rule passes. The needs_human flag is false, which M7's handoff rules allow for a damage claim under $50. Nothing is wrong with the shape of this answer anywhere.

The answer is wrong anyway. Article 305 gives customers 30 days to request a refund and 48 hours from delivery to report damage for a replacement. The model took a real number out of the article it did cite and attached it to the wrong deadline. The customer reads the reply, feels reassured, takes her photo a week later, and by then the replacement window closed five days ago.

The March answer moving down M9's six-rung validation ladder with a tick on every rung. Stop reason is stop, it parses, the schema matches, article 305 is one search returned, the number 30 appears in article 305, and the needs-human flag agrees with the handoff rules. At the bottom the answer ships. To the right, article 305 is shown with two lines highlighted in terracotta, one reading refunds may be requested within 30 days of delivery and one reading damage must be reported within 48 hours of delivery, with an arrow showing the answer picked up the first number and used it for the second rule.
Every rule asks whether the number is in the source. None of them asks what the number is about.

The fix depends on being exact about what went wrong here. The numbers rule checks that a number in the answer appears somewhere in a cited article. It has no way to check that the number is about the thing the answer says it is about, because that would require understanding both the article and the sentence, and a regular expression does neither. This is the shape of every failure that gets past M9's ladder. The output is well formed, the sources are real, the facts are drawn from them, and the claim is still false.

The boundary

Where M9's ladder stops

Two ordinary names for this distinction get used interchangeably in the industry, and this module will keep them apart.

Validation asks whether an output is well formed and consistent with what your code already knows. It runs in microseconds, costs nothing, and gives an answer that is never in doubt. Either the JSON parsed or it did not.

Verification asks whether the task was achieved. It needs a source of truth that lives outside your code, it costs something every time it runs, and its answer is a judgement that can be wrong.

The same answer, two questions
QuestionWho settles itCostCan the answer be wrong
Did it parseYour codefreeNo
Is every cited article one search returnedYour codefreeNo
Does every number appear in a cited articleYour codefreeNo
Is the 30-day figure the right deadline for damageSomething outside your codea call, or a personYes

The first three rows are M9's. The fourth row is this module, and the right-hand column is what makes it a different job. Once the answer to your check can itself be wrong, you have added a component with an error rate to a system you were trying to make more reliable, and everything downstream has to account for that.

M9's ladder had six rungs and explicitly stopped at the fifth, the model-based check, saying that building one and measuring it against human labels belonged here. That is where this module picks up. Rungs one through four stay exactly as M9 built them and they run first, because they are free and certain, and there is no sense paying for a judgement on an answer that does not parse.

Grounding

Check the answer against the chunk it came from

The March answer went wrong in a specific way. It made a claim about a deadline, the passage that claim was supposed to rest on was sitting right there in the retrieval result, and nothing in the system compared the two.

So compare them. Take the passage search returned, take the claim the reply makes, and ask whether the passage supports the claim. That shape has a name in the literature, where the passage is the premise and the claim is the hypothesis, and it is the shape every grounding checker in production uses. RAGAS computes faithfulness this way. Vectara's HHEM takes a premise and a hypothesis and returns a score. Patronus Lynx does the same job with a larger model. None of them asks a model "is this correct" with nothing to check against.

What is left to decide is how you compare, and there are three families to choose from.

Three ways to compare, and what each one can see

Word overlap. ROUGE-L scores the longest common subsequence between two pieces of text, and BLEU scores n-gram precision. Both are counting words in order. Neither has any access to meaning, and the industry habit of calling ROUGE a semantic metric leads people to trust it for exactly the job it cannot do.

Embedding similarity. Encode the passage and the claim with a sentence embedding model and take the cosine between them. This one does see meaning, in the sense that a paraphrase lands near its original. What it measures is whether two pieces of text are about the same thing, and "about the same thing" includes flat contradictions of each other.

Entailment. A natural language inference model reads the premise and the hypothesis together and returns three probabilities, for entailment, neutral and contradiction. This is the purpose-built tool, it is small, and MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli at 184M parameters runs on a laptop.

All three run on article 305 below, against four claims chosen so that truth and word overlap come apart. Every number below comes from research/M15-grounding-checks.py, which downloads both models and reproduces the table.

Article 305 against four claims, scored three ways
The claimTrueROUGE-LCosineEntailment
The March answer, "30 days to report damage"no0.2500.4130.834
"Refunds may be requested within 30 days of delivery"yes1.0000.5500.993
"Let us know about a spoiled bouquet inside two days"yes0.1760.3540.994
"Damaged orders get store credit at twice the price"no0.3330.5100.000

At sensible thresholds, word overlap gets three of the four right, embedding similarity gets two, and entailment gets three. Each of them fails on a different row, and the first thing to take from that is that no single number in the table is a check you could ship.

The third row is what finishes off the first two families. That claim is true, correctly paraphrased from the 48-hour rule, and it scores 0.176 on ROUGE-L and 0.354 on cosine. Both are below any threshold you would pick. A pipeline gated on either one escalates a correct answer, and it does so precisely when the model has done the good thing of answering in its own words.

The first row is worse. The entailment model gives a false claim 0.834 for entailment. The tool built for this job reads "our policy gives you 30 days to report damage" next to a passage that says damage must be reported in 48 hours, and reports that the passage supports it.

How you cut the answer up decides more than which check you pick

The natural conclusion from that table is that you need a stronger model. You do not. You need to hand the same model a different thing.

The same model and the same passage, cut up differently
What the model is givenTrueEntailNeutralContradict
The reply as written, against the whole passageno0.8340.1550.011
One atomic claim, against the whole passageno0.0010.0020.997
One atomic claim, against one passage sentenceno0.0020.0030.996
The true version of that claim, for contrastyes0.9870.0110.002

Row one and row two differ in nothing except how the answer was cut. Row one hands over the reply as the customer would read it, complete with "so you are well within the window". Row two reduces it to the single proposition the reply asserts, that damage must be reported within 30 days of delivery. Same 184M model, same passage, and the verdict goes from 83% entailment to 99.7% contradiction.

What happens is that a conversational sentence carries more than one thing at once. It has a topic, a reassurance, a hedge, and somewhere inside it a claim about a deadline. The model scores the whole bundle against a passage that is also about damage and also mentions 30 days, and lands on "yes, this is broadly supported". Strip the sentence down to the proposition and there is nothing left to be broadly right about.

Row three is worth one note. Narrowing the premise from the whole passage to the one relevant sentence changed almost nothing, 0.997 to 0.996. Retrieval precision was not the problem. The answer's shape was.

Two panels. The left panel shows four claims scored by three checks against article 305, as a grid of numbers with the wrong verdicts marked in terracotta: the March answer scoring 0.250 on ROUGE-L, 0.413 on cosine and 0.834 on entailment while being false, and the correctly paraphrased damage rule scoring 0.176 and 0.354 while being true. The right panel shows the same entailment model given the same passage four different ways, with stacked bars for entailment, neutral and contradiction: the reply as written comes out 83.4% entailment, one atomic claim against the whole passage comes out 99.7% contradiction, one atomic claim against one sentence comes out 99.6% contradiction, and the true version of the claim comes out 98.7% entailment. A note underneath says the model and the passage are identical across all four and only the cut changed.
Left, no single score separates true from false. Right, the same model flips from 83% support to 99.7% contradiction on the strength of how the answer was cut up.

What counts as one claim

Splitting a reply into claims only helps if the splitter and the people reading its output agree on what one claim is, because a loose split hands the entailment model the same kind of bundle it scored at 0.834. The definition this module uses is that an atomic claim is one statement that a single passage could confirm or contradict by itself, with its subject, what it says about that subject, and any condition that limits it all written inside the claim. The condition is the part a policy number depends on, like which kind of request it applies to or when the clock starts. Take the condition off "within 30 days of delivery" and the sentence is true for refunds and false for damage, so it can no longer be checked at all.

The March reply, one row per claim
ClaimChecked againstVerdict
Damage to a delivered order can be reported up to 30 days after deliveryarticle 305contradicted, the article says 48 hours
This customer is still inside the window for reporting damagearticle 305 and the delivery date from the order lookupfails, because it rests on the claim above
A damage report needs a photo of the arrangementarticle 305supported
A damaged order gets a replacementarticle 305supported

The apology at the start of the reply is missing from that table on purpose, because it asserts nothing a passage could confirm. The instruction to reply with a photo is in it, because telling the customer to send one implies the policy asks for one, and that is checkable. The second row is the one to look at twice. "You're well within the window" contains no number, so M9's numbers rule would never look at it, and it is the sentence the customer acts on. It needs two pieces of evidence, the policy and the delivery date, which is why a claim can name more than one source.

Every claim is also rewritten so it stands on its own, with "it" and "the window" replaced by the nouns they point at, since the entailment model sees the claim without the reply around it.

The splitter is itself a model call, so it can drop the one claim that mattered or quietly change a number while rewriting it. A free check covers most of that. Every number and date in the reply has to appear in at least one claim, and every number in a claim has to appear in the reply, which your code can test with the same kind of rule M9 used for article ids.

Ask your AI coding tool

Write split_claims(reply: str) -> list[dict] in verify/claims.py. It calls a model to split a support reply into atomic claims. Each claim is one statement a single passage could confirm or contradict, rewritten to stand alone (no pronouns), and it keeps any condition that limits it, such as the request type or when a deadline starts. Skip apologies and greetings. Keep an instruction to the customer when it implies a policy. Return {"claim": ..., "needs": [...]}, where needs lists the evidence types required, like kb_article or order_lookup. Then add a plain-Python check that every number and date in the reply appears in some claim and every number in a claim appears in the reply, and fail loudly when either direction is violated. Include pytest tests built from the March reply in this module.

What that buys you in the pipeline

The check that works on this example is four steps, and none of them needs a frontier model.

  1. 01Split the reply into atomic claims, one proposition each, by the definition above.
  2. 02Check that every passage a claim will be scored against is current and allowed to state policy, which the next subsection covers.
  3. 03For every claim, run entailment against the passages search returned this turn.
  4. 04Fail the answer when any claim comes back as contradiction, and send anything that lands in neutral to the next rung up.

Step one is the step people skip and it is the one doing the work. It is also the step that costs something, since splitting a reply into propositions is itself a model call, which is why a grounding check is rarely as cheap as the 184M number suggests. RAGAS computes faithfulness by decomposing the answer into statements and verifying each against the context, and FActScore built the same decomposition into atomic facts for biography generation. Both papers landed on decomposition for the reason this table shows.

Neutral deserves its own branch, and folding it into fail throws away what it tells you. An atomic claim the passage neither supports nor contradicts is usually a claim about something search did not retrieve, which is a retrieval problem showing up at the verification step, and M19 is where that gets fixed.

Check the passage before you let it support anything

Entailment tells you whether a passage supports a claim, and it has no way to tell you whether the passage is still true. Suppose that in June the flower company moves its damage window from 48 hours to 72 and publishes a new version of article 305, but the search index keeps serving the old version for a week. The agent tells a customer they have 48 hours, citing article 305, and every step so far passes. The citation is one search returned, the number is in the passage, and the atomic claim scores as entailed, because it is entailed by the passage the agent was given. The customer is still told the wrong thing, and this time the mistake turns away customers who were still owed a replacement, because anyone writing in between hour 48 and hour 72 is told they are too late. The June change is invented, like the rest of the company.

Nothing about the claim can catch that. The check has to run on the evidence, before any claim is scored against it, and most of it is free if the knowledge base stores a few fields next to each article.

  • The passage comes from the current version of its article, meaning its version matches the latest one in the knowledge base and its superseded_by field is empty.
  • The article's effective_from date is on or before the date the answer is about, which for a damage claim is the delivery date.
  • The source type is one allowed to state policy. A policy article can. An old support ticket or a marketing page that quotes a number from last year cannot, even when search ranks it first.
  • No two retrieved passages disagree about the same condition. When one says 48 hours and another says 72, the verifier returns unclear, which leaves the choice to a person, and flags the conflict for whoever owns the knowledge base.

A claim whose only support is a passage that fails one of those checks counts as unsupported, and it gets logged with the reason, stale or conflicting evidence, since that reason decides what happens next. Regenerating the reply cannot fix it, because the generator would read the same stale passage again, so section 10 sends it to a person. Section 11 covers the same policy change from the other side, where the judge's rubric is the thing that went out of date.

If the knowledge base has no version or date fields, adding them belongs to M19, and until then the honest verdict on any policy number is that the evidence behind it has not been checked.

The shortcut of asking the model whether it is right

There is a shortcut that skips all of the above, which is to hand the model its own answer and ask whether it is correct. It comes up in enough blog posts to be worth pricing, and it does not work.

Chen and colleagues put a known-wrong claim inside a model's own reasoning trace, asked it to audit, and counted how often it said out loud that the claim was wrong (Chen et al., 2026). Thirty paired tasks per setting at temperature zero, scored by a frozen judge that matched hand labels at κ = 1.0.

How often a model flags a known error sitting in its own reasoning, maths tasks
ModelCaught it
gpt-oss-20B, reasoning on76.7%
Gemini 2.5 Flash-Lite66.7%
Claude Sonnet 453.3%
Qwen3-30B, reasoning on53.3%
GPT-4o40.0%
Gemma-3-12B30.0%
Phi-4-14B20.0%
Qwen2.5-72B16.7%
Llama-3.3-70B0.0%

Reasoning-trained models do noticeably better, and the best of them still misses a quarter of errors that are sitting in front of it with instructions to go and find them. Compare that against the 99.7% the 184M entailment model returned once it was handed an atomic claim and a passage, and the ranking is not close.

The same study explains part of why. Re-present the byte-identical claim as a user message or a tool response, leaving it out of the model's own reasoning, and the correction rate rises by 23 to 93 percentage points, taking GPT-4o from 40% to 90%. Their reading of it is that the failure "is largely an artifact of how the claim is role-labeled in the chat template". That is the same lesson as the decomposition table arriving by a different route, which is that the model's verdict depends heavily on the packaging of what you hand it.

Two older results fill in the mechanism and both are worth reading in full. Huang and colleagues showed that self-correction improves when an oracle label decides when to stop and degrades without one, with GPT-4 dropping from 95.5% to 89.0% on GSM8K while going from one call to five (Huang et al., 2023). Stechly and colleagues had GPT-4 verify its own graph colourings and found it stopped itself 4 times in 100 and was wrong every time (Stechly et al., 2023). Both are GPT-4 in 2023, and the 2026 numbers above are the ones to quote.

None of this says a model cannot verify. Section 5's fourth rung is a model doing exactly that, and section 7 builds one. What it says is that a model verifies when you give it something to verify against and hand it over in a shape it can act on, and the difference between doing that and not doing it is 83% support against 99.7% contradiction on the same sentence.

This does not get easier as models improve

The reasonable thing to say at this point is that models keep improving, so this is a problem that solves itself. Two 2026 results say the opposite, and one of them says the problem gets worse.

The first is that getting better at a task does not make a model better at checking that task. Chen and colleagues tracked both abilities through training and found "a persistent capability asymmetry between generation and self-verification", where "improving generation does not lead to corresponding improvements in self-verification, even on the same task" (Chen et al., 2026).

The second is worse. Zhou and colleagues studied verifier behaviour across 12 benchmarks, 14 open models from 2B to 72B, and GPT-4o, and found that "weak generators produce errors that are easier to detect than strong generators" (Zhou et al., ICLR 2026). The stronger your generator, the subtler its remaining mistakes, and the harder they are for any verifier to catch.

Frontier scale does not rescue it either. Naik and colleagues had GPT-5.2 and Gemini 3.1 Pro grade olympiad proofs against human-graded ground truth and topped out at 87.7% balanced accuracy, with a 27.3-point swing on GPT-5.2 from changing nothing but the prompt (Naik et al., COLM 2026). A 35B open model with a tuned ensemble of twelve prompts matched them. Which is the same finding as this whole section, one more time: what you hand the verifier decides more than which verifier you picked.

A judge also prefers its own writing

There is one more reason not to send the answer back where it came from, and it survives even when you use a separate model call with a clean context.

Panickssery and colleagues found that language models score their own outputs higher than other people's when human annotators rate the two as equal, and then went looking for why (Panickssery et al., 2024). It turns out the models can tell. GPT-4 was 73.5% accurate at picking out its own summaries from a set that also contained two other models' work and human writing. Fine-tuning GPT-3.5 and Llama 2 on 500 examples pushed self-recognition above 90%, and the strength of the self-preference tracked the self-recognition linearly as it went up.

The practical version of that finding is one line. Use a different model family for the judge than the one that wrote the answer, and when you cannot, stop treating the judge's score as neutral.

What a check is worth

One idea sits underneath every result in this section, and the rest of the module is built on it.

A check is worth exactly what it knows that the writer of the answer did not know, and it can only use what you hand it in a shape it can act on.

The first half explains the easy cases. Business rules work for free because your own code knows which articles search returned this turn and the model does not reliably. Running generated code works because the interpreter is a source of truth no amount of model capability substitutes for. Asking the model to review its own reasoning fails because there is nothing there it did not already have.

The second half is the one this section added, and it is the half people skip. The entailment model had the passage in front of it both times. The information was identical. All that changed was whether the claim arrived as a conversational sentence or as a proposition, and the verdict moved from 83% support to 99.7% contradiction.

So there are two questions to ask about any proposed check. What does it know that the generator did not, and what shape is it getting that in. When the answer to the first is nothing, the check will cost a call and buy a second opinion from the same mind. When the answer to the second is "whatever the model happened to write", the check will be measuring something other than what you think.

The ladder

Pick the check your output can support

That question sorts the available checks into rungs, and the price of a rung tracks how much it knows.

Rung 0, the rules from M9

Parse, schema, business rules and citation checks are free, instant, certain and already built, so they run first on every output. They will catch more than you expect, because a surprising share of failures turn out to be shape failures, and nothing in this module replaces them.

Rung 1, run the output

If what the model produced is code, SQL, a config file, a shell command, or a request against a real API, then the cheapest and most reliable verifier available to you is an executor, whether that means running the test suite, asking the database to explain the query, validating the config against its own schema, or firing the request at a sandbox.

This is the rung engineers skip, and skipping it is close to always a mistake. An interpreter tells you what the output does when you run it, with no judgement involved, and for that one question nothing further up this ladder gets anywhere near it.

The catch is that "what it does" and "whether it answers the question" are two different questions, and an executor on its own only answers the first. Suppose the flower company gives its operations team an internal assistant that writes SQL, and somebody asks it how many refunds went out in August 2026. It writes this.

the assistant's query, which runs without an error
SELECT COUNT(*) FROM refunds
WHERE refunded_at BETWEEN '2026-08-01' AND '2026-08-31'

The query parses and returns 1,169 without an error. The refunded_at column holds a date and a time, though, and SQLite compares those as text, so '2026-08-31 18:30:00' sorts after '2026-08-31' and every refund issued on August 31 falls outside the range. The right count is 1,215, which means the 46 refunds from August 31 are missing and the answer is 3.8% low. That is small enough that nobody reading the number would doubt it. The numbers come from research/M15-sql-date-range.py, which runs both queries on a month of invented refunds in SQLite.

An executor that only reports "ran without an error" passes this query, because running without an error is all it was checking. What catches it is a test with the answer written down in advance. The same script builds a four-row table with one refund at 23:59:59 on July 31, one at midnight on August 1, one at 18:30 on August 31 and one at midnight on September 1, where the right count for August is 2. The assistant's query returns 1 and fails, and the corrected query, refunded_at >= '2026-08-01' AND refunded_at < '2026-09-01', returns 2 and passes.

That is the rule from section 4 again. The executor knows what the query does, and the fixture knows what the answer should be, which is the one thing the model that wrote the query did not know. For generated SQL that usually means a small table of rows placed on the boundaries of whatever the question filters on, like the first and last second of a month or the day a policy changed, with the expected result stored next to it.

The flower company's support agent does not generate code, so it has no rung 1 available, which is worth saying plainly because it is the reason the rest of the module has to work so hard. When you do have an executor and a few tests with known answers, most of the difficulty in this module goes away.

Rung 2, agreement across samples

Ask the same question several times at a non-zero temperature and compare the answers. Manakul and colleagues built a hallucination detector on exactly this, reasoning that "if an LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts. However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another" (Manakul et al., 2023). It needs no external database and no access to the model's probabilities, which is why it works against a provider API.

Be precise about what it buys, because it fits the rule from the last section awkwardly. Sampling gives you information about the model's stability on this question, which is real and useful. It gives you no facts. A model that is confidently and consistently wrong will produce five agreeing samples and sail through.

It is also the most expensive rung on this list per unit of information. Three samples costs two extra full generations, which for the flower company works out at about $0.0033 an answer. That is more than the $0.0025 a judge call costs in the table below, even though the judge runs on a model with five times Luna's price per token, because a generation carries the whole conversation and a judge reads only the answer and its sources.

Rung 3, a specialist verifier model

Small models trained on one verification job exist, they are cheap, and they are frequently better at that job than a frontier model. Vectara's HHEM-2.1-Open is the clearest example. It is a 0.1B-parameter model built on google/flan-t5-base, released under Apache-2.0, that takes a premise and a hypothesis and returns a score from 0 to 1 for whether the premise supports the hypothesis, with anything under 0.5 read as a hallucination. Its model card reports balanced accuracy of 76.55% on AggreFact-SOTA, 64.42% on RAGTruth-Summ and 74.28% on RAGTruth-QA, and states that it "outperforms both GPT-3.5-Turbo and GPT-4 in all three benchmarks" (Vectara, HHEM-2.1-Open).

A tenth of a billion parameters, running on your own hardware, beating GPT-4 at one narrow task. That is what a trained verifier is worth, and Cobbe's team measured the same effect in 2021 on maths word problems, reporting that "the use of verifiers results in approximately the same performance boost as a 30x model size increase" (Cobbe et al., 2021).

Those balanced-accuracy figures are also the honesty check on this rung. Between a quarter and a third of its verdicts are wrong on those benchmarks, so what you get is a score you threshold, and other open options in this category include Patronus Lynx, Prometheus 2, Atla Selene and Flow-Judge.

Rung 4, a judge

A general model, given the answer, the material the answer was supposed to rest on, and a written set of criteria. This is the most flexible rung, because you can check anything you can describe, and the most expensive and the most biased. Section 7 builds one for the flower company, and it is priced here as one call per answer to Claude Haiku 4.5 that returns a verdict on all four of its criteria, with Sonnet 5 as the stronger judge. The judge comes from a different model family from the GPT-5.6 models that write the replies, for the reason section 7 gives. That call reads about 1,560 tokens, most of them the retrieved passages, and writes about 180, a verdict with a short reason for each criterion.

Rung 5, a person

Everything a person knows, applied to a sample of traffic. At M7's $2.50 a touch, reading 2% of the flower company's 60,000 monthly conversations costs $3,000 a month, which is more than a judge on the frontier model would cost to run on every single answer. People are the only rung that produces the labels the other rungs are measured against, so this line never goes to zero no matter how good the automation gets.

Six rungs in a table, each labelled with what it knows that the generator did not, its cost per answer at the flower company's traffic, and its latency. Rung 0 the rules from M9 knows what your own code did this turn, costs nothing and takes microseconds. Rung 1 running the output knows what the interpreter reports, costs cents of compute and takes milliseconds to seconds, and is marked in terracotta as the strongest rung. Rung 2 agreement across samples knows whether the model is stable, costs 0.0033 dollars an answer and takes three turns. Rung 3 a specialist verifier model knows a decision boundary trained on labelled failures, costs nothing self-hosted and takes 10 to 100 milliseconds. Rung 4 a judge knows whatever you put in its context, costs 0.00246 dollars on Haiku 4.5 or 0.00492 on Sonnet 5 and takes one to four seconds. Rung 5 a person knows everything, costs 2.50 dollars a touch and takes minutes.
The rungs are ordered by what each one knows that the generator did not, and the price follows.
Running each rung on every answer, 60,000 conversations a month
RungCost an answerCost a monthLatency
M9's rules$0$0microseconds
Running the outputcompute onlycompute onlyms to s
Agreement over 3 samples$0.00333$9593 turns
HHEM-2.1-open, self-hosted~$0~$010 to 100 ms
A judge on Haiku 4.5, one call$0.00246$708about 1 s
A judge on Sonnet 5, one call$0.00492$1,4172 to 4 s
A person, on 2% of conversations$3,000minutes

The monthly figures assume 4.8 answers per conversation, which is the mean of M12's turn distribution, and that is the detail that catches teams out. Verification cost scales with turns, so a judge that looked affordable against your conversation count is nearly five times that in practice. The judge also costs more per answer than Luna's own generation, $0.0025 against $0.0017, because the cheapest model outside the generator's family is five times Luna's price per token. At $708 a month it comes to about a fifth of M14's entire routed model spend of $3,750, and adding the person's $3,000 brings the cost of checking the work up to about the cost of doing it.

Images and screen state

Checking something that is not text

Everything so far assumed the thing being checked is a piece of text sitting next to another piece of text. A growing share of the work is not that. A computer-use agent clicks something and has to decide whether the application reached the state it was aiming for, and what it has to work with is a screenshot.

The question has the same shape as before. One state the app was supposed to reach, one state it is in, do they match. What changes is that none of the checks from section 4 accept an image.

The obvious move fails faster here than it did for text

Take two screenshots of the same settings panel and measure how different they are. research/M15-visual-checks.py does that three ways over four pairs, generated from research/fixtures/m15-win.html so the whole thing reproduces.

Four screenshot pairs, scored by how different the images are
What changed between the twoSame statePixels movedPerceptual hashCLIP cosine
The clock ticked over one minuteyes0.01%0.0001.000
Night light was switched onno0.57%0.3050.980
New wallpaper, wider window, new clockyes55.02%0.2730.912
A confirmation dialog openedno48.46%0.3910.907
Two screenshots of the same Windows display-settings panel side by side, before and after the Night light toggle is switched on, labelled 0.57% of pixels differ, with a note that a cosmetic change to the same panel moves 55% of the pixels and changes no state. Below, two tables. The left one scores four screenshot pairs by pixel difference, perceptual hash and CLIP cosine: the clock ticking over is the same state at 0.01%, 0.000 and 1.000; Night light switched on is a different state at 0.57%, 0.305 and 0.980; new wallpaper and a wider window is the same state at 55.02%, 0.273 and 0.912; a dialog opening is a different state at 48.46%, 0.391 and 0.907. Correct verdicts are 2 of 4 for pixels, 3 of 4 for perceptual hash and 2 of 4 for CLIP, and a note says CLIP calls the pair that differs in state more similar than the pair that is identical. The right table asks a 2B VLM four ways: open extraction of the resolution gets 4 of 4, a closed question about that same text gets 2 of 2, a closed question about the Night light toggle gets 2 of 3, a closed question about a dialog being open gets 1 of 2, and the holistic same-or-different question gets 1 of 2, with a note that answering no every time still scored 2 of 3 because two of those three screenshots did have the toggle off. A band across the bottom, headed narrow it first then ask once, gives three numbers from RankGround: a fine-tuned 2B reranker picks the right crop out of a screenshot 88.2% of the time before any VLM sees it, the pipeline reaches 71.0% on ScreenSpot-Pro with an 8B model and one VLM call against 68.4% for a three-call 235B pipeline, and selection error is 19.9% on pure icons against 5.5% on elements carrying text.
A real state change moved 0.57% of the pixels. A change that altered nothing moved 55% of them.

At the thresholds an engineer would plausibly pick, pixel diff gets two of the four right, perceptual hash three, CLIP two.

The middle two rows are the whole problem, and they are worse than the text version. A real state change, the one thing the agent was trying to do, moved 0.57% of the pixels. A change that altered nothing about the application's state moved 55% of them. CLIP scored the pair that differs in state at 0.980 and the pair that is identical in state at 0.912, so on the two rows that matter the metric runs backwards.

Moving the threshold cannot fix this, because any number that answers "how different do these look" is answering a different question from "is this a different state", and on a settings panel the two come apart immediately, because the thing you care about is twenty pixels of a toggle and the thing you do not care about is the whole background.

A VLM reads the screenshot. Whether it answers the question depends on the model.

So ask a model. Before any numbers, separate two jobs that get talked about together and behave nothing alike.

Grounding is picking the thing out. You hand the model a screen and an instruction and it tells you where to click. That is hard, and ScreenSpot-Pro exists to measure how hard, across 23 professional applications on high-resolution displays. Good systems sit in the sixties and seventies. GUI-Lens puts the difficulty in one sentence, that a model "may recognize a requested control without locating it precisely enough for interaction" (Fu et al., 2026).

Verifying is confirming a candidate. You hand the model the element you think is right and the thing you were looking for, and it tells you whether they match. That is the easier half of the same problem, it is the half a check needs, and a cheap model does it well.

The structure that gets you there is published. RankGround narrows a screenshot to a dense set of candidate crops, scores them with a fine-tuned 2B reranker, and spends exactly one VLM call on the top-ranked one (Fan et al., 2026). The reranker picks the right crop 88.2% of the time on its own, and the pipeline reaches 71.0% on ScreenSpot-Pro with an 8B model, beating a three-call pipeline built on a 235B model at 68.4% while running four to twelve times faster. Narrowing first and asking one scoped question beats spending more on the model.

I ran that shape on the computer-use agent from M1, the one that reads a screen and decides what to click. The screen gets broken into candidate UI elements, SigLIP embeds them and ranks them against whatever the agent is looking for, and a cheap model in a frontier family then looks at the top-ranked candidate and says whether it is the right element. Over two thousand images it agreed with my labels on every one.

The cases that needed help were the ones where the target carried no text, and what fixed them was writing a description of the element into the prompt alongside the crop. RankGround measures the same thing from the other side, missing on 19.9% of pure icons against 5.5% of elements with text, so an interface of unlabelled icons is where this gets expensive and where a written description of the target is worth adding. It is also the rule this module keeps arriving at, which is that a check is worth what you hand it.

Read that as a report and not as a benchmark, since it is one person's traffic scored against one person's labels in one domain. What it does support is the narrower claim this subsection is built on. Confirming a candidate is a different job from finding one, and the cheap tier is enough for the confirming.

The cost of the scoped call is the other thing to carry across, because it is not the flagship price. In M12's table the cheap model in a frontier family runs at $0.20 per million input tokens against $4.00 for the flagship, which is the difference between a check you can afford on every step and one you cannot.

None of that is what the run below does. It uses SmolVLM-Instruct at 2B on whole unscoped screenshots, which is the shape a lot of people reach for first, and it is included to show what that shape costs you.

SmolVLM-Instruct, 2B, on the same screenshots, asked four ways
The questionAsked aboutRight
Open extraction, "what resolution is shown"text the interface renders4 of 4
Closed, "is the resolution 1280 by 720"text the interface renders2 of 2
Closed, "is the Night light toggle on"a widget's visual state2 of 3
Closed, "is a dialog open on top"an element being present1 of 2
Holistic, "same state or different"both screenshots at once1 of 2

The model is reading the image. Asked what the window is called it answers "Display Settings", asked the resolution it answers 1920 by 1080, asked what colour the desktop is behind the window it answers green, and it gets the changed resolution right on the one screenshot where it changed. None of that is guesswork.

What it cannot do is answer a yes-or-no question about a widget. Every closed question about the toggle came back "no" whatever the toggle was doing, and every closed question about the dialog came back "yes" whether or not a dialog was there. The holistic question came back "same" for both pairs. A 500M model from the same family behaved identically with the answers flipped, saying "yes" to everything. The model is emitting a constant and the image is not reaching the answer.

That last detail is the one to hold on to, and the toggle row shows why. Answering "no" to every screenshot scored 2 out of 3 there, because two of the three screenshots did have the toggle off. A constant lands somewhere between 50% and the base rate of whichever answer it happens to give, and on any dashboard that is not counting per-class it reads as a working check. It is section 9's superficial check, showing up in an image check.

The gap between that and RankGround's 88.2% comes from what each model is handed, since the reranker gets one crop and one question about it, and this run gets a whole screen and a question about everything on it. The next two parts are about closing that gap on the cheap tier, and neither of them costs more than the call you were already making.

Extract, then compare

The fix follows from which column of that table worked. Stop asking the model for a verdict and ask it for the contents, then do the comparing in code.

  1. 01Ask the VLM open questions that read values off the screen, one value per question.
  2. 02Compare those values against what the state was supposed to be, in your own code or with the entailment check from section 4.
  3. 03Where the interface does not render the answer as text, do not use the screenshot at all.

Steps one and two put you back in text, where everything earlier in this module applies, including thresholds, gold sets and self-consistency. Step three is where most of the accuracy comes from, and it is what the research on this has converged on.

When the answer is not on the screen, stop looking at the screen

Shi and colleagues built a benchmark of 321 GUI task trajectories across ten Ubuntu application categories and split the tasks by where the evidence lives (Shi et al., 2026). Visible-state tasks can be settled from a screenshot. Hidden-state tasks need the application's configuration or a system setting. Artifact tasks need somebody to open the file that was supposedly produced. Their framing of the problem is the sentence to take away, that reliable evaluation "often require[s] access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories".

Their evaluator proposes the conditions a completed task would satisfy and then goes and verifies each one with system, application and GUI tools. Against the same backbone, judging from the before and after screenshots alone scored 78.5%, and going and looking scored 86.9%.

Judging a finished GUI task, screenshots against evidence
BackboneFrom the screenshotsGoing and checking the environmentGainTokens
GPT-5.578.5%86.9%+8.43.05K to 14.9K
GPT-5.476.3%85.4%+9.13.11K to 23.8K
Qwen3.6-35B-A3B78.8%85.0%+6.22.98K to 25.9K

Eight points for roughly five times the tokens, on median. The passive evaluators they compare against range from 47.0% to 78.8%, and the weakest of them reaches 100% precision at 0.6% recall, which is a check that says failure to almost everything and is the same failure as section 9's cheapest row.

This is section 5's first rung, arriving for a different kind of output. Running the output was the strongest check available for text and was rarely available, because a support reply has no interpreter. For screen state it is almost always available, since the application's settings are in a config file, the window tree is exposed by the accessibility API, and the file the agent claimed to save either exists or does not. The screenshot is a picture of the state. The state itself is readable, and reading it is both cheaper and more reliable than asking a model to interpret a picture of it.

What this run does and does not show

The small-model numbers above come from one prompt per question and a handful of synthetic screenshots, which makes them a demonstration and not a measurement, and the frontier numbers at the top of this section are the ones to plan against. Chae and colleagues give the reason to expect the difficulty to persist in some form anyway, having built a dataset of the indivisible perception skills that make up visual comprehension and found that current models "struggle with these tasks, despite being trivial for adult humans" (Chae et al., 2025).

What does not depend on model scale is the ordering. A config file is deterministic, an accessibility tree is deterministic, and a screenshot read by a model is a judgement with an error rate you would then have to measure the way section 8 describes. Given a choice between the two, the choice is not close.

The judge

Building a judge you can put in production

Everything in this section comes out of the evidence in section 4, so none of it is a matter of taste.

Give it the material, and a reference answer where you have one

The single most effective change you can make to a judge is handing it something to compare against. Zheng and colleagues measured this on ten maths questions, comparing two models' answers with the positions swapped, so twenty comparisons, and counting a failure as GPT-4 calling an incorrect answer correct (Zheng et al., 2023):

Same judge, same questions, different material in its context
What the judge was givenFailures out of 20
The default prompt14
The default prompt, asked to reason step by step6
A reference answer3

Fourteen down to three, from putting the answer key in the context. Chain of thought helped too and helped less. For a support agent you rarely have a reference answer to hand at request time, but you always have the retrieved passages, which is the same trick in a weaker form, and it is the difference between asking "is this reply correct" and asking "does this passage support this sentence".

Make it answer a closed question

A judge is a classifier, so M13 applies to it without modification. Give it a fixed list of verdicts and it returns one of them with a score you can threshold, and you get a confusion matrix, per-class precision and recall, and a calibration check. Ask it to rate the answer from 1 to 10 and you get a number whose meaning drifts between prompts, between model versions and between the top and bottom of the scale.

This is also how the provider tooling is shaped. OpenAI's Evals API exposes label_model, which returns one of a fixed set of labels and passes when the label is in passing_labels, alongside score_model, which returns a number. The first one is the one to reach for, for the same reason M13 gave.

Better still, break the rubric into criteria and have the judge return a verdict on each one. Four verdicts give you four signals you can act on differently, and they tell you which criterion failed, which is what you need in the log when you are working out why the escalation rate moved.

The flower company's judge returns all four verdicts from one call. Most of what the judge reads is the passages and the reply, so asking about each criterion in its own call would send those same tokens four times, which would make the check four times as expensive and, run one after another, four times as slow. Section 5's cost and latency are for the single call.

Write every criterion as the condition a good reply meets, so that pass means the same thing on every line. A criterion phrased as a question about a defect, like "does the reply leave a question unanswered", makes yes the bad answer on that line and the good answer on the next one, and sooner or later a line of code maps the wrong yes to ship. Each criterion also says when it fails and when it is unclear, because a judge left to decide that for itself will draw the line in a different place for each criterion.

The four criteria, each written as what a good reply does
CriterionPass whenFail whenUnclear when
groundedevery factual claim in the reply is supported by a passagea passage contradicts a claim, or no passage supports ittwo passages disagree about the claim
policyevery deadline or refund term matches the passage for the same kind of requesta term belongs to another kind of request or contradicts its passageno passage states the term
on_topicthe reply answers the request the customer madethe reply answers a different requestthe customer's request could mean two things
completeevery question in the customer's message gets an answer or a handoff to a persona question is left unansweredit is not clear whether a sentence is a question
judge.py — one call, four success criteria, against the sources
RUBRIC = {
    "grounded": "Every factual claim in the reply is supported by a passage. "
                "Fail if a passage contradicts a claim or none supports it. "
                "Unclear if two passages disagree about the claim.",
    "policy":   "Every deadline or refund term in the reply matches the passage "
                "for the same kind of request. Fail if a term belongs to another "
                "kind of request or contradicts its passage. Unclear if no "
                "passage states the term.",
    "on_topic": "The reply answers the request the customer made. Fail if it "
                "answers a different request. Unclear if the request could mean "
                "two things.",
    "complete": "Every question in the customer's message gets an answer or a "
                "handoff to a person. Fail if one is left unanswered. Unclear if "
                "you cannot tell whether a sentence is a question.",
}
VERDICTS = {"pass", "fail", "unclear"}

def judge(reply: str, passages: list[str], question: str) -> dict:
    """Runs on a different model family from the one that wrote the reply."""
    body = "\n\n".join(f"[{i}] {p}" for i, p in enumerate(passages))
    rubric = "\n".join(f"{key}: {text}" for key, text in RUBRIC.items())
    out = call_judge_model(   # one call, returns parsed JSON
        system="For each criterion return JSON with verdict (pass, fail or "
               "unclear), score (0 to 1, how likely the criterion holds), why "
               "(one sentence) and, on fail, claim (the sentence at fault, from "
               "the reply or the customer's message). Key the object by "
               "criterion name.",
        user=f"PASSAGES\n{body}\n\nQUESTION\n{question}\n\n"
             f"REPLY\n{reply}\n\nCRITERIA\n{rubric}",
    )
    verdicts = {}
    for key in RUBRIC:
        v = out.get(key)
        if not isinstance(v, dict) or v.get("verdict") not in VERDICTS:
            v = {"verdict": "unclear", "score": None, "why": "no valid verdict"}
        verdicts[key] = v
    return verdicts

The score always points the same way as the criterion, so 0.9 means the judge rates the reply as likely to meet it, and that is the number you threshold and calibrate against the gold set in section 8. The loop at the end is there for the judge's own failures. A criterion that is missing from the output, or comes back with any word other than the three verdicts, is recorded as unclear, so a malformed judge response sends the answer to the next check and never ships it. Section 10 turns these four verdicts into one action.

Use a different model family than the generator

Section 4 made the case. GPT-4 recognised its own summaries 73.5% of the time, and self-preference rose in step with self-recognition. The flower company's agent answers on Luna and escalates to Sol, both GPT-5.6 models, so its judge runs on Claude Haiku 4.5, with Sonnet 5 for the checks that come back unclear.

Swap the order when you are comparing two things

Position bias is large and it is not a rounding error. Consistency, meaning the share of cases where a judge gives the same verdict after the two answers are swapped, came in at 65.0% for GPT-4, 46.2% for GPT-3.5 and 23.8% for Claude-v1 in the MT-Bench study. Running both orderings and taking the agreement costs a second call and removes the problem.

Verbosity bias belongs in the same paragraph. In the same paper, an attack that appended a repetitive list to an answer fooled Claude-v1 and GPT-3.5 91.3% of the time on 23 answers, and GPT-4 8.7% of the time. Judges reward length, and your generator will find that out before you do.

Give the judge an escape hatch

M13 built an abstain band and this is the same idea. A judge allowed to return "unclear" on the cases it cannot settle is more useful than one forced into a verdict, because the unclear pile is exactly the traffic that should go to a person. Jung and colleagues turned that into a formal result, escalating from a cheap judge to a strong one only when the cheap one was not confident enough, and reported that "on a subset of Chatbot Arena where GPT-4 almost never achieves 80% human agreement, our method, even while employing substantially cost-effective models such as Mistral-7B, guarantees over 80% human agreement with almost 80% test coverage" (Jung et al., 2024). That is M14's cascade, pointed at the judge.

Measuring it

What it costs to know whether the check works

A judge you have not measured is an opinion you are paying for. Measuring it means labelled answers, and the arithmetic of how many is the most useful thing in this module, because it is larger than people expect and it is the reason most teams quietly skip this step.

You want to know two numbers about your check. Recall is the share of genuinely bad answers it flags, and it determines how many bad answers reach customers. Precision is the share of its flags that were right, and it determines how much work you do for nothing. Both are rates, so measuring them to a stated tolerance is ordinary binomial arithmetic.

Labelled examples needed to pin a rate down, 95% confidence
To measure a rate toLabelled examples
± 10 points97
± 5 points385
± 2 points2,401
± 1 point9,604

Those are worst-case numbers, taken at a rate of 50% where the variance is highest, and they get smaller as the true rate moves toward either end. Read the middle row as the working default. Somewhere around 400 labels buys you a number you can make a decision on, and fewer than 100 buys you a number with ten points of slack in it, which is wider than most of the differences you will be trying to detect.

Precision is cheap to measure and recall is not

This asymmetry decides how you collect the labels, and it is not obvious until you try.

To measure precision you need labelled examples of answers the check flagged, and the check hands you those for free. It flagged them. Pull 400 of them out of the log, have somebody read them, count how many deserved to be flagged, and the whole thing is done in an afternoon.

To measure recall you need labelled examples of answers that were genuinely bad, including the ones your check sailed past, and nothing in your system knows where those are. You have to go and find them in traffic that mostly consists of answers that were fine. At the flower company's 14% failure rate on Luna, finding 385 bad answers means reading about 2,750. If your generator is better than that, the numbers get worse quickly.

Reading enough traffic to find 385 genuine failures
True failure rateAnswers you have to label
14%2,750
5%7,700
3%12,834
Two panels above a full-width strip. The left panel, measuring precision, says the check hands you the examples and shows 385 answers to read, sourced from a log query, summarised as one person and one afternoon. The right panel, in terracotta, measuring recall, says you have to go and find the failures it missed, and shows 2,750 answers to read at a 14% failure rate, 7,700 at 5% and 12,834 at 3%. The strip underneath, boxed in black, is headed measuring self-consistency and reads: run the check three times on the same inputs and count how often all three agree, no ground truth, no labellers, no gold set, with a large zero and the caption answers to label. A row at the bottom gives the gold-set sizes 97 for plus or minus 10 points, 385 for 5 points, 2,401 for 2 points and 9,604 for 1 point.
Two of these numbers are bought with somebody's reading time. The third one is three reruns.

Twelve thousand eight hundred reads to measure one number about one check is not a thing anybody is going to do, which is why the practical approach is stratified. Label every answer the check flagged, take a random sample of the ones it passed, and reweight the two groups by how much traffic each represents. That gets you an honest recall estimate out of a few hundred reads, and the cost of it is that your confidence interval on the passed group is set by the sample size you chose, so write that number down next to the estimate.

This is also where M1 comes back. The eval dataset built there, with its known-good answers, becomes a gold set for the judge the moment you point it that way. You already paid for those labels.

A confusion matrix for a check, two by two, with the check's verdict down the side and the truth across the top. Top left, the check passed a good answer, 81.7% of traffic, marked costs nothing. Top right, the check passed a bad answer, 2.1%, marked in terracotta as a bad answer reaches the customer and costs a handoff at 2.50 dollars. Bottom left, the check flagged a good answer, 4.3%, marked as a needless escalation costing 0.1443 dollars. Bottom right, the check flagged a bad answer, 11.9%, marked as working correctly. Underneath, two arrows show that the two error cells cost 17 times different amounts, which is why the threshold does not sit in the middle.
The two mistakes a check can make cost seventeen times different amounts at this company, so they get different thresholds.

The one number you can get without a single label

Precision and recall both measure the judge against the truth. Neither of them asks whether the judge agrees with itself, and that is a separate thing that can be wrong.

Run the judge twice on the same answer, with the same prompt and the same settings, and it can come back with two different verdicts. Naik and colleagues measure this and call it the self-consistency rate, defining it as querying the judge three times on each item with the same prompt and taking "the proportion of examples for which all three verdicts are the same" (Naik et al., COLM 2026). Haldar and Hockenmaier call the same idea self-reliability, and point out that reporting it is routine in essay grading, physical therapy and clinical diagnostics, and absent from almost every LLM evaluation (Haldar and Hockenmaier, 2025).

The reason to care about it before anything else in this section is the price.

What each number about your check costs to establish
The numberWhat it needsAt the flower company
Precision385 labelled examples of what the check flaggedan afternoon
Recall385 labelled genuine failures, found in traffic that is mostly fine2,750 reads at a 14% failure rate, 12,834 at 3%
Self-consistencythe check, run three times on the same inputsno labels at all

Self-consistency needs no ground truth, no labellers and no gold set. You take answers you already have, run the check on each of them three times, and count how often all three verdicts match. It is the cheapest honest number in this module and it is the one almost nobody reports.

What the measurements say

Naik and colleagues found that smaller open models sit only about 10% behind frontier models on balanced accuracy while being up to 25% less self-consistent, so a comparison on accuracy alone understates how far apart they are. Their twelve-prompt ensemble moved Qwen3.5-35B's self-consistency from 76.8% to 92.0% on ProofBench while its balanced accuracy went from 75.5% to 85.2%. A rate of 76.8% means that on roughly one item in four, a rerun would have decided differently.

Yagubyan ran the harder version of the experiment, repeating identical evaluations 50 times per question on 29 tasks across two OpenAI judges (Yagubyan, 2026). Pairwise preferences flipped on average 13.6% of the time, 28% of questions exceeded a 20% flip rate, and one question reached 56%. Agreement between the two judges came to 76%, which is κ = 0.51. Rewriting the prompt template into a semantically equivalent version changed the majority outcome in 25% of the cases tested.

The number from that paper worth writing down is the last one. Recovering the verdict that 50 trials would have given you, with 95% probability, took a majority vote over 11 repeated trials on average, and 15 on the high-variance questions. Eleven judge calls per answer is a different cost structure from the one section 5 priced, and it is what a stable verdict costs.

Temperature zero does not fix it

The standard response to all of this is to set temperature to 0 and assume the problem is gone. It reduces the problem and it does not remove it, for a reason that has nothing to do with sampling.

Tamba tested a real safety-evaluation codebase and found first that the harness called its grader without setting temperature at all, so the provider default of 1.0 applied and items near the decision boundary disagreed with themselves on up to about half of 20 runs (Tamba, 2026). Pinning temperature to 0 across 690 API calls, two providers, three model tiers and five sampling configurations left 1 or 2 of 7 borderline items still non-reproducible under forced greedy decoding. That is a small case study on seven items and it should be read as one, but the mechanism it points at is general, since the residual variation comes from "batch-size-dependent floating-point reductions, mixture-of-experts routing variability, and provider-side load balancing across nonidentical replicas". None of those live in your request. They live in the serving stack, and M11-1 already made the point that the serving stack is not yours.

Two things follow. Claude Opus 4.7 and 4.8 have dropped the temperature parameter altogether, so on those models the standard mitigation is not available to apply. And chasing determinism is the wrong goal anyway, because Haldar and Hockenmaier found that turning off sampling to force a stable rating made agreement with human judgment worse. A judge that always returns the same answer and returns the wrong one is not an improvement.

Their paper ends on a question worth keeping. What does an agreement of 0.8 with human labels mean, when the judge agrees with itself less often than that?

What to do with the number

Measure it on the gold set and on live traffic, three runs, and report it beside precision and recall, since it answers a question neither of those asks. A judge can be reproducible and biased, or unbiased and irreproducible, and the two numbers catch different failures.

Use the disagreements. An answer where three runs of the check split two to one is an answer the check has told you it cannot settle, and that is a better abstain signal than any threshold on the score, because it costs nothing extra to compute once you are already running three times.

And pick the fix that fits the budget. Yagubyan's eleven trials buys stability at eleven times the cost. Naik's prompt ensemble bought 15 points of self-consistency by running twelve different prompts and voting, which is the same trade with the calls spent on diversity. Section 7's four criteria all come out of the same call, so they share its instability, and what they add is a record of where it sits. When three runs split, the log shows which criterion flipped, and that is usually the criterion whose fail and unclear conditions need tighter wording.

The two errors

What a wrong verdict costs in each direction

M14 built a cascade, which answers everything on the cheap model, runs a check, and escalates whatever fails. It assumed the check was perfect, and once the check is allowed to be wrong the arithmetic moves in both directions at once.

A check can be wrong two ways and they are nothing alike. A false pass lets a bad answer through to the customer, and costs whatever M7's failure ladder says a bad answer costs, which for this company is a handoff to a person at $2.50 plus the goodwill nobody put a number on. A false fail escalates an answer that was already fine, which costs the difference between the two models, $0.1443 against $0.0080 a conversation.

Seventeen to one is the ratio that sets the threshold, and what it tells you to do is catch bad answers aggressively and accept that you will escalate a few good ones along the way.

M14's cascade once the check is allowed to be wrong
The checkEscalatedBad answers shippedEscalated for nothingPass rateA monthPer ticket
Perfect, M14's assumption14.0%0.0%0.0%99.6%$2,400$0.798
A good judge, recall 85 precision 9516.2%2.1%4.3%97.4%$2,590$0.839
A judge nobody measured, recall 5516.3%6.3%8.6%93.2%$2,599$0.913
A superficial check, recall 101.4%12.6%0.0%87.4%$1,309$0.993
A jumpy check, recall 95 precision 7039.1%0.7%25.8%98.1%$4,572$0.861
Five cascade configurations as rows, each with a stacked bar showing what share of all conversations the check gets wrong in each direction, bad answers shipped in terracotta and good answers escalated for nothing in taupe, alongside model spend a month and cost per resolved ticket. Perfect ships no bad answers, costs 2,400 dollars and 0.798 a ticket. A good judge at 85 recall and 95 precision ships 2.1% bad, wastes 4.3%, costs 2,590 dollars and 0.839. A judge nobody measured at 55 recall ships 6.3% bad, wastes 8.6%, costs 2,599 dollars and 0.913. A superficial check at 10 recall, highlighted in terracotta, ships 12.6% bad, wastes nothing, costs 1,309 dollars and 0.993. A jumpy check at 95 recall and 70 precision ships 0.7% bad, wastes 25.8%, costs 4,572 dollars and 0.861.
The row with the lowest model spend is the row with the highest cost per resolved ticket.

Every row in that table assumes a check that returns the same verdict every time it sees the same answer, which section 8 established is not a thing that exists. At a self-consistency rate of 77%, about a quarter of the escalations in the second row are a coin landing differently on a rerun, so part of what you are paying for is noise. The ranking between the rows survives that, since the gaps are much larger than the wobble, and the absolute numbers are softer than they look.

The superficial check is the cheapest arrangement on the page at $1,309 a month and the worst outcome at $0.993 a ticket. It almost never fires, so almost nothing escalates, so the model line looks wonderful. Meanwhile 12.6% of conversations ship a wrong answer. If the only dashboard anybody looks at is model spend, this configuration wins, and it is the one that quietly costs the most.

The jumpy check goes the other way. It catches nearly everything, at the price of escalating a quarter of the traffic for no reason, which nearly doubles the bill against the good judge and buys 0.7 points of pass rate. Whether that trade is worth taking depends on what a bad answer costs you, and for a support agent it is not.

Even a mediocre check beats no check

The ratio also cuts the other way, and it is the more encouraging half of this section.

Against answering everything on Luna and shipping it unchecked
Check recallCheck precisionPer ticketSaved
30%95%$0.9594.4%
50%95%$0.9158.8%
70%95%$0.87213.1%
85%95%$0.83916.3%
85%80%$0.86513.7%

A check that catches only three in ten bad answers still takes 4.4% off the cost of a resolved ticket, because escalating costs so much less than a person picking up the pieces. The argument for shipping a rough check today and improving it is strong, as long as you know it is rough, which brings this back to the previous section.

The check that never fires

There is one failure mode this table cannot show you, and it is the one to watch for.

A check with zero recall and perfect precision escalates 0% of traffic and costs $1,188 a month, judge included. A perfect check running over answers that happen to all be good escalates 0% of traffic and costs $1,188 a month. Those two are identical on every operational metric you have. The escalation rate is the same, the spend is the same, the latency is the same, and one of them is catching everything while the other is catching nothing.

Only labelled answers separate them, which is the entire argument for the previous section. MAST found this in the wild, reporting that "many existing verifiers perform only superficial checks, despite being prompted to perform thorough verification, such as checking if the code compiles or if there are leftover TODO comments", and giving as an example a generated chess program that "passes superficial checks (e.g., code compilation) but contains runtime bugs because it fails to validate against actual game rules, rendering the output unusable despite review phases."

Verdict to action

Turn every verdict into one action, and cap the repairs

Section 1 said a verdict ends in one of four actions, which were ship it, repair it, escalate it or give it to a person. The judge from section 7 returns pass, fail or unclear on each of its four success criteria in one call, the grounding check from section 4 returns a verdict per claim, and something in your code has to turn that pile of verdicts into exactly one of those four actions. If that mapping lives only in somebody's head, the same verdict gets handled differently on different days, and nobody can say afterwards why an answer shipped.

The first thing to settle is what counts as one failure. Because every criterion is written as a success condition, a fail always means the reply broke that condition, and the judge's claim field names what broke it. For grounded and policy that is a sentence of the reply, and for complete and on_topic it is the question or request in the customer's message that the reply missed. On the March answer, grounded fails and so does policy, but both point at the same claim, the 30-day damage window, so that is one failure with one piece of evidence behind it. Count failures by the claim or the question they point at, since two criteria tripping over one wrong sentence need one fix.

The harness reads the table below from the top and takes the first row that matches, so an unclear verdict or stale evidence is handled before anybody counts failures.

What each verdict sends the answer to, first match wins
What came backActionWhy
A failure whose only evidence is stale or conflictingGive it to a person, and flag the article for whoever owns the knowledge baseRegenerating reads the same passages and repeats the same answer
The repaired answer fails its re-checkGive it to a person, with both drafts and both verdicts attachedThe repair cap has been reached
Any criterion unclear, or three runs of the judge splitting two to oneRun the check again with Sonnet 5 as the judge, and give it to a person if Sonnet 5 also returns unclearA repair needs a named failure, and unclear names none
Two or more failuresEscalate, by regenerating on Sol from the original requestSeveral failures usually mean the draft went wrong from the start, and patching one per round costs a round each
One failure, with the passage sentence or the customer's question that shows itRepair it, onceThe failure is specific and the evidence to fix it is already in hand
Every criterion passes and every passage passed its currency checksShip itNothing is left for a check to say

The unclear row is the easiest one to get wrong, because the tempting move is to treat unclear as a soft fail and send it to repair. What happens then is that the generator receives a request to fix something with no failed claim and no evidence attached, so it rewrites the reply more or less at random, and the re-check runs on a different answer that is no better supported than the first one.

One repair, worked through on the March answer

A repair request carries the claim that failed and the exact sentence of evidence that contradicts it, both labelled with the criterion that caught them. The first generation had the whole article and nothing pointing at the sentence it got wrong. Without that pointer the repair is the intrinsic self-correction from section 4, where GPT-4 got worse with every extra round.

  1. 01The judge's policy criterion returns fail on the claim that damage can be reported up to 30 days after delivery. The entailment check agrees, scoring that claim against the article's damage sentence at 0.996 contradiction, the row from section 4's table.
  2. 02The harness builds a repair request holding the customer's message and this turn's passages, with the failure record from step 1 attached and an instruction to change only what the failure names.
  3. 03The generator writes a new reply. "Our policy asks for damage to be reported within 48 hours of delivery. Your bouquet arrived yesterday, so you're still inside that window if you send a photo of the arrangement today, and we'll get a replacement sent out."
  4. 04The whole check runs again from the top, meaning M9's rules, the claim split, the currency checks, entailment and the judge call with all four criteria. A repair can break something that passed the first time, like dropping the photo instruction, so re-checking only the failed criterion misses exactly that.
  5. 05The new reply passes and ships. Had it failed again, it would go to a person with both drafts and both verdicts attached, and nothing would loop.
repair_request.json, what the harness sends in step 2
{
  "mode": "repair",
  "customer_message": "My bouquet was delivered yesterday and half the roses...",
  "passages": [{"article": 305, "version": 3, "text": "..."}],
  "failed": [{
    "criterion": "policy",
    "claim": "Damage can be reported up to 30 days after delivery",
    "evidence": {"article": 305,
                 "sentence": "Damage must be reported within 48 hours of delivery."},
    "entailment": {"contradiction": 0.996}
  }],
  "instruction": "Fix only the failed claim and keep every other statement."
}

The repaired reply also tells the customer to act today, which the original had no way to say. Once the generator has the real 48-hour rule next to the delivery date, it can work out that the deadline is tomorrow, and that is the piece of information that decides whether the customer gets a replacement.

Why the cap is one repair

The repair works because the feedback comes from outside the generator, from the passage and from the entailment verdict, and that is the condition the research on self-correction keeps landing on. Kamoi and colleagues surveyed the self-correction literature and concluded that "self-correction works well in tasks where reliable external feedback is available", while finding no prior work showing it succeeding in general tasks with feedback from a prompted model alone (Kamoi et al., 2024). Huang's GSM8K result from section 4 is what happens past that line, with GPT-4 losing points as the rounds went up.

A second failure on the same claim tells you something specific. The generator had the contradicting sentence in front of it and still could not produce a reply that agrees with it, so a third call is paying for the same attempt again. Start the cap at one repair and raise it only if your labelled data shows second repairs succeeding often enough to pay for themselves.

Each round costs about $0.004, which is one Luna generation at roughly $0.0017 plus one Haiku 4.5 judge call at $0.0025, going by section 5's per-answer figures and before the claim splitter's own call. That is tiny next to the $2.50 of a handoff, so what limits repairs in practice is time, because every round adds a generation and a judge call to a turn that already has a deadline from M11-1, so the harness checks the time left before it starts a repair, and when a full round will not fit it escalates straight away.

Log every repair with the failed criterion, the claim, the evidence and whether the re-check passed. The share of repairs that pass on the first try belongs next to the flag rate, because a falling repair pass rate means the failures have moved somewhere the evidence you retrieve does not cover.

In production

Checks go stale, and generators learn to pass them

A verifier is a component in production, so it drifts the same way M13's classifier and M14's router drift, and it has one failure mode of its own.

The rubric ages first. Article 305 gets rewritten, the replacement window moves from 48 hours to 72, and a judge holding the old criterion starts failing answers that are now correct. Every rubric criterion should name the source it came from, and changing that source should raise a flag on the criterion. The passages go out of date on the same day, which is what the currency checks in section 4 catch, so one change to article 305 needs to raise both flags.

The judge model also changes underneath you, since a provider can update the model behind the name and leave your threshold sitting somewhere it no longer belongs. M11-1 covered the general shape of that problem, and what is specific to a judge is that it is the last component you notice moving, because its output is a verdict about somebody else's work. A drifting judge and a drifting generator look identical from outside until you go and check.

The gold set ages out too, because labels collected in February describe February's traffic and nothing else. Refresh a share of it on a schedule and keep the old entries alongside the new ones, so that running both lets you tell a change in the judge apart from a change in the traffic.

The failure mode with no equivalent in M13 or M14 is that the generator learns to pass the check. Any check you optimise against stops measuring what it measured, because the optimisation pressure finds the check's blind spots first. It shows up when you tune prompts against a judge's score, and much more sharply when you train against a verifier, which is why Lightman and colleagues went to the expense of process supervision, labelling each reasoning step in turn where outcome supervision labels only the final answer, and released PRM800K with 800,000 step-level human labels to do it (Lightman et al., 2023). Their process-supervised model solved 78.2% of a representative subset of the MATH test set, and active learning gave a 2.6× improvement in data efficiency on top.

The operational version for a team not training anything is simpler. Keep a set of labelled examples that nothing is ever tuned against, look at it monthly, and treat a gap opening between the tuned score and the held-out score as the signal that the check has been worn smooth.

What to put on the dashboard

Model spend on its own is the metric that makes the superficial check look best, so it needs company.

  • Flag rate, the share of answers the check fails, which is the first thing to move when anything changes.
  • Agreement with the gold set, re-measured on a schedule, reported with its confidence interval.
  • Self-consistency, three runs over the same sample, which needs no labels and catches a drifting serving stack before anything else does.
  • The split between the two errors, since a flag rate that holds steady while precision falls is a different problem from one where recall falls.
  • Judge latency at p95, against whatever turn deadline M11-1 set, because a judge sits in the request path.
  • The unclear rate, which is the check telling you your traffic has moved somewhere its criteria do not cover.

Putting it together

Putting it together

The flower company's agent now has three layers and each one has a job the layer above cannot do.

M9's rules run first on every answer, free and certain, and they catch the malformed replies, the invented article ids and the numbers that appear in no source. They run in microseconds and they never have to be right about anything, only consistent.

A judge runs second, on Claude Haiku 4.5, a different model family from the GPT-5.6 models that write the reply. It reads the reply next to the passages search returned and returns pass, fail or unclear on four success criteria in one call. It costs $708 a month across 60,000 conversations at 4.8 answers each, and it catches the March answer, because the policy criterion passes only when every deadline matches the passage for the same kind of request, and it does not care that the number was lifted from a real article.

What the judge returns decides one action. Before any claim is scored, the passages behind it are checked for being current and allowed to state policy. One failing claim gets one repair that carries the claim and the sentence contradicting it, and then the whole check runs again. Anything unclear goes up to a stronger judge or to a person, a person also gets any answer that fails its repair or rests on stale evidence, and every repair is logged with whether it passed.

A person reads 2% of conversations, which is $3,000 a month and the largest line in this system by some distance. That sample produces the labels, and the labels are what let anybody say what the judge's recall is, so without them nobody can tell whether the judge is catching anything.

What holds the three together is a handful of numbers that somebody re-measures on a schedule. The judge's agreement with the gold set, its precision on the answers it flagged, its recall estimated from a stratified sample, and the flag rate on live traffic. When one of those moves, the thing to work out first is which side moved, since a judge and a generator drifting look identical from outside.

M16 takes the next capability, which is memory, and it changes what verification has to do, because an agent carrying state across turns can be wrong about something it decided twenty minutes ago and there is no passage to check that against. M21 covers training a verifier of your own, which is where Cobbe's 30× and Lightman's 800,000 labels stop being background reading and become a budget.

Checkpoint · recall · 9 questions

What the module said

  1. 01

    What does a verifier read?

  2. 02

    Where does M9's validation ladder stop and this module start?

  3. 03

    In the 2023 Huang study, what did GPT-4 score on GSM8K after two rounds of correcting itself with no oracle?

  4. 04

    What share of failures did MAST attribute to task verification?

  5. 05

    Which rung of the ladder knows the most for the least money, when it is available at all?

  6. 06

    A 2026 replication asked current models to flag a known error sitting in their own reasoning. Roughly how often did a frontier model catch it?

  7. 07

    Which number about your check can you establish without labelling anything?

  8. 08

    Roughly how many labelled examples pin a rate down to within 5 points at 95% confidence?

  9. 09

    Which of these is one atomic claim, by this module's definition?

0 / 9 answered

Checkpoint · understanding · 10 questions

Reason it through

  1. 01

    Your judge scores answers 1 to 10 and you pass anything above 7. A colleague suggests four pass-or-fail criteria returned by the same call. What do you gain?

  2. 02

    Your generator runs on Luna and you are choosing a judge. Why is Luna the wrong answer?

  3. 03

    Three sampled answers agree with each other. What have you learned?

  4. 04

    Why does a judge that costs $0.00246 a call come to $708 a month on 60,000 conversations?

  5. 05

    Your check has recall of 30%. Is it worth shipping?

  6. 06

    Your generator gets a lot better next quarter. What happens to the job of verifying it?

  7. 07

    Your agent clicks a toggle and you compare the before and after screenshots. Pixel diff says 0.57% changed, CLIP says 0.980 similar. What do those numbers tell you?

  8. 08

    Why is precision so much cheaper to measure than recall?

  9. 09

    An assistant writes SQL for "refunds in August 2026". The query runs without an error and returns 1,169. What does rung 1 still need before you trust that number?

  10. 10

    The judge fails two different claims in one reply, one about the deadline and one about the refund amount, each with a contradicting passage. What does the harness do?

0 / 10 answered

Checkpoint · debugging · 11 questions

Debug it

  1. 01

    Your cascade's escalation rate is 1%, model spend is the lowest it has ever been, and complaints are up. What do you check first?

  2. 02

    The judge's flag rate held steady at 16% for months and cost per ticket has been climbing. What moved?

  3. 03

    You have to have a model check something it just produced. What is the cheapest change that helps?

  4. 04

    A teammate adds a judge criterion asking the model whether its own previous answer was correct, in the same call that produced it. What will it do?

  5. 05

    Your judge gives different verdicts on reruns. Somebody sets temperature to 0 and calls it fixed. What is still wrong?

  6. 06

    Your judge's agreement with the gold set dropped from 0.84 to 0.71 and nothing was deployed on your side. Where do you look?

  7. 07

    A VLM check on your agent's screenshots reports 50% accuracy on a balanced set and the escalation rate never moves. What do you suspect first?

  8. 08

    Somebody proposes tuning the answer prompt against the judge's score until the score stops improving. What goes wrong?

  9. 09

    A customer asks about a damaged bouquet. The reply says "You have 30 days to report damage, so reply with a photo whenever you're ready." It cites article 305, which search returned, and 30 appears in article 305. Both the citation rule and the numbers rule pass, and a grounding check that scores the whole reply against the article returns 0.83 supported. What went wrong, and what should the harness do?

  10. 10

    After the damage window moves from 48 hours to 72, the agent tells a customer they have 48 hours. The atomic claim scores 0.99 entailed against article 305 and the judge passes it. What check was missing?

  11. 11

    A teammate adds a fifth judge criterion, "Does the reply promise a refund the policy does not allow?", and the harness ships any answer where every criterion comes back pass. Next week an answer promising a refund on a bouquet delivered three months ago ships. What went wrong?

0 / 11 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.