MLGuerrillaStart with M1 →
Free · in beta·intermediate·M12·42 min read·Prereq: Production, Deployment & Scale (M11-1 and M11-2). The support agent from M6 to M11 is the running example, and its numbers come from the same workload sheet.

LLMOps & Cost

The capability

What LLMOps and cost engineering for an LLM system are

LLMOps for a model-backed feature is the loop that keeps it working after launch, where every change goes from a signal through an experiment and a release gate into production, and this module follows that loop through its cost. Cost engineering is the half that knows what one useful outcome costs and which levers change that number, including the ones that also move quality.

The first part of that is picking a unit somebody outside the team would recognise. Tokens are what the provider bills you in, but cost per resolved ticket is what anyone else can act on, and the gap between those two numbers is where most cost surprises turn out to be hiding.

The second is that every call bills on more than one meter. Input tokens, tokens read from the cache, tokens written to the cache and output tokens each carry a different price, and output is usually the most expensive of the four by a wide margin.

The third is that cost per conversation is a distribution with a long tail. The average conversation and the conversation that sets your bill are rarely the same conversation, which means a forecast built on the mean is a forecast built on the wrong conversation.

The fourth is that a cost belongs to a version. Every experiment and every week of production traffic ran on one specific version of the prompt, the model, the tools and the dataset, so a cost number that doesn't say which can't be compared with anything. The last two sections join the pieces M1, M2, M8 and M11 built into one loop. They run a customer complaint through it and attach a cost to each step.

A ladder of five units stacked from cheapest and least meaningful at the bottom to most meaningful at the top. One token, then one model call, then one turn, then one conversation, then one resolved ticket. Beside each rung sits who cares about it: the provider bills tokens, an engineer debugs a call, a product owner thinks in turns, a support lead thinks in conversations of about 4.8 turns each, and finance thinks in resolved tickets. A marker sits on the resolved-ticket rung, labelled the unit to manage. A note underneath says the multiplier between the bottom rung and the top is the thing nobody estimates.
Every rung is a real number. Only the top one answers the question finance is asking.

You can't forecast this from a price list alone. What sets the bill is how many calls one outcome takes, how much history gets re-sent on each of them, how much of that history was cached, and how long the answers are. This module works out each of those on a real workload and then goes after them in order.

Where it shows up

Every model-backed feature has this problem the moment it stops being a prototype.

  • A support agent whose cost per conversation is 5 to 10 times its cost per model call.
  • A coding assistant where a long session re-sends the whole file tree on every turn.
  • A document pipeline where the per-page price looks fine until somebody uploads a 400-page PDF.
  • An internal chatbot that costs nothing at 50 users and becomes a line item at 5,000.

The arithmetic is the same in all four, and it starts with one number being wrong.

Where this starts

The spreadsheet said $588

It's January of the flower company's second year, and February is the question. Last year the agent ran on half the traffic as a test. This year every ticket goes to it first, which is about 60,000 conversations in the month. Finance wants a number. The company is made up, and so are its numbers, and the prices are real ones read in September 2026.

M7 priced one answer at about a cent. A warm answer on GPT-5.6 Sol, with the system prompt read from the cache and a fresh policy passage written into it, comes to $0.0098. Multiply:

“60,000 conversations × $0.0098 per answer = $588 for February.”

The month's model spend, worked out from what the traces say a conversation does, is about $8,923. That's 15 times the estimate, and none of the gap is a mistake in the per-answer arithmetic. The per-answer number prices one call. February bills conversations.

Six things sit between the two numbers, and each one gets its own section:

A conversation is more than one answer, for a start. The agent averages 4.8 turns, and a turn that calls a tool makes two model calls, so one conversation comes to about 7 calls before anything has gone wrong.

Each of those calls re-sends the whole conversation. M3 established that the model keeps nothing between calls, so every turn carries everything that came before it, which means input grows with the square of the turn count.

Reasoning tokens are billed as output, and there are more of them than of the answer. The visible reply is 200 tokens and the reasoning behind it is another 400.

Cache writes stopped being free. On GPT-5.6 and later, writing a prefix costs 1.25× the uncached rate, and reads cost 0.1× (OpenAI, prompt caching).

Some calls get retried, which bills the input a second time. M9's first-try success rate of 97% works out to about 1.03 expected attempts per call.

And the longest conversations dominate everything. The 7% that run 20 turns or more carry 46% of the spend.

This module builds the model that produces the $8,923, then works through what to do about it. Every number below comes from the same workload sheet the capacity numbers in M11 came from, and every price carries the date it was read.

Unit cost

A conversation costs more than its answer

The first fix is picking the right unit. Tokens are what the provider bills, and they're the wrong thing to manage, because nobody asks for tokens. The ladder runs tokens → one model call → one turn → one conversation → one resolved ticket, and the useful unit is near the bottom. The FinOps Foundation, which writes the vocabulary most finance teams use for cloud spend, describes the same move as going from cost per token to "outcome oriented measures, for example cost per assist, cost per agent action, or cost per case deflected" (FinOps Foundation).

Every call bills four meters

Every call bills four separate things, and a cost model that tracks one number per call will be wrong in both directions:

What one call to GPT-5.6 Sol bills, as of September 2026
MeterWhat it isPrice per million tokens
Uncached inputTokens the model reads that weren't in the cache$4.00
Cache writeTokens stored for reuse, at 1.25× the input rate$5.00
Cache readTokens served from the cache, at 0.1×$0.40
OutputEverything generated, including reasoning tokens$20.00

The providers report these separately, and the field names differ. Anthropic returns cache_creation_input_tokens and cache_read_input_tokens in the response's usage block, beside its own input count (Anthropic, prompt caching). Check which total includes cached tokens before doing arithmetic on it, since a total that already contains them will double-count. Either way the trace from M2 carries the fields, so cost per call is arithmetic on data you already store.

History is re-sent, so input grows with the square of the turns

M3 established that the model holds nothing between calls. Turn 5 re-sends turns 1 through 4, so the input of a conversation is the sum of a growing series, and its total input grows roughly with the square of its turn count.

One conversation on Sol, with caching on
TurnsInput tokensOutput tokensCost
12,960600$0.0176
211,7601,660$0.0502
432,1603,320$0.1048
10145,2008,300$0.2894
20506,40016,600$0.6661
401,876,80033,200$1.6788

Forty turns is 1.9 million input tokens for one customer conversation. That's the whole reason a per-answer estimate can't forecast a month.

Caching is what keeps that in check. The same 40-turn conversation costs $8.17 with no caching at all, against $1.68 with it, because every re-sent token is billed at the full input rate each time.

The multipliers outside the tokens

Three more terms sit between a call's tokens and a ticket's cost:

Attempts are the first. M9 measured a 97% first-try success rate, and with up to two retries the expected attempts per call comes to 1.031, with every retry billing the input again.

Handoffs are the second, and they dwarf the model. M7's design sends about 30% of conversations to a person, and a person costs roughly $2.50 of handling time.

Server-side tools are the third. Tools the provider runs for you, like a hosted web search, get billed per call on top of the tokens, so an agent that searches on most turns pays that fee on most turns.

Putting them together for the flower agent:

“cost per resolved ticket = $0.1443 × 1.031 + 0.30 × $2.50 = **$0.899**”

The model is 17% of that. Uber's engineering team describes the same decomposition for its own agents, as "six terms that multiply," naming price per token, tokens per request and requests per turn among them, and notes that "Every turn re-sends the full conversation history, project context, and tool results" (Uber, August 2026). Anthropic's cost guide gives the same instruction in one line, to "Compare on cost per completed task, not per token" (Anthropic).

Forecasting

The bill is a distribution

The $0.1443 above is a mean, and means are dangerous here. Cost per conversation on this workload looks like this:

Cost per conversation on Sol
StatisticValueWhat it's for
Mean$0.1443The budget, once multiplied by volume
Median$0.0502What a typical customer costs
P95$0.6661The per-conversation guard
Share of spend from the longest 7%46%Where the money goes
A bar chart of conversation length against share of conversations, with each bar's share of total model spend written above it. One turn is 25% of conversations and 3% of spend at $0.02 each. Two turns is 25% and 9% at $0.05. Three turns is 15% and 7% at $0.07. Four turns is 10% and 7% at $0.10. Six turns is 10% and 11% at $0.16. Ten turns is 8% and 16% at $0.29. The last two bars, in terracotta, are 20 turns at 5% of conversations and 23% of spend for $0.67, and 40 turns at 2% of conversations and 23% of spend for $1.68, which together are 7% of conversations and 46% of spend. Three statistics sit underneath: a median of $0.0502 for the typical conversation, a mean of $0.1443 in terracotta for what the month is billed on, and a P95 of $0.6661 for what the guard is set from.
Every bar is the same agent answering the same way. The only difference is how long the customer stayed.

The mean is nearly three times the median, because a handful of long conversations carry the month. Anthropic saw the same shape on a 20-problem benchmark run, where "two problems carried 43% of the spend."

That difference decides M7's ship bar, which says "5 cents or less per conversation" without saying which statistic. The median is 5.0 cents, so the typical conversation passes by a hair. The mean is 14.4 cents, so the budget fails by a factor of three. A ship bar has to name the statistic, or two people can read the same dashboard and disagree about whether the feature shipped.

The forecast itself is three numbers multiplied, and the last one is the one people forget:

“February = 60,000 conversations × $0.1443 mean × 1.031 attempts ≈ $8,923”

Prices belong in the forecast too, with their dates. GPT-5.6 Sol's $4 and $20 are promotional, "available at least through November 21, 2026" (OpenAI pricing, read September 2026). A February forecast written in November has to say which price it assumed, because the number changes under it without anyone deploying anything.

Free levers

Cheaper without touching quality

Cost levers come in two groups, and the order is what keeps a team out of trouble. The first group changes how much work the model does without changing what it can do, so nothing has to be re-evaluated. Anthropic's cost guide opens the same way, telling any workload on any model to "Turn on prompt caching and trim unneeded tokens; both are free" (Anthropic).

The biggest lever is still sending less, and M7 found most of it a year earlier by establishing that 60% of tickets needed no model at all. After that comes everything riding along on each call, which is the tool definitions, the retrieved passages and the history. Dropping retrieved passages out of the history once a turn is finished, on the grounds that the answer already used them, takes the mean conversation on the flower agent from $0.1443 to $0.1112. That is a 23% cut with no change at all in what the model can answer.

Next is the cache hit rate. M5 covered how caching works; the number to watch here is cache reads divided by all input tokens, per route, per day. Both providers now charge a premium to write a prefix, so caching only starts paying once that prefix gets read back. OpenAI's own arithmetic, as of September 2026, is that "Writing a prefix once and fully reusing it once costs 1.35× its ordinary input cost, compared with 2× for processing it twice without caching."

Then there is the shape of the output, which is worth attention because output is the expensive meter and two thirds of the flower agent's output tokens are reasoning nobody reads. Asking for shorter answers cuts it, and so does running at a lower reasoning effort, though that second one belongs in the next section because it changes what the model is able to do.

Treat max_tokens as a backstop and nothing more. It does not lower the bill, since you are billed for what gets generated, and setting it too low ends an answer mid-sentence, which M4 covered. The truncated attempt still gets billed.

And batch anything nobody is waiting for. The nightly summaries from M11 run at half price on the provider's batch API, documented as a "50% cost discount compared to synchronous APIs" in exchange for a 24-hour turnaround.

Cache TTL is an off-peak decision

A cached prefix expires. Whether the next conversation finds it warm depends on how often conversations arrive, and traffic at 3 a.m. is a different problem from traffic at noon. With arrivals spread randomly, the chance the cache is still warm is 1 − e^(−λT) for arrival rate λ and time to live T:

Chance the shared prefix is still cached
Traffic5-minute TTL30-minute TTL60-minute TTL
3 conversations an hour, overnight22%78%95%
60 an hour, midday99%100%100%
1,152 an hour, Valentine's peak100%100%100%

At peak the TTL is irrelevant, and overnight it decides whether every conversation pays the write premium again. Longer time to live costs more to store on providers that price it, so the decision belongs to the quiet hours.

One line at the top of the prompt can undo all of it

The cache matches on an exact prefix. Anything that changes at the top of the prompt invalidates everything below it, so a helpful-looking addition like Current time: {now} turns every call into a full write:

What a timestamp at the top of the system prompt costs
With a stable prefixWith the timestamp
One warm call$0.0098$0.0190
Mean conversation$0.1443$0.5285
February$8,923$32,689

No answer changes, no test fails, and the bill nearly quadruples. Section 11 puts a check for exactly this in the release gate.

Trade-off levers

Cheaper by trading quality, measured

The second group of levers changes what the model can do. Each one runs through the M1 dataset before it ships and the M8 canary after, and each one is judged on cost per resolved ticket, since a cheaper model that fails more sends more conversations to a person at $2.50 each.

Cost per resolved ticket, with failures priced in
ConfigurationCost per conversationPass rateCost per ticket
Sol, medium effort$0.200397%$1.009
Sol, low effort$0.144396%$0.969
Terra, low effort$0.079994%$0.937
Sonnet 5, low effort$0.072196%$0.894
Haiku 4.5$0.022191%$0.930
Luna, low effort$0.008086%$1.003
A scatter chart of six configurations, with cost per resolved ticket across the bottom and pass rate on the 200-question dataset up the side. Sonnet 5 at low effort sits at $0.894 and 96%, and Sol at medium effort at $1.009 and 97%; both are terracotta and joined by a dashed frontier line. Sol at low effort is $0.969 at 96%, Terra at low effort $0.937 at 94%, Haiku 4.5 $0.930 at 91%, and Luna at low effort $1.003 at 86%, all in black below the line. Three cards underneath say that Sonnet 5 low is on the frontier at $0.894 a ticket and Sol medium buys one more point for 12 cents, that Sol low, Terra low and Haiku 4.5 are beaten on both axes, and that Luna costs 18 times less per token than Sol and the most per ticket, because 14% of its conversations reach a person.
Sol at medium effort is on the frontier too. Whether one point of pass rate is worth 12 cents a ticket is a product question, and this chart only prices it.

The pass rates are invented, and the shape they produce is the point. Luna's tokens cost 18 times less than Sol's and its tickets cost more than Sol's, because ten points of pass rate is worth more than the entire model bill. Anthropic measured the same inversion on its own benchmark, where a model with a per-token price five times higher "solved 88.6% of tasks for $0.54 per solved task, against 77.4% for $0.84" from the cheaper one.

Build this table from the M1 dataset, one row per configuration, and read the frontier: the configurations nothing else beats on both cost and quality. Then pick a point on it with a quality floor, and re-run the table after every model migration.

Routing and cascades

Two ways to use more than one model:

  • Route before generating. A classifier picks the model per request. RouteLLM's routers cut cost "by up to 75% as compared to the random router" at the same MT Bench quality. The same paper is honest about where it stops working, since on MMLU "all routers perform poorly at the level of the random router when trained only on Arena dataset," because those questions were outside what the router had seen (RouteLLM). A router is a model trained on a distribution, and your traffic is a different one.
  • Cascade after generating. The cheap model answers, a check decides whether to escalate. On the flower agent's workload, sending everything to Luna first and escalating a tenth of conversations to Sol costs $0.0224 a conversation, an 84% saving, and it only stops saving at 94% escalation. The catch is the cache. Switching models mid-conversation means the second model has never seen the history, so at turn 10 it writes about 14,700 tokens at $0.0734 where the first model would have read them for $0.0059.

Reusing answers is a different thing from caching them

Prompt caching reuses computation on an identical prefix and never changes an answer. A semantic cache reuses a previous answer when a new question looks similar, and similar is a threshold someone picks. The vCache paper measured what those thresholds do, and the trade is brutal:

Semantic cache thresholds, from the vCache paper
Similarity thresholdWrong answersCache hit rate
0.992.5%37%
0.984.1%53%
0.975.2%67%

On the flower agent, a 67% hit rate saves about $6.57 per thousand questions and serves about 52 wrong answers in the same thousand. It breaks even only if cleaning up a wrong answer costs less than $0.126. M7 priced a wrong policy claim well above that, so the answer here is no. Where it can work is non-personalized, repeated questions, scoped by tenant and model version, and never for "where is my order."

Limits and throughput

Capacity is part of the price

M11 sized the peak: about 134 model calls a minute in the busiest hour, roughly 1.7M input tokens and 75k output tokens a minute, with bursts around three times the hourly average. Those numbers decide which tier the company has to be on before February, and buying that tier is this module's problem.

Three things make the decision harder than reading a price:

What the limit counts differs by provider, and M11 covered those mechanics. The money consequence is that the same workload can sit right up against one provider's limit and nowhere near another's, and that max_tokens reserves capacity you may never use.

Ramps carry their own limit on top of that. Traffic that triples over a morning gets throttled even while sitting under the ceiling, which is why the tier has to be arranged in January, the month this lesson is set in.

And reserved capacity is priced by what you hold, with no discount for leaving it idle. The flower company's peak hour runs about 35 times an average hour in a normal week, so capacity sized for that peak sits idle almost all year, which is why the answer for a seasonal business like this one is usually pay-as-you-go, with the tier arranged before the peak.

Reserving also promises less than it sounds. Azure's own documentation on provisioned throughput warns that "Unused quota doesn't guarantee that capacity is available when you want to scale back up your PTU deployment. Provisioned capacity is a finite, dynamically changing resource" (Microsoft). Committed spend buys priority, and not a reservation of physical hardware.

Buy vs run

Rent tokens or rent GPUs

Hosting an open-weight model swaps a per-token price for a per-hour one, and per-hour only wins when the hardware is busy. The break-even is one line:

“break-even utilization = (GPU cost per hour) ÷ (tokens per hour at full load × API price per token)”

Put the flower company's numbers in it. February's model spend is about $8,923 across the whole month, and a pair of high-end GPUs rented by the hour runs into the thousands of dollars a month before anyone serves a request. With a peak hour 35 times the average, hardware sized for February 13th idles through the rest of the year, and hardware sized for the average can't serve the peak.

Self-hosting starts to pay when traffic is steady and high, when the task is narrow enough for a small tuned model, or when data rules make the API impossible. Three costs hide behind the hourly rate: cold starts measured in minutes while weights load, the replicas you keep for availability, and the engineering time M11 described as everything you're now on call for. M26 covers the serving internals.

Attribution

Every dollar has an owner

When the bill jumps 20%, the useful question is which feature, which customer or which change did it, and no provider dashboard can answer that. Cost per request exists only in your own logs.

The fix is tagging at the seam. Every call through complete() carries the route or feature, the tenant, a hashed user id, the model and snapshot, the prompt version, the environment and the trace id. M2's traces already hold the token counts, so cost per call is those counts times a dated price table, and every other view is a group-by:

  • cost per resolved ticket, by route
  • cache hit rate, by route and by day
  • P95 cost per conversation, which is what the guard in the next section is set from
  • spend by prompt version, which turns "the bill moved on Tuesday" into "v14 costs 31% more than v13"

Reconcile your computed total against the provider's invoice daily, because the gap is informative. It comes from untagged traffic, from contracted prices your table doesn't know about, and from retries you didn't count.

Uber's team published what this discipline buys. Holding the model fixed so the numbers reflect their own work, "cost per 1,000 model requests is down almost 34% from its peak, and cost per session is down 52% from its June peak" (Uber, August 2026).

Spend guards

Caps that bound the damage

M9 stopped a single run from looping forever. A spend guard is the same idea at the level of the account, and the arithmetic makes the case for one. A 200-turn conversation on Sol costs about $22 on its own, and ten thousand scripted 40-turn conversations cost about $16,800 in an afternoon.

The first thing to know is that most budget features are alerts. OpenAI's spend controls, as of September 2026, separate the two plainly, where a spend alert "Sends a notification; API traffic continues," and only a hard spend limit stops traffic, with the warning that "Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount" (OpenAI, spend limits).

So guards go in rings, from the call outward, and each one has a different owner and a different speed:

Spend guards, innermost first
GuardSet fromWho enforces it
max_tokens per callThe longest real answerYour code
Input size check before sendingThe window budget from M5Your code
Per-conversation token or dollar budgetP95 cost per conversationYour code
Step and tool-call capsM9's stop conditionsYour harness
Per-user and per-tenant daily quotasNormal usage, with headroomYour API
Project rate limitsThe capacity planThe provider
Hard spend limit on the projectThe month's budgetThe provider, with a lag
Kill switch to a cheaper pathA plain-language degraded replyAn operator, via M11's flag

The inner rings act in milliseconds and the outer ones in minutes to hours, which is the argument for having both. Anomaly detection sits alongside them, watching tokens per conversation and spend per tenant against what the forecast expects. M10 covers the attack side, where the bill is the target.

The lifecycle

One complaint, from the support queue to a rollout

Everything so far priced one configuration of the agent. A running feature changes every few weeks, though. Each change starts from some signal and goes through an experiment and a release, and the production traffic after it produces the next signal. M1, M2, M8, M9 and M11 each built one piece of that loop. This section runs one change through all of it in order and puts a cost on each step.

On March 3rd a customer writes in about a bouquet that arrived a day late. The customer had asked the agent whether same-day delivery reached their village, and the agent said yes. The village is outside the same-day area. The complaint and everything that follows are made up, like the company's other numbers, and the prices are the dated ones from section 3.

Six numbered cards in two rows, with a dashed bar under them that loops back to the start. Card one is the complaint, a same-day promise to a village outside the area, where the trace names release 2026-02-24.1 and retrieval returned 4 passages with the delivery-areas article ranked 5th. Card two is the reviewed test case, where the support lead writes the expected behavior and four more conversations from 30 days of traces join it, taking the dataset from v8 to v9 with 205 test cases. Card three re-runs production on v9, where it passes 192 of 205, and notes that its 96% on v8 doesn't compare. Card four is two candidates, A retrieving 8 passages a turn and B adding check_delivery_area, both with a delivery_promise field, and about $55 of eval spend for five runs each. Card five, in terracotta, is the gate, where both pass 197 of 205 and anything above +5% cost needs a named approver, so A at +27.3% and $2,400 a month needs approval and B at +3.8% and $340 a month passes. Card six is the canary and then everyone, where the field check finds no out-of-area promises against 41 from the text search before, delivery questions were 26% of traffic where the sheet assumed 20%, and observed cost is +4.7% or $420 a month. The dashed bar says the workload sheet gets delivery-question share as a measured input and the next complaint starts from this release record.
Every box writes something into the release record, and the record is what the next change starts from.

The complaint becomes a reviewed test case

The support lead finds the conversation from the id on the ticket, and the trace from M2 shows what happened. The agent searched the help center and got back the general same-day policy, and it never saw the article that lists the areas, because that article ranked fifth and retrieval returns four passages. The trace also carries the release id, 2026-02-24.1, which says exactly which bundle answered.

A complaint isn't a test case yet, because complaints can be wrong. Of the 23 conversations support tagged as a wrong answer in February, review found 9 where the agent had been right and the customer had misread the policy. So a person who knows the policy writes down the expected behavior before anything goes into the dataset. Here that's the support lead, and the expected behavior is that the agent checks the postcode before it promises same-day delivery, and offers next-day when the postcode is outside the area. Grading that behavior takes some care, because the reply is free text. A same-day promise can be worded as "yes, we can do today" or "it'll be with them this afternoon", and no string match or regular expression can find every version of it. So the fix gives the promise a structured place to live. Prompt v16 has every reply come back as structured output, the M4 technique where the model fills in a JSON schema, with the text in one field and a delivery_promise field beside it that holds the service promised and the postcode it was promised for, or null when the reply promises nothing. Code compares that field with the area table, which is exact and costs nothing to run.

The field can't catch a reply whose text promises one thing while the field says another, like text that says "this afternoon" next to a null field. So the eval also has a judge read the text for any delivery promise, and a test case passes only when the field is right and the judge finds no same-day promise in the text for an out-of-area postcode. The judge can miss a promise worded in a way it hasn't seen, which is why M1's rule applies and a person checks it, and here that's cheap, since the support lead can read the replies to the delivery test cases by hand.

One complaint is one input, and the failure behind it is usually wider. Production doesn't have the field yet, so finding the same failure in the last 30 days means reading text. A small model reads the replies in every conversation that mentions delivery, about 12,000 of them, and pulls out any delivery promise with its postcode, for about $50. Code checks those postcodes against the area table, and a person reads the 58 out-of-area promises it flags and confirms 41. The model can miss oddly worded promises, so 41 is a lower bound, and none of those customers complained. The lead picks four of them that differ in how the question was asked, like a town name with no postcode, and they join as test cases too. They're picked by how the question was asked, because M1's rule is that test cases chosen by which version gets them right tilt the dataset toward that version. Dataset v8, with 200 test cases, becomes v9 with 205.

Then the version running now is re-run on v9, because M1's other rule is that a score only compares with a score from the same dataset version. It passes 192 of 205, which is 93.7%. Its 96% on v8 is still true, and the two numbers don't compare.

Every score names the versions that produced it

M11 listed what runs in a release (the image, the prompts, the model id, the tool schemas, the index and the flags) and wrote them into a release manifest. A score needs two more entries, because 96% means nothing without the dataset it was measured on and the judge that graded it. Put together, that's the version tuple, and every eval run and every production trace carries all of it:

The version tuple, before and after the fix
PartRunning now, 2026-02-24.1Candidate, 2026-03-09.1
Codeimage sha256:4c7e…image sha256:b812…, adds check_delivery_area
Promptpolicy-answer v15policy-answer v16, adds the delivery_promise field and a line on when to call the tool
Model id and effortgpt-5.6-sol-2026-08, lowunchanged
Tool schemasv7v8
Retrieval settingskb-live → build 43, 4 passagesunchanged
Eval datasetv8, 200 test casesv9, 205 test cases
Online judgev5v6, adds a check on delivery-area promises
Price tableread 2026-02-02read 2026-03-02

The judge line is the one teams forget. M9's online judge grades a 5% sample of live conversations, and v6 adds a criterion, so last week's score was graded by different instructions. Re-grade last week's sample with v6 before the release goes out, and use that as the baseline the canary compares against. M9's weekly labels from a person are how you'd notice v6 grading the old criteria differently from v5.

The price table is in the tuple because every cost number in this module is tokens times a dated price. When Sol's promotional price ends, a cost comparison that spans the date has two prices in it, and the table's date is what shows it.

Two candidates, and a gate that prices them

Two fixes are worth trying. Candidate A raises retrieval from four passages a turn to eight, so the delivery-areas article makes the cut. Candidate B adds a tool, check_delivery_area, which takes a postcode and returns whether same-day delivery reaches it, plus one line in the prompt saying to call it before promising a delivery time.

Each candidate is an experiment first, and experiments cost money too. The eval runs the 205 test cases five times, the way M9 runs it, and one pass costs about $3.60 in model calls before the judge. Two candidates plus the baseline comes to about $55. That goes on the release record, because a team trying forty candidates a week has an eval bill worth watching, even though it's small next to what the wrong choice costs in production.

The timestamp example from section 5 changed no answers and nearly quadrupled the bill. Nothing in a quality eval catches that, which is the argument for putting cost in the same gate. The gate runs the dataset on every change to any part of the tuple and compares against the baseline on the same dataset version, on four numbers:

  • pass rate, and which individual test cases changed verdict
  • tokens and cost per test case
  • P95 latency
  • cost per resolved ticket, using the handoff rate the change implies

Tools support this directly. promptfoo has assertions for cost ("Inference cost is below a threshold") and latency beside its correctness checks, so a pull request that raises cost per test case above a threshold fails the same way a wrong answer does (promptfoo). The cost per test case then goes through the workload sheet from section 3 to become cost per conversation, which is the number the forecast uses.

The gate on dataset v9
Running nowA, eight passagesB, the area tool
Pass rate192 of 205 (93.7%)197 of 205 (96.1%)197 of 205 (96.1%)
Test cases that changed verdictbaseline5 fixed, 0 broken5 fixed, 0 broken
Mean cost per conversation$0.1443$0.1837$0.1498
Change in costbaseline+27.3%+3.8%
Cost per resolved ticket$0.899$0.939$0.904
Extra spend in a February-size monthbaselineabout $2,400about $340

The quality columns can't separate the two, so the cost rows decide. A's four extra passages are about 1,600 tokens on each of the 4.8 turns, and they're fresh text on every turn, so they get written to the cache at $5 per million. B adds one model call to the 20% of conversations that ask about delivery, plus 150 tokens of tool definition read from the cache on every call. Both carry the delivery_promise field, about 10 output tokens a turn at $20 per million, which adds about $0.001 to every conversation, so the cost of making the promise checkable shows up in the same rows.

The company's gate says that any change raising mean cost per conversation by more than 5% needs a named person to accept the extra spend in the release record. A would need that approval and nobody would give it, since B fixes the same five test cases for about a seventh of the extra spend, so B is the one that goes to the rollout.

The rollout checks the prediction

B goes out through the staged rollout from M11, and the canary reads two signals it didn't read before. One is the field check, which runs on every live conversation because comparing a field with a table costs nothing. It found no out-of-area same-day promises in the first week, where the text search had confirmed 41 in the 30 days before. Those two numbers come from different methods, and the field check only sees what the model put in the field, so the online judge's delivery criterion from v6 keeps reading the text on its 5% sample, and any reply where the text and the field disagree goes to a person. The other is mean cost per conversation, grouped by release id.

Cost came in at +4.7%, where the gate predicted +3.8%. The extra 0.9 points came from traffic, because delivery questions were 26% of conversations in the first week of March and the workload sheet assumed 20%. At 26%, the same arithmetic gives +4.7%. So the thing to fix is the sheet, which now takes the share of delivery questions as a measured input and updates it monthly.

Cost on the release record for 2026-03-09.1
StageCost
Experiments, three bundles run five times on v9About $55 before the judge
Predicted by the gate+3.8%, about $340 in a February-size month
Observed in the first week+4.7%, about $420 in a February-size month

After a few releases, the gap between the last two rows shows whether the team's predictions run high or low, and that's what lets finance trust the next forecast.

Ownership and retirement

Every dependency has an owner and a reason to change

The fix in section 11 made the agent depend on something the agent team doesn't run. The delivery areas come from the logistics team's table, so when logistics adds a village next month, check_delivery_area starts answering differently without anyone on the agent team deploying a thing. The five new test cases may then expect an out-of-date answer. Sculley and colleagues at Google described this in 2015 as an unstable data dependency, an input signal owned by another system that can "qualitatively or quantitatively change behavior over time," and they suggested keeping "a versioned copy of a given signal" that only moves to a new version once someone has checked it (Sculley et al., NeurIPS 2015).

So every dependency gets a named owner, and a named event that makes the owner revise it:

Who owns what the agent depends on
DependencyOwnerWhat makes them revise it
Prompt versionsThe agent teamA reviewed test case fails, or a policy changes
Model id and effortThe agent teamA deprecation notice, or a price change
The area table behind check_delivery_areaLogistics, with a versioned copy the agent readsAn area opens or closes, and logistics tells the dataset owner
Help-center articlesThe help-center editorA policy changes, or a trace shows an article saying the wrong thing
Retrieval settingsThe agent teamA trace shows retrieval missing the article it needed
Eval datasetThe support lead reviews test cases, and the agent team keeps the fileEvery reviewed complaint, and every policy change
Online judgeThe agent team, checked against the support lead's weekly labelsThe judge and the person start disagreeing
Price tableWhoever reconciles the invoice, from section 9A provider changes a price, like Sol's promotion ending after November 21, 2026
Spend guardsThe agent teamP95 cost per conversation moves

The last column is the one that's easy to leave empty. A dependency with an owner and no trigger gets looked at when something breaks, which for the area table means the next complaint.

What each trigger sets off

Each trigger ends in one of three moves, and under pressure they're easy to mix up:

What each signal triggers
SignalMoveCovered in
A reviewed complaint fails in the datasetRevise, with a candidate through the gateSection 11
A canary metric moves the wrong way after a releaseRoll back, fastest lever firstM11
Observed cost exceeds the prediction by more than a point for a weekRevise the workload sheet, then the release if the sheet was rightSection 11
A provider deprecation noticeMigrate, with a full re-run of the frontier tableSection 6
A policy or a delivery area changesRevise the test cases that encode it, and retire the ones that no longer applyThis section
A route costs more per resolved ticket than a person doesRetire the route and send those conversations to a personSection 3
A prompt version has served no traffic for 30 daysRetire it from the registry and keep its historyThis section

Deprecations come on the provider's schedule. Anthropic gives "at least 60 days' notice before model retirement for publicly released models," and after the retirement date "Requests to retired models will fail" (Anthropic, model deprecations, as of September 2026). A migration is a full re-run of the frontier table from section 6, since tokenizers and prices change together, so 60 days is enough only if the dataset and the gate already exist.

A migration that passes the gate still goes through a canary, because the dataset only holds the inputs someone thought to put in it. Shopify fine-tuned a model for its Flow agent and reached benchmark parity, then found at 1% of traffic that "the fine-tuned model's workflow activation rate … came in 35% lower than the prompt-based agent" (Shopify, April 2026). Cost per ticket belongs among the canary's guardrail metrics for the same reason.

Retiring something has its own rules. A test case whose expected answer encodes an old policy gets marked retired, with the date and the reason, and it stays in the file, because scores recorded on earlier dataset versions were measured with it. Retiring it makes a new dataset version, the same as adding one does. A prompt version that nothing has served for 30 days comes off the registry's list of choices, and its text stays, because last quarter's traces still point to it by id.

The same paper names what happens when nobody retires anything. Its term is "underutilized data dependencies," inputs that leave a system "unnecessarily vulnerable to change" even though they "could be removed with no detriment." On an agent, that's a tool nobody calls anymore, whose definition is still sent with every call and still bills input tokens each time.

Putting it together

Putting it together

January's forecast started as one number from a spreadsheet, and by March it's a page anyone can argue with, attached to a release process that keeps it current. Every line has a source and a date:

The February cost plan
LineThe numberWhere it comes from
UnitCost per resolved ticketSection 3
Model cost per conversationMean $0.1443, median $0.0502, P95 $0.6661Section 4
Attempts per call1.031, from a 97% first-try success rateM9
Handoff rate and cost30% at $2.50M7
Cost per resolved ticket$0.899, of which the model is 17%Section 3
February model spendAbout $8,923 at 60,000 conversationsSection 4
Prices assumedSol at $4 / $20 per million, promotional through at least November 21, 2026Section 4
Free levers appliedDrop stale passages from history, cache reads above 85%, batch for the nightly jobSection 5
Trade-off levers consideredSonnet 5 at low effort is cheapest per ticket, and the eval decidesSection 6
CapacityTier arranged in January for 1.7M tokens a minute at the peak, pay-as-you-goSection 7
GuardsPer-conversation budget from P95, per-tenant daily quota, hard spend limit on the projectSection 10
Version tupleCode, prompts, model id, tool schemas, retrieval settings, eval dataset, judge and price table, on every eval run and traceSection 11
GatePass rate and cost per conversation against a baseline on the same dataset version, with a named approver above +5% costSection 11
Cost on each releaseExperiment spend, next to the predicted and the observed changeSection 11
Owners and triggersOne owner per dependency, each with the event that makes them change or retire itSection 12

The forecast is wrong the moment a price changes or the traffic mix moves, which is why it names its assumptions and expects to be revised. The March complaint showed what a revision looks like. The fix went through the gate with its cost predicted. Production came in 0.9 points higher, so the sheet gained a measured input it didn't have in January.

M13 starts the capability modules, where the same loop runs on a classifier's cost profile, and M26 goes under the API for anyone whose numbers make self-hosting look reasonable.

Checkpoint · recall · 5 questions

What the module said

  1. 01

    Why is a per-answer price a bad basis for a monthly forecast?

  2. 02

    Which four meters does one call bill?

  3. 03

    M7's ship bar says "5 cents or less per conversation." What's wrong with it?

  4. 04

    A release record says the candidate passed 96%. Why does it also need the dataset version and the judge version?

  5. 05

    What does a spend alert do when the amount is reached?

0 / 5 answered

Checkpoint · understanding · 5 questions

Reason it through

  1. 01

    Someone adds Current time: {now} to the top of the system prompt to help with delivery questions. What happens to the bill, and why?

  2. 02

    The cheapest model per token produces the most expensive tickets. How?

  3. 03

    Your overnight traffic is about 3 conversations an hour. What does that say about cache time to live?

  4. 04

    A router promises 75% savings at the same quality. What do you check before believing it on your traffic?

  5. 05

    Two candidate fixes both pass 197 of 205 test cases. One raises mean cost per conversation by 27.3%, the other by 3.8%. Why does the gate need a cost threshold to choose between them?

0 / 5 answered

Checkpoint · debugging · 4 questions

Debug it

  1. 01

    Spend per conversation rose 40% overnight. Answers and latency are unchanged, traffic is flat, and no deploy went out. Where do you look?

  2. 02

    The invoice is 15% higher than your own computed total for the same month. What explains the gap?

  3. 03

    The area tool ships on Monday. On Tuesday the online judge's score is three points lower than last week's, but the delivery_promise field check is clean and no complaints came in. What's the first thing to check?

  4. 04

    A pull request changes only the system prompt's wording. The eval's pass rate is unchanged, so it's approved. Two days later spend is up 30%. What should the gate have checked?

0 / 4 answered

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.