The capability
What LLMOps and cost engineering for an LLM system are
LLMOps for a model-backed feature is the loop that keeps it working after launch, where every change goes from a signal through an experiment and a release gate into production, and this module follows that loop through its cost. Cost engineering is the half that knows what one useful outcome costs and which levers change that number, including the ones that also move quality.
The first part of that is picking a unit somebody outside the team would recognise. Tokens are what the provider bills you in, but cost per resolved ticket is what anyone else can act on, and the gap between those two numbers is where most cost surprises turn out to be hiding.
The second is that every call bills on more than one meter. Input tokens, tokens read from the cache, tokens written to the cache and output tokens each carry a different price, and output is usually the most expensive of the four by a wide margin.
The third is that cost per conversation is a distribution with a long tail. The average conversation and the conversation that sets your bill are rarely the same conversation, which means a forecast built on the mean is a forecast built on the wrong conversation.
The fourth is that a cost belongs to a version. Every experiment and every week of production traffic ran on one specific version of the prompt, the model, the tools and the dataset, so a cost number that doesn't say which can't be compared with anything. The last two sections join the pieces M1, M2, M8 and M11 built into one loop. They run a customer complaint through it and attach a cost to each step.

You can't forecast this from a price list alone. What sets the bill is how many calls one outcome takes, how much history gets re-sent on each of them, how much of that history was cached, and how long the answers are. This module works out each of those on a real workload and then goes after them in order.
Where it shows up
Every model-backed feature has this problem the moment it stops being a prototype.
- A support agent whose cost per conversation is 5 to 10 times its cost per model call.
- A coding assistant where a long session re-sends the whole file tree on every turn.
- A document pipeline where the per-page price looks fine until somebody uploads a 400-page PDF.
- An internal chatbot that costs nothing at 50 users and becomes a line item at 5,000.
The arithmetic is the same in all four, and it starts with one number being wrong.
Where this starts
The spreadsheet said $588
It's January of the flower company's second year, and February is the question. Last year the agent ran on half the traffic as a test. This year every ticket goes to it first, which is about 60,000 conversations in the month. Finance wants a number. The company is made up, and so are its numbers, and the prices are real ones read in September 2026.
M7 priced one answer at about a cent. A warm answer on GPT-5.6 Sol, with the system prompt read from the cache and a fresh policy passage written into it, comes to $0.0098. Multiply:
“60,000 conversations × $0.0098 per answer = $588 for February.”
The month's model spend, worked out from what the traces say a conversation does, is about $8,923. That's 15 times the estimate, and none of the gap is a mistake in the per-answer arithmetic. The per-answer number prices one call. February bills conversations.
Six things sit between the two numbers, and each one gets its own section:
A conversation is more than one answer, for a start. The agent averages 4.8 turns, and a turn that calls a tool makes two model calls, so one conversation comes to about 7 calls before anything has gone wrong.
Each of those calls re-sends the whole conversation. M3 established that the model keeps nothing between calls, so every turn carries everything that came before it, which means input grows with the square of the turn count.
Reasoning tokens are billed as output, and there are more of them than of the answer. The visible reply is 200 tokens and the reasoning behind it is another 400.
Cache writes stopped being free. On GPT-5.6 and later, writing a prefix costs 1.25× the uncached rate, and reads cost 0.1× (OpenAI, prompt caching).
Some calls get retried, which bills the input a second time. M9's first-try success rate of 97% works out to about 1.03 expected attempts per call.
And the longest conversations dominate everything. The 7% that run 20 turns or more carry 46% of the spend.
This module builds the model that produces the $8,923, then works through what to do about it. Every number below comes from the same workload sheet the capacity numbers in M11 came from, and every price carries the date it was read.
Unit cost
A conversation costs more than its answer
The first fix is picking the right unit. Tokens are what the provider bills, and they're the wrong thing to manage, because nobody asks for tokens. The ladder runs tokens → one model call → one turn → one conversation → one resolved ticket, and the useful unit is near the bottom. The FinOps Foundation, which writes the vocabulary most finance teams use for cloud spend, describes the same move as going from cost per token to "outcome oriented measures, for example cost per assist, cost per agent action, or cost per case deflected" (FinOps Foundation).
Every call bills four meters
Every call bills four separate things, and a cost model that tracks one number per call will be wrong in both directions:
| Meter | What it is | Price per million tokens |
|---|---|---|
| Uncached input | Tokens the model reads that weren't in the cache | $4.00 |
| Cache write | Tokens stored for reuse, at 1.25× the input rate | $5.00 |
| Cache read | Tokens served from the cache, at 0.1× | $0.40 |
| Output | Everything generated, including reasoning tokens | $20.00 |
The providers report these separately, and the field names differ. Anthropic returns cache_creation_input_tokens and cache_read_input_tokens in the response's usage block, beside its own input count (Anthropic, prompt caching). Check which total includes cached tokens before doing arithmetic on it, since a total that already contains them will double-count. Either way the trace from M2 carries the fields, so cost per call is arithmetic on data you already store.
History is re-sent, so input grows with the square of the turns
M3 established that the model holds nothing between calls. Turn 5 re-sends turns 1 through 4, so the input of a conversation is the sum of a growing series, and its total input grows roughly with the square of its turn count.
| Turns | Input tokens | Output tokens | Cost |
|---|---|---|---|
| 1 | 2,960 | 600 | $0.0176 |
| 2 | 11,760 | 1,660 | $0.0502 |
| 4 | 32,160 | 3,320 | $0.1048 |
| 10 | 145,200 | 8,300 | $0.2894 |
| 20 | 506,400 | 16,600 | $0.6661 |
| 40 | 1,876,800 | 33,200 | $1.6788 |
Forty turns is 1.9 million input tokens for one customer conversation. That's the whole reason a per-answer estimate can't forecast a month.
Caching is what keeps that in check. The same 40-turn conversation costs $8.17 with no caching at all, against $1.68 with it, because every re-sent token is billed at the full input rate each time.
The multipliers outside the tokens
Three more terms sit between a call's tokens and a ticket's cost:
Attempts are the first. M9 measured a 97% first-try success rate, and with up to two retries the expected attempts per call comes to 1.031, with every retry billing the input again.
Handoffs are the second, and they dwarf the model. M7's design sends about 30% of conversations to a person, and a person costs roughly $2.50 of handling time.
Server-side tools are the third. Tools the provider runs for you, like a hosted web search, get billed per call on top of the tokens, so an agent that searches on most turns pays that fee on most turns.
Putting them together for the flower agent:
“cost per resolved ticket = $0.1443 × 1.031 + 0.30 × $2.50 = **$0.899**”
The model is 17% of that. Uber's engineering team describes the same decomposition for its own agents, as "six terms that multiply," naming price per token, tokens per request and requests per turn among them, and notes that "Every turn re-sends the full conversation history, project context, and tool results" (Uber, August 2026). Anthropic's cost guide gives the same instruction in one line, to "Compare on cost per completed task, not per token" (Anthropic).
Forecasting
The bill is a distribution
The $0.1443 above is a mean, and means are dangerous here. Cost per conversation on this workload looks like this:
| Statistic | Value | What it's for |
|---|---|---|
| Mean | $0.1443 | The budget, once multiplied by volume |
| Median | $0.0502 | What a typical customer costs |
| P95 | $0.6661 | The per-conversation guard |
| Share of spend from the longest 7% | 46% | Where the money goes |

The mean is nearly three times the median, because a handful of long conversations carry the month. Anthropic saw the same shape on a 20-problem benchmark run, where "two problems carried 43% of the spend."
That difference decides M7's ship bar, which says "5 cents or less per conversation" without saying which statistic. The median is 5.0 cents, so the typical conversation passes by a hair. The mean is 14.4 cents, so the budget fails by a factor of three. A ship bar has to name the statistic, or two people can read the same dashboard and disagree about whether the feature shipped.
The forecast itself is three numbers multiplied, and the last one is the one people forget:
“February = 60,000 conversations × $0.1443 mean × 1.031 attempts ≈ $8,923”
Prices belong in the forecast too, with their dates. GPT-5.6 Sol's $4 and $20 are promotional, "available at least through November 21, 2026" (OpenAI pricing, read September 2026). A February forecast written in November has to say which price it assumed, because the number changes under it without anyone deploying anything.
Free levers
Cheaper without touching quality
Cost levers come in two groups, and the order is what keeps a team out of trouble. The first group changes how much work the model does without changing what it can do, so nothing has to be re-evaluated. Anthropic's cost guide opens the same way, telling any workload on any model to "Turn on prompt caching and trim unneeded tokens; both are free" (Anthropic).
The biggest lever is still sending less, and M7 found most of it a year earlier by establishing that 60% of tickets needed no model at all. After that comes everything riding along on each call, which is the tool definitions, the retrieved passages and the history. Dropping retrieved passages out of the history once a turn is finished, on the grounds that the answer already used them, takes the mean conversation on the flower agent from $0.1443 to $0.1112. That is a 23% cut with no change at all in what the model can answer.
Next is the cache hit rate. M5 covered how caching works; the number to watch here is cache reads divided by all input tokens, per route, per day. Both providers now charge a premium to write a prefix, so caching only starts paying once that prefix gets read back. OpenAI's own arithmetic, as of September 2026, is that "Writing a prefix once and fully reusing it once costs 1.35× its ordinary input cost, compared with 2× for processing it twice without caching."
Then there is the shape of the output, which is worth attention because output is the expensive meter and two thirds of the flower agent's output tokens are reasoning nobody reads. Asking for shorter answers cuts it, and so does running at a lower reasoning effort, though that second one belongs in the next section because it changes what the model is able to do.
Treat max_tokens as a backstop and nothing more. It does not lower the bill, since you are billed for what gets generated, and setting it too low ends an answer mid-sentence, which M4 covered. The truncated attempt still gets billed.
And batch anything nobody is waiting for. The nightly summaries from M11 run at half price on the provider's batch API, documented as a "50% cost discount compared to synchronous APIs" in exchange for a 24-hour turnaround.
Cache TTL is an off-peak decision
A cached prefix expires. Whether the next conversation finds it warm depends on how often conversations arrive, and traffic at 3 a.m. is a different problem from traffic at noon. With arrivals spread randomly, the chance the cache is still warm is 1 − e^(−λT) for arrival rate λ and time to live T:
| Traffic | 5-minute TTL | 30-minute TTL | 60-minute TTL |
|---|---|---|---|
| 3 conversations an hour, overnight | 22% | 78% | 95% |
| 60 an hour, midday | 99% | 100% | 100% |
| 1,152 an hour, Valentine's peak | 100% | 100% | 100% |
At peak the TTL is irrelevant, and overnight it decides whether every conversation pays the write premium again. Longer time to live costs more to store on providers that price it, so the decision belongs to the quiet hours.
One line at the top of the prompt can undo all of it
The cache matches on an exact prefix. Anything that changes at the top of the prompt invalidates everything below it, so a helpful-looking addition like Current time: {now} turns every call into a full write:
| With a stable prefix | With the timestamp | |
|---|---|---|
| One warm call | $0.0098 | $0.0190 |
| Mean conversation | $0.1443 | $0.5285 |
| February | $8,923 | $32,689 |
No answer changes, no test fails, and the bill nearly quadruples. Section 11 puts a check for exactly this in the release gate.
Trade-off levers
Cheaper by trading quality, measured
The second group of levers changes what the model can do. Each one runs through the M1 dataset before it ships and the M8 canary after, and each one is judged on cost per resolved ticket, since a cheaper model that fails more sends more conversations to a person at $2.50 each.
| Configuration | Cost per conversation | Pass rate | Cost per ticket |
|---|---|---|---|
| Sol, medium effort | $0.2003 | 97% | $1.009 |
| Sol, low effort | $0.1443 | 96% | $0.969 |
| Terra, low effort | $0.0799 | 94% | $0.937 |
| Sonnet 5, low effort | $0.0721 | 96% | $0.894 |
| Haiku 4.5 | $0.0221 | 91% | $0.930 |
| Luna, low effort | $0.0080 | 86% | $1.003 |

The pass rates are invented, and the shape they produce is the point. Luna's tokens cost 18 times less than Sol's and its tickets cost more than Sol's, because ten points of pass rate is worth more than the entire model bill. Anthropic measured the same inversion on its own benchmark, where a model with a per-token price five times higher "solved 88.6% of tasks for $0.54 per solved task, against 77.4% for $0.84" from the cheaper one.
Build this table from the M1 dataset, one row per configuration, and read the frontier: the configurations nothing else beats on both cost and quality. Then pick a point on it with a quality floor, and re-run the table after every model migration.
Routing and cascades
Two ways to use more than one model:
- Route before generating. A classifier picks the model per request. RouteLLM's routers cut cost "by up to 75% as compared to the random router" at the same MT Bench quality. The same paper is honest about where it stops working, since on MMLU "all routers perform poorly at the level of the random router when trained only on Arena dataset," because those questions were outside what the router had seen (RouteLLM). A router is a model trained on a distribution, and your traffic is a different one.
- Cascade after generating. The cheap model answers, a check decides whether to escalate. On the flower agent's workload, sending everything to Luna first and escalating a tenth of conversations to Sol costs $0.0224 a conversation, an 84% saving, and it only stops saving at 94% escalation. The catch is the cache. Switching models mid-conversation means the second model has never seen the history, so at turn 10 it writes about 14,700 tokens at $0.0734 where the first model would have read them for $0.0059.
Reusing answers is a different thing from caching them
Prompt caching reuses computation on an identical prefix and never changes an answer. A semantic cache reuses a previous answer when a new question looks similar, and similar is a threshold someone picks. The vCache paper measured what those thresholds do, and the trade is brutal:
| Similarity threshold | Wrong answers | Cache hit rate |
|---|---|---|
| 0.99 | 2.5% | 37% |
| 0.98 | 4.1% | 53% |
| 0.97 | 5.2% | 67% |
On the flower agent, a 67% hit rate saves about $6.57 per thousand questions and serves about 52 wrong answers in the same thousand. It breaks even only if cleaning up a wrong answer costs less than $0.126. M7 priced a wrong policy claim well above that, so the answer here is no. Where it can work is non-personalized, repeated questions, scoped by tenant and model version, and never for "where is my order."
Limits and throughput
Capacity is part of the price
M11 sized the peak: about 134 model calls a minute in the busiest hour, roughly 1.7M input tokens and 75k output tokens a minute, with bursts around three times the hourly average. Those numbers decide which tier the company has to be on before February, and buying that tier is this module's problem.
Three things make the decision harder than reading a price:
What the limit counts differs by provider, and M11 covered those mechanics. The money consequence is that the same workload can sit right up against one provider's limit and nowhere near another's, and that max_tokens reserves capacity you may never use.
Ramps carry their own limit on top of that. Traffic that triples over a morning gets throttled even while sitting under the ceiling, which is why the tier has to be arranged in January, the month this lesson is set in.
And reserved capacity is priced by what you hold, with no discount for leaving it idle. The flower company's peak hour runs about 35 times an average hour in a normal week, so capacity sized for that peak sits idle almost all year, which is why the answer for a seasonal business like this one is usually pay-as-you-go, with the tier arranged before the peak.
Reserving also promises less than it sounds. Azure's own documentation on provisioned throughput warns that "Unused quota doesn't guarantee that capacity is available when you want to scale back up your PTU deployment. Provisioned capacity is a finite, dynamically changing resource" (Microsoft). Committed spend buys priority, and not a reservation of physical hardware.
Buy vs run
Rent tokens or rent GPUs
Hosting an open-weight model swaps a per-token price for a per-hour one, and per-hour only wins when the hardware is busy. The break-even is one line:
“break-even utilization = (GPU cost per hour) ÷ (tokens per hour at full load × API price per token)”
Put the flower company's numbers in it. February's model spend is about $8,923 across the whole month, and a pair of high-end GPUs rented by the hour runs into the thousands of dollars a month before anyone serves a request. With a peak hour 35 times the average, hardware sized for February 13th idles through the rest of the year, and hardware sized for the average can't serve the peak.
Self-hosting starts to pay when traffic is steady and high, when the task is narrow enough for a small tuned model, or when data rules make the API impossible. Three costs hide behind the hourly rate: cold starts measured in minutes while weights load, the replicas you keep for availability, and the engineering time M11 described as everything you're now on call for. M26 covers the serving internals.
Attribution
Every dollar has an owner
When the bill jumps 20%, the useful question is which feature, which customer or which change did it, and no provider dashboard can answer that. Cost per request exists only in your own logs.
The fix is tagging at the seam. Every call through complete() carries the route or feature, the tenant, a hashed user id, the model and snapshot, the prompt version, the environment and the trace id. M2's traces already hold the token counts, so cost per call is those counts times a dated price table, and every other view is a group-by:
- cost per resolved ticket, by route
- cache hit rate, by route and by day
- P95 cost per conversation, which is what the guard in the next section is set from
- spend by prompt version, which turns "the bill moved on Tuesday" into "v14 costs 31% more than v13"
Reconcile your computed total against the provider's invoice daily, because the gap is informative. It comes from untagged traffic, from contracted prices your table doesn't know about, and from retries you didn't count.
Uber's team published what this discipline buys. Holding the model fixed so the numbers reflect their own work, "cost per 1,000 model requests is down almost 34% from its peak, and cost per session is down 52% from its June peak" (Uber, August 2026).
Spend guards
Caps that bound the damage
M9 stopped a single run from looping forever. A spend guard is the same idea at the level of the account, and the arithmetic makes the case for one. A 200-turn conversation on Sol costs about $22 on its own, and ten thousand scripted 40-turn conversations cost about $16,800 in an afternoon.
The first thing to know is that most budget features are alerts. OpenAI's spend controls, as of September 2026, separate the two plainly, where a spend alert "Sends a notification; API traffic continues," and only a hard spend limit stops traffic, with the warning that "Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount" (OpenAI, spend limits).
So guards go in rings, from the call outward, and each one has a different owner and a different speed:
| Guard | Set from | Who enforces it |
|---|---|---|
max_tokens per call | The longest real answer | Your code |
| Input size check before sending | The window budget from M5 | Your code |
| Per-conversation token or dollar budget | P95 cost per conversation | Your code |
| Step and tool-call caps | M9's stop conditions | Your harness |
| Per-user and per-tenant daily quotas | Normal usage, with headroom | Your API |
| Project rate limits | The capacity plan | The provider |
| Hard spend limit on the project | The month's budget | The provider, with a lag |
| Kill switch to a cheaper path | A plain-language degraded reply | An operator, via M11's flag |
The inner rings act in milliseconds and the outer ones in minutes to hours, which is the argument for having both. Anomaly detection sits alongside them, watching tokens per conversation and spend per tenant against what the forecast expects. M10 covers the attack side, where the bill is the target.
The lifecycle
One complaint, from the support queue to a rollout
Everything so far priced one configuration of the agent. A running feature changes every few weeks, though. Each change starts from some signal and goes through an experiment and a release, and the production traffic after it produces the next signal. M1, M2, M8, M9 and M11 each built one piece of that loop. This section runs one change through all of it in order and puts a cost on each step.
On March 3rd a customer writes in about a bouquet that arrived a day late. The customer had asked the agent whether same-day delivery reached their village, and the agent said yes. The village is outside the same-day area. The complaint and everything that follows are made up, like the company's other numbers, and the prices are the dated ones from section 3.

The complaint becomes a reviewed test case
The support lead finds the conversation from the id on the ticket, and the trace from M2 shows what happened. The agent searched the help center and got back the general same-day policy, and it never saw the article that lists the areas, because that article ranked fifth and retrieval returns four passages. The trace also carries the release id, 2026-02-24.1, which says exactly which bundle answered.
A complaint isn't a test case yet, because complaints can be wrong. Of the 23 conversations support tagged as a wrong answer in February, review found 9 where the agent had been right and the customer had misread the policy. So a person who knows the policy writes down the expected behavior before anything goes into the dataset. Here that's the support lead, and the expected behavior is that the agent checks the postcode before it promises same-day delivery, and offers next-day when the postcode is outside the area. Grading that behavior takes some care, because the reply is free text. A same-day promise can be worded as "yes, we can do today" or "it'll be with them this afternoon", and no string match or regular expression can find every version of it. So the fix gives the promise a structured place to live. Prompt v16 has every reply come back as structured output, the M4 technique where the model fills in a JSON schema, with the text in one field and a delivery_promise field beside it that holds the service promised and the postcode it was promised for, or null when the reply promises nothing. Code compares that field with the area table, which is exact and costs nothing to run.
The field can't catch a reply whose text promises one thing while the field says another, like text that says "this afternoon" next to a null field. So the eval also has a judge read the text for any delivery promise, and a test case passes only when the field is right and the judge finds no same-day promise in the text for an out-of-area postcode. The judge can miss a promise worded in a way it hasn't seen, which is why M1's rule applies and a person checks it, and here that's cheap, since the support lead can read the replies to the delivery test cases by hand.
One complaint is one input, and the failure behind it is usually wider. Production doesn't have the field yet, so finding the same failure in the last 30 days means reading text. A small model reads the replies in every conversation that mentions delivery, about 12,000 of them, and pulls out any delivery promise with its postcode, for about $50. Code checks those postcodes against the area table, and a person reads the 58 out-of-area promises it flags and confirms 41. The model can miss oddly worded promises, so 41 is a lower bound, and none of those customers complained. The lead picks four of them that differ in how the question was asked, like a town name with no postcode, and they join as test cases too. They're picked by how the question was asked, because M1's rule is that test cases chosen by which version gets them right tilt the dataset toward that version. Dataset v8, with 200 test cases, becomes v9 with 205.
Then the version running now is re-run on v9, because M1's other rule is that a score only compares with a score from the same dataset version. It passes 192 of 205, which is 93.7%. Its 96% on v8 is still true, and the two numbers don't compare.
Every score names the versions that produced it
M11 listed what runs in a release (the image, the prompts, the model id, the tool schemas, the index and the flags) and wrote them into a release manifest. A score needs two more entries, because 96% means nothing without the dataset it was measured on and the judge that graded it. Put together, that's the version tuple, and every eval run and every production trace carries all of it:
| Part | Running now, 2026-02-24.1 | Candidate, 2026-03-09.1 |
|---|---|---|
| Code | image sha256:4c7e… | image sha256:b812…, adds check_delivery_area |
| Prompt | policy-answer v15 | policy-answer v16, adds the delivery_promise field and a line on when to call the tool |
| Model id and effort | gpt-5.6-sol-2026-08, low | unchanged |
| Tool schemas | v7 | v8 |
| Retrieval settings | kb-live → build 43, 4 passages | unchanged |
| Eval dataset | v8, 200 test cases | v9, 205 test cases |
| Online judge | v5 | v6, adds a check on delivery-area promises |
| Price table | read 2026-02-02 | read 2026-03-02 |
The judge line is the one teams forget. M9's online judge grades a 5% sample of live conversations, and v6 adds a criterion, so last week's score was graded by different instructions. Re-grade last week's sample with v6 before the release goes out, and use that as the baseline the canary compares against. M9's weekly labels from a person are how you'd notice v6 grading the old criteria differently from v5.
The price table is in the tuple because every cost number in this module is tokens times a dated price. When Sol's promotional price ends, a cost comparison that spans the date has two prices in it, and the table's date is what shows it.
Two candidates, and a gate that prices them
Two fixes are worth trying. Candidate A raises retrieval from four passages a turn to eight, so the delivery-areas article makes the cut. Candidate B adds a tool, check_delivery_area, which takes a postcode and returns whether same-day delivery reaches it, plus one line in the prompt saying to call it before promising a delivery time.
Each candidate is an experiment first, and experiments cost money too. The eval runs the 205 test cases five times, the way M9 runs it, and one pass costs about $3.60 in model calls before the judge. Two candidates plus the baseline comes to about $55. That goes on the release record, because a team trying forty candidates a week has an eval bill worth watching, even though it's small next to what the wrong choice costs in production.
The timestamp example from section 5 changed no answers and nearly quadrupled the bill. Nothing in a quality eval catches that, which is the argument for putting cost in the same gate. The gate runs the dataset on every change to any part of the tuple and compares against the baseline on the same dataset version, on four numbers:
- pass rate, and which individual test cases changed verdict
- tokens and cost per test case
- P95 latency
- cost per resolved ticket, using the handoff rate the change implies
Tools support this directly. promptfoo has assertions for cost ("Inference cost is below a threshold") and latency beside its correctness checks, so a pull request that raises cost per test case above a threshold fails the same way a wrong answer does (promptfoo). The cost per test case then goes through the workload sheet from section 3 to become cost per conversation, which is the number the forecast uses.
| Running now | A, eight passages | B, the area tool | |
|---|---|---|---|
| Pass rate | 192 of 205 (93.7%) | 197 of 205 (96.1%) | 197 of 205 (96.1%) |
| Test cases that changed verdict | baseline | 5 fixed, 0 broken | 5 fixed, 0 broken |
| Mean cost per conversation | $0.1443 | $0.1837 | $0.1498 |
| Change in cost | baseline | +27.3% | +3.8% |
| Cost per resolved ticket | $0.899 | $0.939 | $0.904 |
| Extra spend in a February-size month | baseline | about $2,400 | about $340 |
The quality columns can't separate the two, so the cost rows decide. A's four extra passages are about 1,600 tokens on each of the 4.8 turns, and they're fresh text on every turn, so they get written to the cache at $5 per million. B adds one model call to the 20% of conversations that ask about delivery, plus 150 tokens of tool definition read from the cache on every call. Both carry the delivery_promise field, about 10 output tokens a turn at $20 per million, which adds about $0.001 to every conversation, so the cost of making the promise checkable shows up in the same rows.
The company's gate says that any change raising mean cost per conversation by more than 5% needs a named person to accept the extra spend in the release record. A would need that approval and nobody would give it, since B fixes the same five test cases for about a seventh of the extra spend, so B is the one that goes to the rollout.
The rollout checks the prediction
B goes out through the staged rollout from M11, and the canary reads two signals it didn't read before. One is the field check, which runs on every live conversation because comparing a field with a table costs nothing. It found no out-of-area same-day promises in the first week, where the text search had confirmed 41 in the 30 days before. Those two numbers come from different methods, and the field check only sees what the model put in the field, so the online judge's delivery criterion from v6 keeps reading the text on its 5% sample, and any reply where the text and the field disagree goes to a person. The other is mean cost per conversation, grouped by release id.
Cost came in at +4.7%, where the gate predicted +3.8%. The extra 0.9 points came from traffic, because delivery questions were 26% of conversations in the first week of March and the workload sheet assumed 20%. At 26%, the same arithmetic gives +4.7%. So the thing to fix is the sheet, which now takes the share of delivery questions as a measured input and updates it monthly.
| Stage | Cost |
|---|---|
| Experiments, three bundles run five times on v9 | About $55 before the judge |
| Predicted by the gate | +3.8%, about $340 in a February-size month |
| Observed in the first week | +4.7%, about $420 in a February-size month |
After a few releases, the gap between the last two rows shows whether the team's predictions run high or low, and that's what lets finance trust the next forecast.
Ownership and retirement
Every dependency has an owner and a reason to change
The fix in section 11 made the agent depend on something the agent team doesn't run. The delivery areas come from the logistics team's table, so when logistics adds a village next month, check_delivery_area starts answering differently without anyone on the agent team deploying a thing. The five new test cases may then expect an out-of-date answer. Sculley and colleagues at Google described this in 2015 as an unstable data dependency, an input signal owned by another system that can "qualitatively or quantitatively change behavior over time," and they suggested keeping "a versioned copy of a given signal" that only moves to a new version once someone has checked it (Sculley et al., NeurIPS 2015).
So every dependency gets a named owner, and a named event that makes the owner revise it:
| Dependency | Owner | What makes them revise it |
|---|---|---|
| Prompt versions | The agent team | A reviewed test case fails, or a policy changes |
| Model id and effort | The agent team | A deprecation notice, or a price change |
The area table behind check_delivery_area | Logistics, with a versioned copy the agent reads | An area opens or closes, and logistics tells the dataset owner |
| Help-center articles | The help-center editor | A policy changes, or a trace shows an article saying the wrong thing |
| Retrieval settings | The agent team | A trace shows retrieval missing the article it needed |
| Eval dataset | The support lead reviews test cases, and the agent team keeps the file | Every reviewed complaint, and every policy change |
| Online judge | The agent team, checked against the support lead's weekly labels | The judge and the person start disagreeing |
| Price table | Whoever reconciles the invoice, from section 9 | A provider changes a price, like Sol's promotion ending after November 21, 2026 |
| Spend guards | The agent team | P95 cost per conversation moves |
The last column is the one that's easy to leave empty. A dependency with an owner and no trigger gets looked at when something breaks, which for the area table means the next complaint.
What each trigger sets off
Each trigger ends in one of three moves, and under pressure they're easy to mix up:
| Signal | Move | Covered in |
|---|---|---|
| A reviewed complaint fails in the dataset | Revise, with a candidate through the gate | Section 11 |
| A canary metric moves the wrong way after a release | Roll back, fastest lever first | M11 |
| Observed cost exceeds the prediction by more than a point for a week | Revise the workload sheet, then the release if the sheet was right | Section 11 |
| A provider deprecation notice | Migrate, with a full re-run of the frontier table | Section 6 |
| A policy or a delivery area changes | Revise the test cases that encode it, and retire the ones that no longer apply | This section |
| A route costs more per resolved ticket than a person does | Retire the route and send those conversations to a person | Section 3 |
| A prompt version has served no traffic for 30 days | Retire it from the registry and keep its history | This section |
Deprecations come on the provider's schedule. Anthropic gives "at least 60 days' notice before model retirement for publicly released models," and after the retirement date "Requests to retired models will fail" (Anthropic, model deprecations, as of September 2026). A migration is a full re-run of the frontier table from section 6, since tokenizers and prices change together, so 60 days is enough only if the dataset and the gate already exist.
A migration that passes the gate still goes through a canary, because the dataset only holds the inputs someone thought to put in it. Shopify fine-tuned a model for its Flow agent and reached benchmark parity, then found at 1% of traffic that "the fine-tuned model's workflow activation rate … came in 35% lower than the prompt-based agent" (Shopify, April 2026). Cost per ticket belongs among the canary's guardrail metrics for the same reason.
Retiring something has its own rules. A test case whose expected answer encodes an old policy gets marked retired, with the date and the reason, and it stays in the file, because scores recorded on earlier dataset versions were measured with it. Retiring it makes a new dataset version, the same as adding one does. A prompt version that nothing has served for 30 days comes off the registry's list of choices, and its text stays, because last quarter's traces still point to it by id.
The same paper names what happens when nobody retires anything. Its term is "underutilized data dependencies," inputs that leave a system "unnecessarily vulnerable to change" even though they "could be removed with no detriment." On an agent, that's a tool nobody calls anymore, whose definition is still sent with every call and still bills input tokens each time.
Putting it together
Putting it together
January's forecast started as one number from a spreadsheet, and by March it's a page anyone can argue with, attached to a release process that keeps it current. Every line has a source and a date:
| Line | The number | Where it comes from |
|---|---|---|
| Unit | Cost per resolved ticket | Section 3 |
| Model cost per conversation | Mean $0.1443, median $0.0502, P95 $0.6661 | Section 4 |
| Attempts per call | 1.031, from a 97% first-try success rate | M9 |
| Handoff rate and cost | 30% at $2.50 | M7 |
| Cost per resolved ticket | $0.899, of which the model is 17% | Section 3 |
| February model spend | About $8,923 at 60,000 conversations | Section 4 |
| Prices assumed | Sol at $4 / $20 per million, promotional through at least November 21, 2026 | Section 4 |
| Free levers applied | Drop stale passages from history, cache reads above 85%, batch for the nightly job | Section 5 |
| Trade-off levers considered | Sonnet 5 at low effort is cheapest per ticket, and the eval decides | Section 6 |
| Capacity | Tier arranged in January for 1.7M tokens a minute at the peak, pay-as-you-go | Section 7 |
| Guards | Per-conversation budget from P95, per-tenant daily quota, hard spend limit on the project | Section 10 |
| Version tuple | Code, prompts, model id, tool schemas, retrieval settings, eval dataset, judge and price table, on every eval run and trace | Section 11 |
| Gate | Pass rate and cost per conversation against a baseline on the same dataset version, with a named approver above +5% cost | Section 11 |
| Cost on each release | Experiment spend, next to the predicted and the observed change | Section 11 |
| Owners and triggers | One owner per dependency, each with the event that makes them change or retire it | Section 12 |
The forecast is wrong the moment a price changes or the traffic mix moves, which is why it names its assumptions and expects to be revised. The March complaint showed what a revision looks like. The fix went through the gate with its cost predicted. Production came in 0.9 points higher, so the sheet gained a measured input it didn't have in January.
M13 starts the capability modules, where the same loop runs on a classifier's cost profile, and M26 goes under the API for anyone whose numbers make self-hosting look reasonable.
Checkpoint · recall · 5 questions
What the module said
- 01
Why is a per-answer price a bad basis for a monthly forecast?
- 02
Which four meters does one call bill?
- 03
M7's ship bar says "5 cents or less per conversation." What's wrong with it?
- 04
A release record says the candidate passed 96%. Why does it also need the dataset version and the judge version?
- 05
What does a spend alert do when the amount is reached?
0 / 5 answered
Checkpoint · understanding · 5 questions
Reason it through
- 01
Someone adds
Current time: {now}to the top of the system prompt to help with delivery questions. What happens to the bill, and why? - 02
The cheapest model per token produces the most expensive tickets. How?
- 03
Your overnight traffic is about 3 conversations an hour. What does that say about cache time to live?
- 04
A router promises 75% savings at the same quality. What do you check before believing it on your traffic?
- 05
Two candidate fixes both pass 197 of 205 test cases. One raises mean cost per conversation by 27.3%, the other by 3.8%. Why does the gate need a cost threshold to choose between them?
0 / 5 answered
Checkpoint · debugging · 4 questions
Debug it
- 01
Spend per conversation rose 40% overnight. Answers and latency are unchanged, traffic is flat, and no deploy went out. Where do you look?
- 02
The invoice is 15% higher than your own computed total for the same month. What explains the gap?
- 03
The area tool ships on Monday. On Tuesday the online judge's score is three points lower than last week's, but the
delivery_promisefield check is clean and no complaints came in. What's the first thing to check? - 04
A pull request changes only the system prompt's wording. The eval's pass rate is unchanged, so it's approved. Two days later spend is up 30%. What should the gate have checked?
0 / 4 answered
