MLGuerrillaStart with M1 →
Free · in beta·intermediate·M7·26 min read·Prereq: None. The support agent from Context Engineering (M5) and Data & State (M6) is useful background.

AI Product Framing

The capability

What is AI product framing?

Most AI features start as a request, something like "can we use AI for our support tickets?" Product framing is how you go from a request like that to knowing what you're going to build. It happens before anyone writes code, and what you end up with is usually a single page that says which problem you're solving, how you'll know it worked and what happens when the model gets something wrong.

Now you might be asking why this is an engineer's job at all. Well, because a lot of these questions need someone who knows what models are good and bad at. Whether a task needs a model in the first place, and what a wrong answer would cost, are engineering questions, and they decide whether the project is worth starting.

Framing can't tell you whether the model will be good enough. You only find that out by building it and measuring it on a dataset, which is what M1 was about. What framing does is get everyone to agree on what good enough means before you measure, so when the result comes in, it settles the question.

Where it shows up

This comes up any time someone asks for an AI feature before anyone's looked closely at the problem it's supposed to solve. A request for a chatbot on the help page and a plan to have a model sort incoming invoices both go through the same questions, and sometimes the answer is that plain code does the job better.

Where this starts

A working demo can't tell you whether to ship

Take a company that delivers fresh flowers across the country, for which February is far and away the biggest month of the year. In the week before Valentine's Day its support tickets run at about eight times a normal week, and last year its twelve-person support team fell so far behind that customers waited more than a day for a first reply, many of them asking about orders that had to arrive on the 14th or not at all. The company is made up, and so are its numbers, but the situation is a common one.

In November, the head of support asks engineering for an AI agent that can take over support before February. An engineer builds a demo in about a week, and it is the same kind of agent the last two modules worked with, a model that answers from the help-center articles and can look up an order. In the demo it does well. Asked "where are my roses?", it looks up the order and gives the delivery window. Asked whether the delivery date can change, it quotes the policy. The head of support wants it live by January.

Then the planning meeting asks questions the demo can't answer:

Which tickets would it take once it was live? The demo handled the questions the engineer typed into it, and nobody has looked at what customers send in February.

Is it better than the team? The agent answers in seconds where the team answers in hours, so speed was never the thing in doubt. What nobody knows is how often the agent's answers are right, or how often the team's are.

What does a mistake cost? A wrong answer about vase sizes is cheap. A missed delivery on February 14th cannot be fixed on the 15th, and a refund issued by mistake on a large order is real money leaving the business.

And what does each conversation cost? Every turn is billed by the token, and February brings tens of thousands of conversations.

Two panels. On the left, the one-week demo: six questions the engineer typed, such as where are my roses and can I change the delivery date, all marked as looking right. On the right, what shipping needs to know about February's tens of thousands of real tickets: which tickets the agent would take, whether it is right more often than the support team, what a mistake costs, and what each conversation costs, with none of them measured, priced, or estimated yet. An arrow between the panels is labeled can't answer these.
The rest of this module works through the four questions on the right, starting with which tickets customers send.

A demo shows that a model can give believable answers to the questions someone tried. It can't show which questions customers send, or how often the answers to those are right. M1 described the same problem from the testing side, an AI component that works on the inputs you checked by hand and fails on the ones you didn't. The demo also never raises a more basic question, which is whether a model is the right tool for each of those tickets in the first place.

The work that answers these questions happens before anyone builds, and it's called framing. It comes down to four decisions:

  • what problem the team is solving, and for whom
  • which parts of that problem need a model, and which parts plain code handles better
  • the smallest version worth shipping first
  • what success means in numbers, agreed before any results come in

These decisions apply to any AI feature, from a coding assistant to a document pipeline. The flower company's support agent is the example this module works through. Hiring teams call this skill product sense, and it shows up in about a third of the AI engineering job postings this course is built from.

Is this an AI problem?

Start from the problem, then check whether it needs a model

Before deciding what to build, the engineer pulls 500 tickets from last February and sorts them by what each customer wanted. Labeling them by hand takes an afternoon, and it's the same dataset-building step from M1, used here to decide what to build. The shares below are illustrative:

  • Where's my order? 45% of tickets, most of them asking whether the flowers will arrive in time.
  • Change the delivery date or address. 15%, most of them sent before the order goes out for delivery.
  • Arrived damaged or wilted. 20%, usually with a photo attached.
  • Refunds and charges. 10%, like a customer who was charged twice.
  • Everything else. 10%, from policy questions to corporate orders.

For each group, four questions decide whether a model belongs there:

The first is whether the input is free-form language. A question typed in the customer's own words needs something that can read language, whereas a button press or a form field does not.

The second is whether somebody could write the rule. If a person can write down exactly how to get to the answer, then code can follow that rule every time and will not get it wrong.

The third is what a wrong answer costs, because the higher that cost goes the more checking the answer needs, whoever or whatever produced it.

The fourth is how many of these arrive. A large group is worth fixing first even when the fix is a simple one.

The answers point to one of these options, listed from cheapest and most predictable to most flexible:

  1. 01plain code, like a tracking page or a form
  2. 02search over the help-center articles
  3. 03a classifier, which is a model that sorts each ticket into a fixed set of labels
  4. 04an LLM that writes an answer
  5. 05an LLM agent that also calls tools, with a person approving anything costly

For February's five groups, the answers come out this way:

Take "where's my order" first. The courier's tracking data already contains the answer, so a tracking link in the confirmation email, plus an automatic text when a delivery is running late, answers most of these before anybody types anything.

Delivery changes are similar. The rules are known, including the cutoff after which an order has gone out for delivery and can no longer change, so a self-serve form applying those rules handles them, again with no model involved.

Damaged or wilted flowers are the first group where a model earns its place. Somebody has to look at the photo and decide on a refund, and a vision-language model, meaning one that reads images along with text, can do the looking. A person then approves the refund.

Refunds and charges need more than that. "I was charged twice" requires something that reads the customer's words and checks their account, so this group gets an agent with tools, with approval before any money moves.

Everything else is free-form policy questions, which need search and an LLM working together. Search finds the help-center article that covers the question, and the LLM writes the answer from it.

A classifier sits in front of all five groups, sorting each new ticket into its group before any of these paths runs. Later modules on classification and routing cover that step.

A card with five rows, one per ticket type from a hand-labeled sample of 500 February tickets. Where's my order is 45 percent and is handled by a tracking link plus a late-delivery text. Delivery date or address changes are 15 percent and handled by a self-serve form with cutoff rules. These two rows are bracketed as needing no model, 60 percent of volume. Damaged or wilted flowers are 20 percent and handled by a vision model that checks the photo with a person approving the refund. Refunds and charges are 10 percent and handled by an agent that reads the account with a person approving. Everything else is 10 percent and handled by search plus an LLM answering from help articles. These three rows are bracketed as needing a model, 40 percent. A strip below lists the options from cheapest to most flexible, plain code, search, classifier, LLM, and an agent with tools.
The two plain-code groups make up 60% of the volume, so they can ship before any model work starts.

The first two groups add up to 60% of February's volume, and neither needs a model. That changes the plan, because the tracking texts and the change form can ship first and should take most of that 60% off the support queue before February. The agent is left with the 40% that does need a model.

It's tempting to send every ticket through the LLM anyway, since it can answer all five kinds. For the first two groups that adds a cost to every ticket and a few seconds of waiting, and it adds a way to be wrong. The tracking system already knows when the order will arrive, and a model that reads the tracking data and rewrites it in its own words can misread it. M4 also showed that the same prompt can come back differently from one run to the next. For questions where a rule gives the exact answer, plain code is both cheaper and more reliable.

Knowing which 40% needs a model still leaves a lot to build, so the next decision is which piece to ship first.

Scoping v1

Pick the smallest version worth shipping

The 40% that needs a model is still three groups, and building all three before February would mean shipping everything at once and then finding out about everything at once. So one group goes first, and four things decide which one it is.

How much volume a group covers matters, since a group with more tickets pays the work back sooner.

What a mistake costs matters more. A wrong sentence about a policy can be corrected in the next reply, while a wrong refund has already moved money.

How easily an answer can be checked decides how much of the work you can automate afterwards. A policy answer can be compared against the article it came from, whereas a judgment about a wilted bouquet needs a person to look at the photo.

And whether the data already exists decides how long it takes to start. Past tickets and help-center articles are sitting there already. Labeled photos of damaged flowers are not.

Lined up against those, the three groups come out in a clear order. Policy questions are the cheapest to get wrong and the easiest to check. They also run on data the company already has. Damaged flowers need photo judgment and a refund decision. Refunds and charges touch money directly and need approval before anything moves.

So v1 answers policy questions and hands everything else to the support team. v2 adds photo triage for damaged orders, with a person approving each refund. v3 takes on refunds and charges. Only v1 has to be ready for February.

Writing down what v1 leaves out matters as much as what it includes, because non-goals are what keep a two-month project from becoming a six-month one:

  • no refunds, and no promises about money
  • no decisions from photos
  • no commitments about delivery times, which come from the tracking system

Handing off to a person belongs in v1 from the start. When the agent has no article that covers the question, or the customer asks for a refund, the conversation goes to the support team with the ticket history attached. Without a handoff, v1 would have to be right about everything.

Three release rows beside a non-goals panel. v1 in January covers policy questions with search and an LLM over help articles, marked ready for February. v2 after February covers damaged or wilted photos with a vision model and a person approving, once v1 has numbers. v3 later covers refunds and charges with an agent and a person approving, where money moves. The panel lists what is not in v1: no refunds or promises about money, no decisions from photos, and no commitments about delivery times. A strip underneath describes the v1 handoff, where no matching article or any mention of a refund sends the conversation to the support team with its history attached.
Only v1 has to be ready for February. v2 and v3 wait until v1 has produced numbers.

A smaller v1 also makes the rest of the work smaller. The eval set from M1 only has to cover policy questions. The traces from M2 only have to explain one path. And the decision to expand rests on numbers from real February traffic.

The head of support still wants the whole thing by January, and framing gives an answer to that. v1 ships in January, and February measures it under the heaviest load of the year. v2 starts once there are numbers to argue from.

Success metrics

Define success before you build

Once v1 is running, someone will pull a number out of it and call the project a success. Deciding which number that is belongs before the build, while nobody knows yet how the results will come out. Afterward, whichever number looks best tends to become the one that mattered all along.

Four kinds of numbers do different jobs, and v1 needs all four:

The outcome metric is what the support team cares about. For v1 that is the share of policy questions the agent answers without a person stepping in, and how long a customer waits for a first reply.

The quality metric is correctness against a labeled set of questions, built the way M1 describes. For v1 that means whether the answer matches the help-center article it should have come from.

The guardrail metrics are the things that are not allowed to get worse while the outcome metric improves. Wrong claims about policy sit here, and so do refund requests the agent failed to hand off.

And the ship bar turns all of those into thresholds that decide whether v1 goes live at all, with that decision made before launch.

For the flower company, an illustrative ship bar for v1 reads like this:

  • at least 95% of answers on a 200-question eval set match the help-center article
  • every ticket that mentions a refund is handed to a person, with no exceptions
  • cost stays under five cents per conversation
  • first-reply time in February beats last February's thirty hours

Two of those come from the eval set before launch, and two can only come from live traffic in February. Splitting them apart is worth doing early, because the ones that need live traffic decide how long v1 has to run before anyone can judge it. Running that comparison properly, against what the team would have done anyway, is what M8 covers.

The number to be most careful with is the one that looks most like success. Counting tickets the agent "deflected" rewards any conversation where the customer stopped replying, including the customer who gave up and drove to a florist. A resolved ticket is one where the customer got what they asked for and didn't come back about it a day later, and that's the version worth putting in the ship bar.

Four stacked rows and a side panel. Outcome is what the support team wants, policy questions answered without a person and first-reply time, measured on live traffic. Quality is whether the answers are right, matching the help-center article on a labeled set, measured on the eval set. Guardrails are what can't get worse, wrong policy claims and refund requests never handed off, measured on both. The ship bar, outlined in terracotta, lists at least 95 percent of 200 eval answers matching the article, 100 percent of refund mentions handed to a person, five cents or less per conversation, and a first reply faster than last February's thirty hours. The side panel shows tickets deflected struck out, replaced by got what they asked for and didn't come back a day later.
Everything here is agreed in November, so February's numbers only have to be read off.

The last part of defining success is agreeing on it out loud. The engineer and the head of support both sign off on the same numbers in November, so February's argument can be about whether the bar was met.

Baseline and cost

Compare against what happens today

A ship bar with a cost in it needs two numbers behind it. One is what an answer costs today and the other is what it would cost through the agent, and neither of those is obvious until somebody sits down and works it out.

Today's number comes from the support team, where a policy question takes about six minutes of an agent's time. At a fully loaded cost of around 25 dollars an hour, that comes to about 2.50 dollars per ticket. Those figures are invented for the example, and every company has its own, usually sitting in a spreadsheet the support lead already keeps.

The agent's number comes from the tokens. A policy answer sends the stable instructions along with a few retrieved passages and the customer's question, which lands around 3,000 input tokens, and it writes about 200 tokens back. Most of that input is the stable prefix M5 described, so after the first call it's served from the prompt cache. On OpenAI's published prices for GPT-5.6 Sol as of September 2026, input is $4.00 per million tokens, cached input is $0.40, and output is $20.00. The 1,000 fresh tokens are billed as a cache write at $5.00, which is 1.25× the input price, because OpenAI stores them for the next call to reuse:

One answered policy question
TokensCountPrice per 1M tokensCost
Input, cached2,000$0.40$0.0008
Input, written to the cache1,000$5.00$0.0050
Output200$20.00$0.0040
Total per answer$0.0098

Under a cent per answer, and a conversation that takes two turns is still under two cents. The first call of the day, before anything is cached, writes all 3,000 input tokens at $5.00 and comes to about $0.019, so an answer in a quiet hour, after the cache has gone cold, costs about twice what it does in a busy one. Those Sol prices are also a promotion that OpenAI says runs at least through November 21, 2026, so check the pricing page before you reuse the arithmetic. Against 2.50 dollars of agent time, none of that changes the ratio enough to matter.

What makes the comparison honest is the tickets the agent hands off. If v1 answers seven policy questions in ten and sends three to the team, those three still cost the full 2.50 dollars each, plus the pennies the agent spent before giving up. Blended across the group, a policy ticket costs about 77 cents where it used to cost 2.50, which is a real saving and a smaller one than the per-answer number suggests.

Three bars comparing the cost of one policy ticket. Answered by the support team today, six minutes at 25 dollars an hour, 2 dollars 50. One answer from the agent, tokens only with nothing handed off, under a cent at 0.0098 dollars, a bar too small to see. A blended v1 ticket with seven answered and three handed to the team, 77 cents, drawn in terracotta. A side panel shows where the 0.0098 comes from: 2,000 cached input tokens at 40 cents per million, 1,000 fresh input tokens written to the cache at 5 dollars per million, and 200 output tokens at 20 dollars per million. A strip underneath lists the costs outside the arithmetic: writing and updating the eval set, reading traces through February, and two months of one engineer.
The per-answer price and the per-ticket price are different numbers, and the ship bar should name which one it means.

Three costs sit outside that arithmetic and belong in the same conversation:

Somebody has to build and keep the eval set, which means writing the 200 labeled questions and then updating them every time the help-center articles change.

Somebody has to watch it in production, using the traces from M2, and that somebody has to be a real person reading them during February.

And there is the engineering time itself, which is two months of one engineer, competing for that engineer against the tracking texts and the change form that handle 60% of the tickets.

That last one is the point most cost conversations miss. The tracking texts remove more tickets per engineering day than the agent does, which is exactly why they shipped first in the plan. Cost framing almost always compares the model against people, when the comparison that changes the plan is against the other things the same engineer could build in those two months.

Failure and handoff

Decide what happens when the model is wrong

A model that answers policy questions will sometimes answer one wrong. M4 covered why the same prompt can come back differently on different runs, and no ship bar of 95% says otherwise, since what a 95% bar says is that the other 5% is coming. What framing decides is what that 5% costs you when it does.

Sorting v1's failures by what each one costs puts them in an order:

At the cheap end is an unhelpful answer. The customer asks again, or asks for a person, and it costs a little patience.

Above that is a wrong detail. The agent says the cutoff for changing an order is 24 hours when the policy says 48, so the customer plans around a wrong number and the team fixes it later.

Worse is a missed handoff. Someone asks for a refund and the agent answers the policy question without passing it on, so the customer sits waiting for money that nobody is working on.

Worst is an invented policy. The agent states a rule the company does not have, the customer acts on it, and now the company is arguing about a promise it never made.

That last one has a price tag somebody has already paid in public. Air Canada's website chatbot told a passenger he could apply for a bereavement fare retroactively, which its own policy didn't allow. The airline argued it shouldn't be responsible for what the chatbot said. A British Columbia tribunal disagreed and ordered it to pay the difference, CA$650.88, in Moffatt v. Air Canada. A support agent's answers are the company's answers.

Each level of cost buys a different control, and picking one per failure is the framing decision:

The loosest is to let it answer freely, which is fine wherever being unhelpful is the worst thing that can happen.

Tighter than that is answering with the source, where the reply links the help-center article it came from. That lets the customer check the claim and lets the team see where a wrong answer originated.

Tighter again is answering only from what was retrieved. If no article covers the question, the agent says so and hands off, and that single rule is what keeps invented policies out.

Tighter still is asking a person first, so that anything moving money waits for approval, the way M6's refund run waited.

And the tightest is leaving it to people entirely, which is why refunds are a non-goal in v1.

Four failure rows with a cost rating and the control v1 applies to each. An unhelpful answer, where the customer asks again, rates one dot and is answered freely. A wrong detail, saying 24 hours when the policy says 48, rates two dots and is answered with a link to the article it used. A missed handoff, a refund request answered as a policy question, rates three dots and is handled by routing refund words to a person. An invented policy, a rule the company does not have, rates four dots in terracotta and is prevented by answering only from retrieved articles or handing off. A strip underneath records that Air Canada's chatbot promised a bereavement fare its policy did not allow and a tribunal ordered the airline to pay CA$650.88, in Moffatt v. Air Canada, 2024 BCCRT 149.
The control column is also the list of rules the agent enforces at runtime, and each one shows up in February's guardrail numbers.

The handoff rules follow from the same list. A refund request goes to the support team, and so does a question with no matching article or a customer who asks the same thing twice. Every reply carries a way to reach a person. Each of those rules is also a line in the guardrail metrics from earlier, so February's numbers show how often each one fired.

One more thing makes a wrong answer fixable. Every answer is logged with the article ids it used, in the traces M2 described, so a complaint about a wrong cutoff leads back to the passage the agent read and the article that needs rewriting.

The framing doc

Write the framing down on one page

Everything decided so far lives in people's heads and in a few meetings. In November that feels like enough. By February, the head of support remembers a promise about handling refunds while the engineer remembers agreeing to policy questions. Nobody can point at the version they agreed on. One page fixes that, and it's short enough that people read it.

Seven headings cover it:

  • The problem. Who has it, and what it costs them today.
  • The tasks. The ticket groups with their shares, and which of them need a model.
  • Scope. What v1 covers, and the non-goals that keep it there.
  • Success. The outcome, quality and guardrail numbers, with the ship bar.
  • Baseline and cost. Today's cost per ticket and the expected cost per ticket.
  • When it's wrong. The controls and the handoff rules.
  • Risks. What could make the plan wrong, and what would show it early.

Filled in for the flower company, it fits on one screen:

Support agent v1, framing, November
HeadingWhat it says
ProblemFebruary runs 8x normal ticket volume. Last year first replies took 30 hours during Valentine's week, on orders that must arrive on the 14th.
TasksSample of 500 February tickets: 45% where's my order, 15% delivery changes, 20% damaged photos, 10% refunds, 10% other. 60% needs no model, so the tracking texts and the change form ship first.
Scopev1 answers policy questions from help-center articles and hands off everything else. Non-goals: no refunds, no photo decisions, no delivery promises.
SuccessOutcome: share of policy questions answered without a person. Quality: at least 95% of 200 eval answers match the source article. Guardrails: 100% of refund mentions handed off, CSAT not below last February, 5 cents or less per conversation. Ship bar: all four, measured before launch where possible.
CostToday about $2.50 per policy ticket, 6 minutes of agent time. v1 about $0.77 blended, assuming 3 in 10 still hand off.
When wrongAnswer only from retrieved articles, link the source, route any refund wording to a person, always offer a human.
RisksHelp-center articles are out of date, and answers inherit that. February volume may differ from last year, so re-check the shares in January. A handoff rate above 5 in 10 makes the saving disappear, so revisit the scope.

The value of the page is in the arguments it ends. When someone asks in January why the agent can't issue refunds, the non-goals line answers it. When February's numbers come in, the ship bar says whether v1 passed, using thresholds written before anyone saw a result.

It's also the artifact interviews ask about. A question like "how would you decide whether to build this with an LLM" is asking for the reasoning on this page. Answering it with ticket shares and a ship bar is what product sense sounds like in practice.

The page keeps working long after November, which is the point of writing it down. The shares get re-checked in January against fresh tickets, the ship bar gets its results written next to it in March, and v2's scope then starts from whatever February turned out to show.

Putting it together

Putting it together

The module started with a demo that answered six questions well and a room full of people who couldn't say whether to ship it. Everything since has been the work that answers that question before the build.

Sorting a sample of real tickets came first, because it turned "an AI agent for support" into five groups with shares attached, and showed that 60% of February's volume needed no model at all. Scoping came next, picking the group with the cheapest mistakes and the easiest checking, and wrote down the parts v1 leaves alone. Then the numbers that decide shipping, agreed in November so that February can only confirm or fail them. Then the cost of a ticket today next to the cost of a ticket through the agent, with the handoffs counted honestly. Then a control for each way the agent can be wrong, sized to what that mistake costs. All of it fits on one page that the team can argue with in January.

Later modules pick up each thread. M1 builds the eval set the quality number depends on. M8 runs the comparison that tells you whether February's improvement came from the agent. M13 and M14 cover the classifier that routes each ticket to its group, and M19 covers making the search behind the policy answers good. M25 works through the architecture choices this module only sketched, and M29 covers carrying decisions like these through a company.

The habit underneath all of it is small. Before building an AI feature, find out what people are asking for and decide which parts of it need a model. Then say what number would make it worth shipping.

Checkpoint · recall · 5 questions

What the module said

  1. 01

    What does framing decide, before anyone starts building?

  2. 02

    February's tickets were sorted into five groups. Which option handled "where's my order?"

  3. 03

    What is a ship bar?

  4. 04

    Why does v1 write down its non-goals?

  5. 05

    What did the Air Canada tribunal decision establish about the airline's chatbot?

0 / 5 answered

Checkpoint · understanding · 5 questions

Reason it through

  1. 01

    A team wants an LLM to answer "where is my order?" using data from the tracking system. Why is plain code the better choice?

  2. 02

    After launch, tickets "deflected" by the agent are up sharply, while customer satisfaction has dropped. What is the most likely explanation?

  3. 03

    One answered policy question costs under a cent, but the blended cost per policy ticket is about 77 cents. Where does the difference come from?

  4. 04

    Three groups need a model: policy questions, damaged-flower photos, and refunds. Why does v1 take policy questions?

  5. 05

    What keeps a support agent from stating a policy the company doesn't have?

0 / 5 answered

Checkpoint · debugging · 4 questions

Debug it

  1. 01

    A feature shipped in January. In March, one person says it worked and another says it didn't, and both point at real numbers. What was skipped?

  2. 02

    In January, a stakeholder asks the team to add refunds to v1 "since the agent is nearly done." What in the framing answers that?

  3. 03

    February's real cost per policy ticket comes in far above the estimate, though the token cost per answer matches the projection. What changed?

  4. 04

    A customer asks about a policy the help center has never covered, and the agent answers with a confident rule that doesn't exist. Which rule was missing?

0 / 4 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.