The capability
What fine-tuning is
As you already know from M4, the models you build with were trained by somebody else. A lab collected an enormous amount of data, trained on it for weeks on thousands of GPUs, and published the result, which is a big file of numbers called the weights. When you call the model, those weights are what turn your input into an output, and so far in this course you've changed the output by changing the input, with a better prompt or with documents pulled in by retrieval.
Fine-tuning changes the weights themselves. You take the pretrained model and keep training it on a dataset of your own examples, where each example is an input paired with the output you wanted for it. The training loop is the same one the lab ran, just on far less data and starting from a model that's already good. The model reads an example and produces an output, and your training code compares that output with the one you wanted and turns the difference into a single number called the loss. Then it nudges every weight a tiny amount in whichever direction would have made the loss smaller. It goes through every example in the dataset that way, usually several times, and one full pass over the dataset is called an epoch.
What comes out is a new copy of the weights, or with some methods a small extra file that sits on top of the original weights, which is called an adapter. Either way you still have the same kind of model with the same inputs and outputs, so a fine-tuned text classifier still takes text and returns a label, and a fine-tuned LLM still takes a prompt and writes text back. The difference is in how it behaves on inputs that look like your examples.

Now you might be asking what "the output you wanted" even looks like, since an LLM writes text and a classifier returns a label. It depends on the model, and it's the biggest difference between fine-tuning one kind of model and another. For an LLM, it's the text you wanted it to write. For an embedding model, it's which pairs of inputs should end up close together and which should end up far apart. For a classifier it's the correct label, and for a detector it's the correct box. That target, together with the loss that compares the model's output against it, is called the training objective. This module covers what stays the same whatever the objective is, and M41 to M43 each go deep on one objective.
Fine-tuning is good at changing how a model behaves, like the format it answers in or the labels it picks, but it can't reliably teach a model new facts. Gekhman et al. fine-tuned an LLM on questions whose answers it hadn't learned in pretraining, and found that those examples "are learned significantly slower than those consistent with the model's knowledge", and that once the model did learn them, they "linearly increase the model's tendency to hallucinate" (Gekhman et al., 2024). That was a controlled study on closed-book question answering, so it doesn't cover every kind of fine-tuning, but it's a good reason not to count on fine-tuning for facts. If you want a model to know your product catalogue or this week's prices, that's a retrieval problem, which M19-1 covers.
I learned this the hard way when LLMs first came out. In 2023 I tried to fine-tune one on legal documents, because I wanted a model that could write a contract from start to finish and wouldn't forget what was in those documents. It didn't go well. With a lot more data it might have picked up the right words and the right structure, but most of what makes a contract right is knowledge, and that knowledge has to stay up to date, which isn't something fine-tuning puts into a model. What surprised me was that the fine-tuned model ended up doing worse than the regular LLM I'd started from. That's one project from 2023, so it doesn't prove much on its own, but it lines up with what Gekhman et al. measured a year later.
Fine-tuning also learns whatever your dataset says, mistakes included, so a dataset with mislabeled examples gives you a model that makes those same mistakes, and makes them confidently. And once you've fine-tuned a model, you own it, which means testing it on every change and retraining it when your data changes or when the base model you started from gets retired. Somebody also has to host it, either you or a provider you pay.
Some places it shows up
- A classifier on your own labels, like sorting product photos into a marketplace's own categories, which no general image model has seen named that way. M13 shows the cheapest version, a small linear model trained on top of a frozen encoder.
- A search embedding model trained on your users' queries and the results they clicked, so that "similar" means what it means for your users. M42 covers how.
- A small LLM that writes in one fixed format, trained on outputs from a bigger model so it does the same job faster and for less money. Training a small model on a bigger model's outputs is called distillation, and M41 covers it.
- A detector for your own objects, like the Blender icons in M20, trained on boxes a VLM drew. M43 covers how.
These are examples, and the pattern is the same in all of them. You start from a pretrained model that's close to what you need and a dataset of examples showing the behavior you want, and then an eval tells you whether the new weights do better than the old ones on the inputs you care about.
The running example
Relist's pipeline failed in three places, and each was a different kind of model
Relist is a secondhand marketplace for furniture and home goods. It's invented, and so are its numbers. A seller uploads a few photos and a line or two of text, and a pipeline turns that upload into a listing.
- It picks a category from Relist's own list of 240, like "nightstand" or "mid-century sideboard", by looking at the photos.
- It finds similar items already on the site, for the "you might also like" row under every listing.
- It writes a title in Relist's house format, like "IKEA Hemnes nightstand, pine, white".
The first version of the pipeline trained nothing. SigLIP, the image encoder M13 showed for zero-shot image classification, scored each photo against the 240 category names and picked the closest one. SigLIP's embeddings also found the similar items, by looking up the listings whose photos had the closest vectors, and a large hosted LLM wrote each title from the seller's photos and text, with a prompt that had five example titles in it. That version worked well enough to launch, so Relist shipped it.
Then the problems showed up, one per job. On a test dataset of 2,000 photos that the team labeled by hand, SigLIP picked the right category 71% of the time, and most of its mistakes were between categories that Relist separates and the rest of the world doesn't, like nightstand and side table. The "you might also like" row put things next to each other because they looked alike, so a brass lamp got a brass vase and a brass candle holder next to it, when buyers who click on a lamp want other lamps. And the titles were good, with the large LLM getting the house format right on 97% of listings, but Relist gets about 60,000 uploads a day, so the title step alone came to roughly $7,000 a month and added four seconds to every upload. A small model with the same prompt was cheap and fast and got the format right only 78% of the time.
So an engineer on the team proposed fine-tuning all three, which is three separate projects, because each of those models learns from a different kind of target. The classifier learns from photos paired with their correct category. The embedding model learns from pairs of listings that belong next to each other, and the title writer learns from uploads paired with the title each one should have gotten.

Before any of that, though, it's worth checking whether each one needs fine-tuning at all, because some of those problems have cheaper fixes.
Before you train
Try the fixes that don't need training first, and fine-tune what's left
Fine-tuning is the most expensive fix you can reach for. You need a labeled dataset before you start, then a training run, then an eval to find out whether it worked, and after all that you own a model you have to keep testing and hosting. Most of the cheaper fixes take an afternoon. So the usual order goes from cheapest to most expensive, and you stop at the first step that gets you the number you need.
- 01Improve the prompt or the label descriptions, and put a few examples in the context.
- 02Change the system around the model so it has less to get right, like filtering its candidates in code or asking for structured output.
- 03Add retrieval, if what's missing is information the model doesn't have.
- 04Try a bigger or newer model, which tells you whether the task can be done at all. If the strongest model does it well, you know it's possible, and distilling that model into a small one becomes an option.
- 05Train a small head on frozen features. The head is a small final layer that turns the model's vectors into your output, and training only that is the cheapest kind of training, because the pretrained model doesn't change at all.
- 06Fine-tune the model itself, starting with an adapter, and train every weight only if the adapter isn't enough.

Relist went up that list one job at a time, and each job stopped somewhere different.
The category classifier got most of the way with a head on frozen features
Relist's first move was to rewrite the category names as descriptions, the way M13 did for zero-shot classification, so "nightstand" became "a photo of a nightstand, a small bedside table with a drawer". That moved SigLIP from 71% to 75% on the 2,000 hand-labeled test photos. A larger SigLIP checkpoint added two more points, to 77%. Neither touched the main problem, because nightstand and side table are Relist's own categories with Relist's own rule, which is that anything under 70 centimetres with a drawer counts as a nightstand. No description in a prompt gets a general model to apply a rule like that reliably.
So the team labeled 30 photos per category, 7,200 in all, ran every one through SigLIP to get its embedding, and trained a linear classifier on those embeddings, the frozen-features approach from M13. It trained in about a minute on a laptop and got 88%. That's a big jump for almost no cost, and for a lot of teams it would be good enough to stop there.
The errors that were left had a pattern, though. Photos of nightstands and side tables got nearly the same vectors from SigLIP, because SigLIP was never trained to care about the difference. A linear head can only draw boundaries between the vectors it's handed, so when two categories get the same vectors, no head can separate them. When that's the problem, the encoder itself has to change, which is fine-tuning.
The similar-items row needed a filter before it needed training
The "you might also like" row didn't need training to fix its worst problem. The brass lamp got brass vases because the search was finding things that looked alike across the whole site. Once the row only searched within the listing's predicted category, lamps got lamps, and the share of buyers who clicked something in the row went from 2.1% to 3.0%, and the filter is a few lines of code.
What the filter couldn't fix was what Relist's buyers mean by similar, which turns out to be the same style at a similar price. A general embedding model doesn't know that, but Relist has eighteen months of records of which listings the same buyer clicked in one session, and that's the kind of data an embedding model is fine-tuned on. M42 covers how.
The title writer needed structured output, and then distillation
The small model's format problem went away without training. The team stopped asking it for a finished title and asked for a JSON object with one field per part of the title, checked against a schema, and the code put the title together from the fields. The format was then right on every listing, because code wrote it.
With the format handled, the team measured how often every field was right, and the large model got all of them right on 94% of listings while the small one got 79%, mostly by guessing brands and materials from the photos. A bigger prompt didn't close it. What Relist did have was a large model that already did the job well, and that's the setup for distillation, where the large model labels thousands of uploads and the small model is fine-tuned to copy it. The small model then runs at the small model's price and speed.
Signs that fine-tuning is the right fix
| What you see | Fine-tune? | Why |
|---|---|---|
| The model needs facts or anything that changes, like prices | No | Retrieval puts the current facts in front of it, and fine-tuning learns facts slowly and badly |
| You can't yet say what a good output looks like, or you have no test dataset | Not yet | Without an eval you can't tell whether training helped or hurt |
| You have fewer than about 50 good examples | Not yet | Try prompting with the examples, and label more first |
| Labels or boundaries that are yours alone, and a head on frozen features has stopped improving | Yes | The model's features don't separate what you need separated |
| A large model does the job well, but it costs too much or takes too long | Yes | Distill it into a small model |
| Your inputs look nothing like what the model was trained on, like X-rays or UI icons | Usually | The encoder has to learn which details count in your kind of input |
What to train
Change as little of the model as gets you the result
Most pretrained models you'll fine-tune have two parts. The backbone is the big pretrained part that turns an input into vectors, which are lists of numbers describing the input, and the head is a small final layer that turns those vectors into the output you want, like a category. M32 covers the backbone-and-head idea in depth. When you fine-tune, you choose how much of that you let training change, anywhere from only the head to every weight. Changing less makes training cheaper and safer for what the model could already do, but it also limits how much the model can learn.

Train only a new head, and keep the backbone frozen
This is the frozen-features approach from section 3 and M13, and it's also called linear probing. The backbone runs once over every example to get its vectors, and then you train just the head on those vectors, which takes seconds to minutes and needs no GPU. Since the backbone never changes, nothing it could already do gets worse. The limit is the one Relist hit, which is that the head can only work with the vectors the backbone produces, so if the backbone puts two of your categories in the same place, the head can't pull them apart.
Train a small adapter next to the frozen weights
An adapter is a small group of new weights added beside the frozen ones, and the most common kind is LoRA, short for low-rank adaptation (Hu et al., 2021). Inside a transformer, most of the weights sit in big square tables of numbers called weight matrices. LoRA leaves each of those frozen and adds two thin matrices next to it, which multiply together into a matrix of the same size. During training only the thin matrices change, and when the model runs, their product is added to the original matrix's output, so the layer behaves like the original plus a small learned correction. How thin they are is a setting called the rank, often 8 or 16, and it decides how much the adapter can learn.
For GPT-3, with 175 billion weights, the LoRA paper reports 10,000 times fewer trainable weights than full fine-tuning and three times less GPU memory. A follow-up called QLoRA stores the frozen model in 4-bit numbers and trains LoRA adapters on top, which let its authors "finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance" (Dettmers et al., 2023). The adapter you get out is a small file, usually tens of megabytes, so you can keep one base model and swap adapters for different jobs.
Now you might be asking why anyone trains every weight, then. Well, because an adapter learns less. Biderman et al. compared LoRA with full fine-tuning on code and on maths, and found that "in the standard low-rank settings, LoRA substantially underperforms full finetuning", and also that "LoRA better maintains the base model's performance on tasks outside the target domain" (Biderman et al., 2024). Those were big training runs that were trying to teach a model a whole new domain. For a narrow job, like Relist's categories or a fixed output format, an adapter is usually enough, and keeping what the base model could already do is worth a lot.
Train every weight
Full fine-tuning lets every weight change. It can learn the most, and it costs the most, because during training the GPU has to hold the weights plus two more things for every weight, a gradient that says which way to nudge it and the optimizer's running statistics. Together that comes to several times the size of the model. It's also the easiest way to damage what the model could already do.
Kumar et al. measured that on image classifiers across ten datasets where the test images came from a different distribution than the training images. Full fine-tuning got "on average 2% higher accuracy ID but 7% lower accuracy OOD than linear probing", where ID means test images like the training ones and OOD means test images unlike them (Kumar et al., 2022). Their explanation is that while training is still learning the head from scratch, the backbone layers are changing at the same time, and that distorts features that were already good. The fix they tested is to train the head first, with the backbone frozen, and only then fine-tune everything, which they call LP-FT, and it did better than either method alone on those datasets.
Which one Relist picked
| Job | What trains | Why |
|---|---|---|
| Category classifier | The head first, then a LoRA adapter on SigLIP's image encoder | The head alone stopped at 88%, and an adapter changes the features while the base encoder stays untouched |
| Title writer | A LoRA adapter on a small open LLM | A narrow format and field job, where an adapter is usually enough |
| Similar-items embeddings | Covered in M42 | Contrastive training on co-click pairs has its own setup |
The classifier choice follows the LP-FT idea, so the adapter starts from a head that already works. Starting with the cheapest option that can change the features also means that if the adapter isn't enough, full fine-tuning is still there to try, with the adapter's numbers as the baseline to beat.
The training dataset
Most of the work goes into the dataset
A fine-tuned model learns whatever its dataset shows it, so the dataset decides most of the result before training even starts. The training settings in section 6 make much less difference than whether you have enough examples and whether their labels agree with each other. How you split the data counts just as much, because that's where a lot of teams fool themselves about how well the model works.
How many examples you need
It depends on how much you're changing, and the numbers people quote range widely.
- A head on frozen features can work with a few hundred labeled examples in total, or a few dozen per class, which is what M13 found for its classifiers.
- OpenAI's fine-tuning guide says "the minimum number of examples you can provide for fine-tuning is 10" and recommends "starting with 50 well-crafted demonstrations" for an LLM, as of October 2026 (OpenAI).
- LIMA fine-tuned a 65-billion-weight LLM on 1,000 carefully chosen examples and got a strong assistant out of it, which led its authors to conclude that "almost all knowledge in large language models is learned during pretraining" (Zhou et al., 2023).
- The BERT authors noticed that datasets with "100k+ labeled training examples" were "far less sensitive to hyperparameter choice than small data sets" (Devlin et al., 2018), so with more data, the training settings matter less.
What those add up to is that a few hundred good examples is a reasonable start for most narrow jobs. Then let the eval tell you where to add more. A cheap way to see whether more data would help is to train on a quarter of your dataset, then half, then all of it. If the test score is still climbing at the full dataset, more examples will probably help, and if it's flat, they probably won't.
Labels have to be consistent before they can be right
When two people label the same examples differently, the model gets taught both answers and learns neither. So before labeling thousands of examples, have two people label the same couple of hundred and count how often they agree. Relist did this with 300 photos and the two labelers agreed on 91% of them, and a quarter of the disagreements were nightstand against side table. Both labelers knew Relist's 70-centimetre rule. One of them applied it to a table with a shelf and no drawer, and the other didn't. The team wrote the rule out in a labeling guide with photos of the borderline examples, and agreement went up to 97%. That 97% is roughly the ceiling on what any model can score against those labels, since the test labels carry the same 3% of disagreement.
Split by listing, so the test is new to the model
You split your labeled data into three datasets. The training dataset is what the model learns from. The validation dataset is what you check during training, to pick settings and decide when to stop. The test dataset is what you measure the finished model on once, at the end, and it has to contain things the model has never seen in any form.
That last part is easy to get wrong. Relist's listings have about four photos each, and the first training run split the photos at random. Photos of the same nightstand ended up in training and in validation, so validation was partly testing whether the model recognised a nightstand it had already seen from another angle. Validation said 96%. On the next week's new uploads, the same model got 89%. Leakage is the name for any way information from the test side gets into training, and this kind, where near-copies sit on both sides of the split, is the most common.

The fix is to split by whatever groups the near-copies together. For Relist that was the listing, and then the seller, because one seller tends to photograph everything in the same room with the same light, and a model can learn to recognise the room. After splitting by seller, validation said 92.8% and new uploads came in at 92.5%, so the validation number could finally be trusted. If your data changes over time, it also helps to hold out the most recent few weeks as the test dataset, since that's closest to what the model will see after launch.
Write a Python script that splits labels.csv into train, validation and test datasets of 80%, 10% and 10%. Each row has photo_id, listing_id, seller_id, category and uploaded_at. Split by seller_id so that every photo from one seller lands in exactly one dataset. Keep each category's share roughly the same in all three. Print the number of photos and sellers per dataset, the categories with fewer than 5 test photos, and a check that no seller_id appears in two datasets.
Check what comes back. The script should fail loudly if a seller shows up in two datasets, and the category counts will tell you whether rare categories have enough test photos to measure at all.
Cover the rare categories and the confusable ones
Relist's most common category, dining chairs, had 3,100 labeled photos, and its rarest, piano benches, had 40. Trained on that as it is, a model learns that guessing dining chair is usually safe. You can fix that by sampling the rare categories more often during training or by weighting their mistakes more in the loss, and either way you should report accuracy per category, because the average hides the rare ones.
The examples that teach the most are the confusable ones, the side tables that look like nightstands and the other way round. M20 called these hard negatives, wrong answers that look almost right, and a dataset with plenty of them teaches the model the boundary you care about. A dataset made of easy examples teaches it things it already knew.
Labels from a bigger model need checking too
For the title writer, nobody at Relist was going to write 20,000 titles by hand. They had the large LLM fill in the title fields for 20,000 past uploads, which is the distillation setup from section 3, and those outputs became the small model's training dataset. The catch is that the small model learns the large model's mistakes along with everything else. A person reviewed 300 of those outputs at random and found at least one wrong field in about 6% of them, mostly brands guessed from the photo when the seller never named one. The team dropped every example where the brand didn't appear in the seller's text or in Relist's list of known brands, which removed most of the wrong brands, and they kept the reviewed 300 aside as part of the test dataset.
The training run
Watch the validation loss, and keep the checkpoint where it was lowest
A training run has a handful of settings, and the learning rate is the one that most often decides whether it works. It sets how big each nudge to the weights is. If it's too small the model barely moves, and if it's too large the weights overshoot and the loss jumps around or blows up. The batch size is how many examples the model sees before each nudge, and the number of epochs is how many times it goes through the whole dataset. Fine-tuning uses much smaller learning rates than pretraining, because the weights are already good and you only want to move them a little. BERT was pretrained with a learning rate of 1e-4, and its authors fine-tuned it with 5e-5, 3e-5 or 2e-5, for two to four epochs, with batches of 16 or 32, and found those ranges "work well across all tasks" (Devlin et al., 2018). Start from what the model's authors or the library's defaults suggest, because guessing these from scratch wastes runs.
The thing to watch while it trains is two curves. The training loss is the loss on the examples the model is learning from, and it almost always goes down, because the model is getting better at those exact examples. The validation loss is the loss on the validation dataset, which the model never trains on, and it's the one that tells you whether the model is getting better at the job. Early in training both go down together. At some point the validation loss stops falling and starts to rise while the training loss keeps going down, and that's overfitting, which means the model has started memorising its training examples, and that makes it worse on anything new.

So you save a copy of the weights, called a checkpoint, at the end of every epoch or every few hundred steps, and at the end you keep the one with the lowest validation loss. Stopping the run once the validation loss has risen for a while is called early stopping, and it saves the compute you'd spend on the epochs you were going to throw away. Relist's classifier adapter reached its lowest validation loss at epoch 3 of 8, so the team kept the epoch-3 checkpoint, and the next run was configured to stop two epochs after the validation loss stopped improving.
Most problems in a training run show up in those two curves before they show up anywhere else.
| What you see | What it usually means | What to try |
|---|---|---|
| Training loss barely moves | Learning rate too small, or the weights you meant to train are frozen | Check which weights are trainable, then raise the learning rate |
| Loss jumps around, or turns into NaN, meaning not a number | Learning rate too large | Lower it, and add a warmup that starts it small |
| Training loss falls, validation loss rises | Overfitting | Keep the best checkpoint, and use fewer epochs or more data |
| Training loss drops to nearly zero in the first few steps | The answer is leaking into the input, or the task is trivial | Look at a few inputs as the model sees them |
| Validation looks great and new data looks bad | Leakage between training and validation, or validation is unlike real traffic | Split by group, and test on recent data |
Run it twice before you trust a small difference
Training involves randomness, in the order the examples come in and in how new weights like the head start out, and that randomness is controlled by a number called the seed. Two runs with different seeds can land at noticeably different scores, and the smaller the dataset, the bigger the gap. Relist ran the classifier adapter with two seeds and got 92.6% and 93.1% on validation, so a change that moves the score by half a point could be nothing more than the seed. M8 covers how to tell a real improvement from noise, and the short version is to run the comparison more than once and only believe differences bigger than the spread between seeds.
Write down everything about the run
A month later, someone will ask how this model was made, and "I think it was the April data" isn't an answer. Training libraries and experiment trackers can log most of what you need, and the rest goes in a short text file saved next to the weights.
- The base model and its exact version
- A fingerprint of every data file, like a SHA-256 hash, and the split method
- Every training setting, and the seed
- The checkpoint you kept, with its validation and test scores
Write a training script that fine-tunes a LoRA adapter (rank 16) on the image encoder of google/siglip-base-patch16-224, starting from the linear head saved in head.pt, for the 240 categories in labels.csv. Use the train and validation datasets from splits/. Log training and validation loss every 100 steps, save a checkpoint at the end of each epoch, and stop when validation loss hasn't improved for two epochs. At the end, write run.json with the base model id, the SHA-256 of each data file, every setting, the seed, the best epoch and its validation accuracy per category.
When it comes back, read it before running anything. The backbone's original weights should be frozen, so only the adapter and the head train. The validation loss should be computed without updating any weights. And the per-category numbers should come from the validation dataset, never the training one.
Evaluating the result
Beat the best cheap fix, and check what got worse
The number a fine-tuned model has to beat is the best one you got without fine-tuning. For Relist's classifier that's the head on frozen features at 88.4%. Zero-shot's 71% doesn't count, because comparing against it would credit fine-tuning with a jump that the frozen head already gave you for a minute of training. Measure both on the same test dataset, the 2,000 hand-labeled photos from sellers that appear nowhere in training, and look at slices of it as well as the average, because the average hides the categories that fail.
| Slice | Test photos | Zero-shot | Head on frozen features | LoRA adapter |
|---|---|---|---|---|
| All photos | 2,000 | 71.0% | 88.4% | 93.1% |
| Nightstand and side table | 140 | 52% | 64% | 86% |
| Categories with under 100 training photos | 190 | 58% | 70% | 71% |
The adapter did what it was trained for. The nightstand and side table slice went from 64% to 86%, which is the boundary the frozen head couldn't draw. The rare categories barely moved, though, which tells the team where the next labeling effort should go, since no setting is going to teach the model a category it has seen 40 photos of.
Fine-tuning can break things the model did before
A model you fine-tune is often doing more than one job, and training only measures the one you're training for. Catastrophic forgetting is the name for a model losing abilities it had before fine-tuning, because the weights those abilities depended on moved.
That happened to Relist. Before the adapter, an engineer tried full fine-tuning on the SigLIP image encoder and got a classifier that scored 93.6%, a bit better than the adapter would later get, and shipped it. The similar-items search used the same encoder's embeddings, and nobody had listed it as something the change could affect. Fine-tuning for categories pulled the embeddings toward "which category is this", which meant style and price signals got weaker. The team measured search with recall at 10, which means taking pairs of listings the same buyer clicked in one session and checking how often the second listing shows up in the first one's top 10 similar items. It dropped from 0.41 to 0.29, and buyers' clicks on the row fell with it.

The fix had two parts. The adapter approach let search keep using the base encoder while only the classifier loaded the adapter, so search went back to 0.41. And the team added search's recall at 10 to the eval that runs before any model change ships. The general version of that is to list every job the weights do before you change them, and keep an eval for each one. That's the regression check from M1, running the same test cases again after a change to find anything that used to pass and now fails.
For an LLM, the jobs you didn't train for include following instructions and refusing harmful requests. Qi et al. found that fine-tuning GPT-3.5 Turbo on just 10 adversarially designed examples, for under $0.20, broke its safety guardrails, and that fine-tuning on ordinary, harmless datasets also weakened safety (Qi et al., 2023). Relist's title model only fills in fields and never talks to anyone, so a general chat eval doesn't apply to it. If your fine-tuned LLM talks to users, though, its regression checks need a safety eval too, with requests the model should refuse, run before and after fine-tuning.
The distilled title writer
The title writer is measured the same way, against the best cheap fix on the same test dataset, which for titles includes the 300 reviewed examples from section 5.
| Model | All fields right | Cost per 1,000 uploads | Time per upload |
|---|---|---|---|
| Large hosted LLM, structured output | 94% | $3.90 | 4.1 s |
| Small open LLM, structured output, prompted | 79% | $0.20 | 0.6 s |
| Small open LLM with a LoRA adapter, distilled | 91% | $0.20 | 0.6 s |
The distilled model got within three points of the large one at about a twentieth of the cost per call. Whether three points is acceptable is a product question, and Relist answered it by sending the uploads where the small model's brand isn't on the known-brands list to the large model, which is about one in ten, so the hard ones still get the better model. M14 covers that kind of routing.
Hosted or open weights
Decide who trains it, and who keeps the weights
You can have a provider fine-tune a model for you or train one yourself, and the difference is mostly about who ends up holding the weights. With hosted fine-tuning, you upload your examples to a provider, usually as a JSONL file with one example per line, and the provider trains the model on its own hardware and serves it to you behind the same API as its other models. You never see the weights. With open weights, you download a model whose weights are published, like Qwen or Gemma, train an adapter on it yourself on a GPU you rent, and keep the resulting file.
Hosted is the easier start, since you don't need a GPU of your own and the fine-tuned model is one more model name in your code. The cost is that the model only exists for as long as the provider keeps it running. OpenAI's own guide says, as of October 2026, that "OpenAI is winding down the fine-tuning platform", that "the platform is no longer accessible to new users", and that "all fine-tuned models will remain available for inference until their base models are deprecated" (OpenAI). So a team that fine-tuned on OpenAI keeps its model until the base model is retired, and then has nothing to move to, because the weights were never theirs. Other providers still offer hosted tuning, like Google's supervised tuning for Gemini models on Vertex AI (Google Cloud) and Amazon Bedrock, which added reinforcement fine-tuning for open-weight models like Qwen and OpenAI's gpt-oss in February 2026 (AWS). The same question applies to all of them, which is what happens to your model when the base model goes away.
| Hosted fine-tuning | Open weights, trained by you | |
|---|---|---|
| Who trains it | The provider, from your uploaded examples | You, usually a LoRA adapter on a rented GPU |
| Who holds the weights | The provider | You |
| How you serve it | The provider's API, at its prices | Your own GPU, or a service that hosts open models |
| What happens when the base model is retired | The fine-tuned model goes with it, and you retrain on whatever replaces it | Nothing, until you choose to move |
| What you're responsible for | Your dataset and your evals | Everything on the left, plus training and serving |
Relist picked open weights for both of its fine-tuned models. The title writer is a LoRA adapter on a small open LLM, trained on the 18,000 examples left after filtering, in a few hours on one rented GPU. The classifier adapter is a few megabytes on top of SigLIP. Both run on a GPU server Relist rents by the month, which costs a fraction of the $7,000 a month the large model's titles had cost. These numbers are made up, like Relist's other numbers. M40 covers how to serve models like these.
Plan for retraining from the first run
A fine-tuned model is a snapshot of your data on the day you trained it, and the world keeps moving. Relist adds categories a few times a year, new brands show up every month, and the base models themselves get new versions. So keep the dataset and both scripts, the split one and the training one, in version control, next to the run.json from section 6, so that retraining is a rerun with new data and not a project.
Then decide what triggers a retrain before you need one. Relist labels 200 new uploads a week and runs the classifier on them, and a retrain is due when accuracy on those drops two points below the test score, or when a new category ships. Watching a live model's numbers over time is M12's subject, so for now make sure the retraining path exists and has been run at least once before anyone depends on it.
Hands-on
Fine-tune one encoder, and prove it was worth it
Pick an image classification job where a general model gets one boundary wrong. Good ones are your own photos sorted into categories you care about, or a public dataset with two classes that look alike, like two dog breeds or two kinds of leaf disease. A free notebook GPU is enough for everything below, and the point is to come out with numbers you can defend in an interview.
Build the datasets first
Label at least 30 images per class, with at least four classes, two of which are easy to confuse. Write a one-page labeling guide, have a second person label 100 of the images, and record how often you agree. Split into training, validation and test by whatever groups your near-copies, like the photographer or the day, and check that no group appears in two datasets.
Climb the ladder, and record every rung
Measure zero-shot SigLIP with plain class names, then with written-out descriptions. Then train a linear head on frozen SigLIP embeddings. Then train a LoRA adapter on the image encoder, starting from that head, with checkpoints every epoch and early stopping. Record the test accuracy of every rung, overall and on the confusable pair, plus the time and cost of each.
Look for what got worse
Pick a second job the base encoder does, like image-to-image search over your images, and measure it with recall at 10 before and after. Then fully fine-tune the encoder once, without an adapter, and measure both jobs again. Write down what changed and by how much.
What to hand in
- The labeling guide, and the agreement number between the two labelers.
- The split script's output, showing that no group appears in two datasets.
- A table of every rung on the test dataset, overall and on the confusable pair, with time and cost.
- The training and validation loss curves for the adapter run, with the checkpoint you kept marked.
- The adapter run repeated with a second seed, and the spread between the two.
- The second job's recall at 10 for the base encoder, the adapter and full fine-tuning.
- The
run.jsonfor the run you'd ship. - A paragraph saying which rung you'd ship and why, and what new evidence would change your mind.
Putting it together
Putting it together
Fine-tuning keeps training a model someone else pretrained, on your own examples, so the weights move toward the behavior you showed it. It's good at behavior, like labels only you use or a fixed output format, and it can teach a small model to copy what a bigger model does at a fraction of the price. It's bad at facts, which it learns slowly and which make it hallucinate more once it does. That's why a model fine-tuned on contracts can end up worse than the model it started from, and why facts go in retrieval.
Most of the work happens before training. Relist fixed the worst problem in two of its three jobs with a category filter and structured output, and got from 71% to 88% on the third with a head that trained in a minute. Fine-tuning was for what was left, and the dataset decided most of how that went. Its labels had to agree with each other before they could be right, and its split had to keep near-copies of the same item on one side. The labels that came from the large model had to be checked as well, because the small model learns the large model's mistakes too.
During training, the validation loss says when to stop, and the checkpoint where it was lowest is the one to keep. After training, the model has to beat the best cheap fix on a test dataset it has never seen, slice by slice. It also has to leave every other job the weights do no worse than before, which is what the shared encoder taught Relist. Adapters make that easier, because the base model stays as it was and the adapter only loads where it's wanted. And whoever holds the weights decides how long the model lives, which is a good reason to hold the weights yourself, next to the data and scripts that made them.
| Job | Cheap fix first | What was fine-tuned | Result on the test dataset |
|---|---|---|---|
| Category classifier | Descriptions, then a head on frozen SigLIP features, 88.4% | A LoRA adapter on SigLIP's image encoder, from the head | 93.1%, with nightstand and side table up from 64% to 86% |
| Similar-items row | A category filter, clicks on the row from 2.1% to 3.0% | Nothing yet, contrastive training on co-click pairs comes next | Recall at 10 kept at 0.41 by leaving the base encoder alone |
| Title writer | Structured output, format right on every listing | A LoRA adapter on a small open LLM, distilled from the large one | 91% of listings with every field right, at about a twentieth of the cost |
M41 covers fine-tuning LLMs in depth, including preference tuning. M42 covers embedding models and the co-click training Relist's search needs, and M43 covers vision models, from classifiers to detectors. M8 covers how to tell a real improvement from noise, and M40 covers serving the models you trained.
Checkpoint · recall · 5 questions
What the module said
- 01
What does fine-tuning change?
- 02
Why is fine-tuning a poor way to give an LLM your product catalogue?
- 03
What does LoRA train?
- 04
During training, the training loss keeps falling and the validation loss starts rising. What is happening?
- 05
Which number should a fine-tuned classifier be compared against?
0 / 5 answered
Checkpoint · understanding · 5 questions
Reason it through
- 01
Relist's linear head on frozen SigLIP features stopped at 88%, and most remaining errors were nightstand against side table. Why couldn't a better head fix that?
- 02
A team splits 20,000 photos at random and gets 96% on validation, but new uploads score 89%. What most likely explains the gap?
- 03
Why did Kumar et al. find LP-FT did better out of distribution than full fine-tuning from a random head?
- 04
Relist's title writer was distilled from a large LLM. Why did the team filter the large model's outputs before training?
- 05
Why does it matter who holds the weights of a fine-tuned model?
0 / 5 answered
Checkpoint · debugging · 4 questions
Debug it
- 01
After shipping a fully fine-tuned SigLIP encoder for categories, Relist's dashboard shows clicks on the "you might also like" row falling. The classifier's test accuracy went up. What's going on?
- 02
A training run's loss turns into NaN after a few hundred steps. The data loads fine and the model runs on a single example. What's the most likely cause?
- 03
A team's fine-tuned classifier beats the frozen head by 0.4 points on validation. They ran each once. Should they ship the fine-tuned one?
- 04
Two weeks after launch, a new category, "bar cart", ships. Uploads of bar carts are all classified as side tables. The model's test accuracy hasn't changed. What's wrong?
0 / 4 answered
That's the last one written so far
Pick your next module from the board.
