MLGuerrillaStart with M1 →
Free · in beta·intermediate·M41·40 min read·Prereq: Fine-Tuning Fundamentals (M21), which covers the training loop, adapters, dataset splits and reading a training run. The Model as a Component (M4) covers tokens and calling an LLM.

Fine-Tuning LLMs

The capability

What fine-tuning an LLM is

As you already know from M4, an LLM writes text one token at a time, and each token is its best guess at what comes next given everything before it. Pretraining taught it to guess the next token on a huge pile of internet text, and then the lab trained it further on conversations so that it answers questions and follows instructions. That second stage is fine-tuning too, and it's the same thing you do when you fine-tune an LLM yourself, just with your own examples and far fewer of them.

M21 covered what's the same for every kind of model, meaning the training loop, adapters, how to split your data and how to read a training run. What makes LLMs different is the kind of training signal you can give them.

The first is supervised fine-tuning, usually shortened to SFT. Each training example is a prompt and the response you wanted, and the model is trained to make that response more likely, token by token. It's the most common kind by far, and it's what you use when you can show the model what a good answer looks like.

The second is preference tuning. Each example is a prompt with two responses, one marked as better than the other, and the model is trained to make the better one more likely than the worse one. You use it when people can't easily write the perfect answer but can reliably say which of two answers they prefer, which turns out to be most questions of style and tone.

The third is training against a grader, often called reinforcement fine-tuning. The model writes its own responses, a piece of code or another model scores each one, and training pushes the model toward the responses that scored well. You use it when correctness can be checked automatically, like whether generated SQL runs and returns the right rows.

Three kinds of training example for an LLM, side by side. Supervised fine-tuning shows one prompt, describe this nightstand, paired with one written response, and the model is trained to make that response more likely. Preference tuning shows the same prompt with two responses, a short specific one marked better and a generic one marked worse, and the model is trained to prefer the better one. Training against a grader shows the model writing four responses of its own, a grader scoring them pass or fail, and training pushing toward the ones that passed, which is highlighted.
Most projects start with supervised fine-tuning, and many never need anything else.

All three change how the model behaves, and none of them reliably teach it new facts, for the reasons M21 gave. They're also all ways of steering a model toward one job, so a model fine-tuned hard on product descriptions can get worse at things you didn't train it on, which M21 called forgetting.

Some places it shows up

  • Cutting cost by distillation, where a large model's answers become the training data for a small one that then does the same job for a fraction of the price. Relist's title writer in M21 was one.
  • A house voice or format that a prompt only gets right most of the time, like support replies in a company's tone or reports with a fixed structure.
  • Tool calls for your own API, where a small model learns which of your functions to call and with which arguments, which M18 covers from the prompting side.
  • Domain language, like clinical notes or legal clauses, where the words and the structure are unusual, as long as what the model needs to know still comes from retrieval.

The running example

Relist's description writer said the right things in the wrong way

Relist is the secondhand furniture marketplace from M21, where a pipeline turns each seller's upload into a listing. It's invented, and so are its numbers. By the end of M21 it had a fine-tuned image classifier for categories and a small open LLM with a LoRA adapter that wrote titles, distilled from a large hosted model.

The next feature was descriptions. Sellers write a line or two, often less, and buyers scroll past listings that don't say what the item is made of or how big it is. So Relist had the large model write a description for every upload from each upload's photos and text, and in a two-week test, listings with those descriptions got 11% more saves. The problem was cost again. A description is about 120 words, and at 60,000 uploads a day the large model's descriptions alone would have cost about $20,000 a month. So the team decided to distill the large model into the same small LLM that already wrote titles, as a second adapter.

The first version had two bugs that had nothing to do with the data. Its training script formatted every example as plain text, with a line saying "Seller text:" followed by the text and a line saying "Description:" followed by the answer, while the server sent requests through the model's chat format, so the model never saw at serving time the layout it had learned. And none of the training examples ended with the token that marks the end of a reply, so the model never learned to stop. Its descriptions ran until they hit the 400-token limit, repeating the last sentence over and over.

Once those were fixed, the descriptions were accurate, with the facts right on 96% of a 500-upload test dataset, and they followed the house rules. But when Relist's merchandisers, the people who decide how listings look, compared them blind against the large model's, they preferred the large model's 64% of the time. The facts in the small model's descriptions were right. What the merchandisers didn't like was that they ran long and leaned on phrases like "a perfect addition to any home", and that they opened with the colour when buyers want the size first. None of that was easy to fix by writing better training examples, because the large model's own outputs, which were the training examples, had some of the same habits.

The first two problems came from how an LLM training example has to be formatted, which section 3 covers. The third is what preference tuning is for, and section 5 covers it.

The training example

Format every example the way the model will see it, and train only on the answer

An LLM never sees your conversation as a list of messages. Before the model reads anything, the conversation gets flattened into one sequence of tokens, with special tokens that mark where each message starts and ends and who wrote it. The rules for that flattening are called the chat template, and every chat model ships with its own. Hugging Face's documentation shows two models fine-tuned from the same base model, Mistral-7B, where one wraps the user's message in [INST] and [/INST] and the other uses <|user|> and <|assistant|>, and it warns that "with the wrong control tokens, these models would have drastically worse performance" (Hugging Face).

That's what went wrong with Relist's first run. The training examples were plain text with "Seller text:" and "Description:" labels, so the adapter learned to write descriptions after those labels. The server sent each request through the model's chat template, with the special tokens around a user message, which the adapter had never been trained on. So the model behaved much like the model it started from, and the training barely showed in its descriptions. The fix is to build training examples with the same chat template the server uses, which in practice means storing each example as a list of messages and letting the tokenizer's apply_chat_template do the formatting, both in training and at serving time.

The loss should count the answer and nothing else

Once an example is a token sequence, the model is trained the same way it was pretrained, to predict each token from the ones before it. The question is which tokens count toward the loss. The system message and the seller's text are in every example, and the model doesn't need to learn to write them, since at serving time they're given to it. Only the reply is what you want it to learn to produce. So the usual setup is to mask the prompt tokens, which means they're still in the input so the model can read them, but they're left out of the loss, so the training only pushes on the tokens of the answer.

A Relist training example shown as a row of tokens after the chat template. The system message tokens and the user message tokens, holding the seller's text, are grey and labeled read but not trained on. The assistant reply tokens, a description beginning Solid pine nightstand, 45 cm wide, are highlighted and labeled counted in the loss. The last token, the end-of-turn token, is highlighted more strongly and labeled this is how the model learns to stop.
The template's special tokens differ between model families. The split between read and trained on is the same everywhere.

Training libraries handle this if you give them data in the right shape. In Hugging Face's TRL library, for example, a dataset with separate prompt and completion fields computes the loss "on the completion tokens only, ignoring the prompt tokens" by default, and for a dataset of whole conversations a setting called assistant_only_loss restricts the loss to the assistant's messages (TRL documentation). Check it anyway, because when the masking is off, a lot of every training step goes into learning to reproduce prompts that are the same in every example, and short answers after long prompts barely get trained at all.

The end-of-turn token is part of the answer

Every reply in a chat template ends with a special token that means "this message is over", and when the model generates that token at serving time, the server stops. The model only learns to produce it if the training examples contain it, inside the part the loss counts. Relist's first examples didn't, so the model had learned to write descriptions and had never once been shown where one ends, and it kept going until it hit the 400-token cap.

The opposite mistake happens too. Some tokenizers add their own start and end tokens when they turn text into tokens, so if you format with the chat template and then tokenize the result as ordinary text, those tokens can end up in the sequence twice. Hugging Face's documentation warns that adding extra special tokens on top of the template "is often incorrect or duplicated, hurting model performance". The way to catch all of this is to look at the tokens.

Ask your AI coding tool

Load the tokenizer for our base model and the first three examples from train.jsonl, where each example is a messages list. Format each one with apply_chat_template exactly the way our training script does, then print it two ways. First as text, with the special tokens visible. Then as a table of token id, token text and whether the token counts in the loss, using the same masking the training script uses. Flag any example where the last counted token isn't the end-of-turn token, and any special token that appears twice in a row.

Read that output before every first training run on a new model. It takes five minutes, and it would have caught both of Relist's bugs before a single GPU hour was spent.

Supervised fine-tuning data

Build the training dataset from the outputs you'd want to ship

For supervised fine-tuning, every training example is a demonstration, meaning a prompt and the exact response you'd want the model to give. The model learns to imitate those responses, habits and all, so the dataset is a description of the model you'll get. Where the responses come from decides most of what that model is like.

Where SFT responses usually come from
SourceWhat you getWhat to watch for
People write themThe behavior you want, as precisely as you can describe itSlow and expensive, and different writers write differently
Your own logs, after people corrected themReal inputs, with the fixes your team already madeOnly the corrected ones are good examples, so keep the corrected version
A larger model, which is distillationThousands of examples in an afternoonThe small model learns the large model's mistakes and habits too

Distillation is the most common source now, and the reason is in the numbers. Hsieh et al. trained a 770-million-weight T5 model on a larger model's outputs, including its written-out reasoning, and found that it "outperforms the few-shot prompted 540B PaLM model using only 80% of available data on a benchmark" (Hsieh et al., 2023). That was one benchmark in 2023, so read it as a sign of what's possible. It's also the idea behind Relist's title writer in M21, where a distilled small model got within three points of the large one on a single narrow job.

Relist built its description dataset that way. The large model wrote descriptions for 9,000 past uploads, using the same prompt that had won the two-week test. Then the team filtered them with checks that run in code. They dropped any description over 160 words or with a repeated sentence, and any that contained a number not found in the seller's text or the title fields, which left 8,200. A merchandiser rewrote 300 of the remaining ones by hand to fix tone, and those went in as well. The dataset was split by seller, for the reasons M21 gave, and 500 uploads from sellers outside training became the test dataset.

What makes a good SFT dataset

  • The prompt in every example matches production exactly, including the system message and the order of the fields, so the model learns from the inputs it will get.
  • The inputs cover the range you'll see, like uploads with no seller text at all or categories with few listings, as well as the typical upload.
  • The responses are ones you'd ship. A response you'd have to fix before showing a buyer teaches the model to produce exactly that.
  • Duplicates are removed, since a thousand near-identical examples mostly teach the model that one example.

How many examples is the same question M21 covered, and for a narrow job like Relist's, a few thousand good ones is plenty. LoRA on a small open LLM is the usual starting point, with the library's suggested learning rate and two or three epochs, and the validation loss curve from M21 tells you when to stop. With LLMs, though, the validation loss is less useful than it was for the classifier, because a lower loss means the model predicts the training-style text better, which isn't the same as writing descriptions people like. So alongside the loss, generate answers for a couple of dozen validation prompts at each checkpoint and read them.

Preference tuning

Use preference pairs when people can judge an answer better than they can write one

Relist's merchandisers couldn't have written 8,000 perfect descriptions, but shown two descriptions of the same nightstand, they could say which one was better in a few seconds, and they mostly agreed with each other. That's the situation preference tuning is for. Each training example is a prompt with a chosen response and a rejected response, and training makes the chosen one more likely relative to the rejected one.

This is how the chat models you use were made in the first place. OpenAI's InstructGPT work fine-tuned GPT-3 on written demonstrations and then on people's rankings of the model's outputs, and in their tests "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters" (Ouyang et al., 2022). Their method, reinforcement learning from human feedback or RLHF, trained a separate reward model on the rankings and then used reinforcement learning to push the LLM toward outputs that reward model scored highly, which is a lot of machinery to get right.

DPO trains on the pairs directly

Direct preference optimization, or DPO, gets to a similar place without the reward model. Rafailov et al. showed that the same goal can be reached "with only a simple classification loss" over the pairs, and that it's "stable, performant, and computationally lightweight" compared with the RLHF pipeline (Rafailov et al., 2023). That's why DPO, and the family of similar methods that followed it, is what most teams use when they do preference tuning themselves.

DPO starts from two copies of your SFT model. One is frozen and serves as the reference, and the other is the one being trained. For the chosen description and the rejected one, it asks both copies how likely they think that description is. The loss is small when the model being trained has raised the chosen description's likelihood more than the reference did, and lowered the rejected one's, so training pushes the gap in that direction. A setting called beta controls how far the trained model can move away from the reference. In TRL it defaults to 0.1, and TRL's documentation says "higher β means less deviation from the reference model".

One of Relist's preference pairs. The prompt is a seller's upload of a pine nightstand. The chosen description, written by a merchandiser editing the model's draft, opens with Solid pine nightstand, 45 cm wide, with one drawer, and is 62 words long. The rejected description, the model's original draft, opens with This charming white nightstand is the perfect addition to any bedroom, and is 121 words long. An arrow labeled raise likelihood points at the chosen one, which is highlighted, and an arrow labeled lower likelihood points at the rejected one. Below both, a frozen reference model is drawn as an anchor, with beta labeled as how far the trained model may move from it.
The pair came free, because the merchandisers were already editing a daily sample of descriptions.

Where the pairs come from

Relist didn't run a labeling project for its pairs. The merchandisers already edited a sample of each day's descriptions before they went live, so every edited description gave a pair, with the merchandiser's version as chosen and the model's draft as rejected. Three weeks of that gave 3,400 pairs. Other common sources are an A/B test where buyers' behavior picks the winner, or showing a reviewer two drafts from the model at different settings and recording which one they pick.

Watch the length of what wins

Preference data has a well-known bias, which is that the longer answer tends to win. Singhal et al. studied this in RLHF and found the improvements "to largely be driven by increasing response length", and that "even a purely length-based reward reproduces most downstream RLHF improvements over supervised fine-tuned models" (Singhal et al., 2023). So before training, compare the average length of the chosen responses with the rejected ones. If chosen is much longer, the model may learn to write more, whether or not that's what people liked.

Relist's pairs leaned the other way, because the merchandisers' edits mostly cut words, and that was the point. After DPO on the 3,400 pairs with beta at 0.1, the average description went from 118 words to 86. In a fresh blind comparison against the large model, merchandisers preferred the small model's description 49% of the time, up from 36%, and the facts were still right on 96% of the test uploads, which mattered, because a model that learns to write shorter can also learn to drop the dimensions.

Reinforcement fine-tuning

Train against a grader when code can check the answer

Relist's description writer had one more problem that neither kind of training fixed. In about 3% of descriptions it stated a measurement the seller never gave, like "60 cm tall" for a nightstand whose seller only mentioned the width. The large model did it too, which is how it got into the training data before the filter caught most of it, and preference pairs didn't help, because the merchandisers' daily sample was too small to contain many.

This kind of error has a useful property, which is that code can find it. Every number in a description can be checked against the seller's text and the title fields, and a description with a number that isn't in either one fails. With a check like that, you can train a third way. The model writes several descriptions for the same upload and the check scores each one. Then training raises the likelihood of the ones that passed and lowers the ones that failed. That's reinforcement learning with a grader, and when the grader is code, it's often called reinforcement learning with verifiable rewards.

It's the method behind the recent jump in reasoning models. DeepSeek trained R1's reasoning mostly this way, rewarding answers to maths and coding problems that could be checked automatically, and reported that reasoning behaviors like checking its own work came out of that training without people writing examples of them (DeepSeek-AI, 2025). Hosted platforms offered it too, and OpenAI's reinforcement fine-tuning ran on its reasoning models before the fine-tuning platform started winding down, as M21 covered.

Which kind of training fits which problem
You haveUseRelist's example
Examples of exactly the right outputSupervised fine-tuning8,200 filtered descriptions from the large model
People who can say which of two outputs is betterPreference tuning, like DPO3,400 merchandiser edits as chosen and rejected pairs
Code that can check whether an output is rightTraining against a graderEvery number in a description must appear in the seller's text or the title fields

It's also the most expensive and the easiest to get wrong. The model generates many responses per prompt during training, which makes each step far slower than SFT. And the model will find any way to score well that the grader allows, so a check that only looks for invented numbers would reward a description with no numbers at all. A real grader for Relist would have to require the seller's dimensions to appear as well as forbid new ones.

Relist looked at that and decided not to train against the grader. It ran the same check in production, and when a description failed it, the pipeline asked the model for another one, which fixed nearly all of the 3% for one extra call on 3% of uploads. That's the usual outcome, and it's worth checking every time. If the grader can run at serving time and failures are rare, regenerating is far cheaper than training, and you keep the grader as a guard either way. Training against it makes sense when failures are common enough that regenerating would cost more, or when a retry is expensive.

Two ways to use the same grader, side by side. On the left, train against it, the model writes four descriptions per upload during training, the grader marks each pass or fail, and training raises the likelihood of the passes, labeled slow training and the grader must close every loophole. On the right, guard in production, which is highlighted, the model writes one description, the grader checks it, a pass is published and a fail triggers one regeneration, labeled one extra call on about 3% of uploads. Both panels use the same grader, every number in the description must appear in the seller's text or the title fields.
The grader is worth writing either way. The question is only whether it also becomes training data.

Evaluating a fine-tuned LLM

Judge the outputs, and check what else changed

A fine-tuned classifier has one number, accuracy, and you compare it with the baseline. A fine-tuned LLM writes text, so its evaluation needs a few kinds of check, and the rule from M21 still applies, which is to compare against the best thing you had without fine-tuning, on the same test dataset.

Checks that run in code

Start with everything a program can check, because those checks are cheap and exact. For Relist's descriptions that meant the facts check above, a length limit, a check that the reply ended on its own before the token cap, and a check that no sentence repeats. The two bugs from section 3 would both have shown up here, as replies that hit the cap.

A judge for what code can't check

Tone and helpfulness need a judge. The cheapest one is another LLM, asked to compare two descriptions of the same upload and pick the better one against written criteria, which M15 covers in depth. Two details make it trustworthy enough to use. Show each pair twice with the order swapped, and count a win only when both orders agree, because LLM judges are known to favour answers by their position, as Zheng et al. found when they measured judge biases (Zheng et al., 2023). And check the judge against people before trusting it. Relist had merchandisers label 200 pairs, and the LLM judge agreed with them on 85% of those, which was good enough to run the judge on every new checkpoint and keep the merchandisers for a final check.

Relist's description writer on the 500-upload test dataset
ModelFacts rightAverage lengthPreferred over the large modelCost per 1,000 uploads
Large hosted model, prompted97%104 wordsbaseline$11.00
Small model, prompted88%131 words21%$0.40
Small model with SFT adapter96%118 words36%$0.40
Small model with SFT, then DPO96%86 words49%$0.40

What else the adapter changed

The description adapter is only loaded for descriptions, so Relist's other jobs on the same base model weren't affected, which is one of the reasons M21 preferred adapters. If your fine-tuned LLM does more than one job, or talks to users directly, run a general eval before and after as well, with instruction-following and safety prompts. Qi et al. showed that even harmless fine-tuning data can weaken a model's refusals, as M21 covered, and a model tuned toward one kind of output can get worse at other things it used to do, which is the forgetting M21 measured on Relist's image encoder.

Serving

Serve one base model with an adapter per job

Relist now had two LoRA adapters on the same small open LLM, one for titles and one for descriptions. There are two ways to serve that. You can merge each adapter into a copy of the base model's weights, which gives you two ordinary models, each the full size of the base model, that run exactly like any other model. Or you can keep one copy of the base model in GPU memory and load the adapters next to it, picking the adapter per request.

The second way is what most serving engines support now, and it's why adapters are cheap to run as well as to train. vLLM, a widely used open-source serving engine, takes the base model and a list of named adapters when it starts, and each request names the adapter it wants in the model field, the same field you'd use to pick a model on a hosted API. Research systems have pushed this a long way, and S-LoRA was built to serve "thousands of concurrent LoRA adapters" from one base model (Sheng et al., 2023).

One GPU server holding a single copy of Relist's small base model, with two small adapter files beside it, one for titles and one for descriptions, which is highlighted. Two kinds of request come in. A title request names the titles adapter and gets the base model plus the titles adapter. A description request names the descriptions adapter and gets the base model plus the descriptions adapter. A note says both adapters share the base model's memory, and each adds a few tens of megabytes.
Merging an adapter into the weights is the other option, and it costs a full copy of the model per job.
Start vLLM with the base model and both adapters
vllm serve relist/base-llm \
  --enable-lora --max-lora-rank 16 \
  --lora-modules titles=adapters/titles-v3 descriptions=adapters/descriptions-v2

There are two things to get right when you deploy. The server has to apply the same chat template the training script used, which is the problem from section 3 again, so the safest setup is to load the template from the same tokenizer files in both places. And every adapter has a version, like descriptions-v2, recorded with the base model it was trained on, since an adapter only works on the exact base model it was trained against. When the base model changes, every adapter on it has to be retrained, which is the retraining plan from M21.

Hands-on

Fine-tune a small LLM to copy a big one, and then make it better

Pick a narrow text job where a large hosted model does well and a small open model doesn't, like writing product descriptions from a few fields, turning support tickets into a fixed summary format or extracting fields from emails into JSON. A free notebook GPU is enough for a model of one to three billion weights with LoRA.

Build the datasets

Write or collect 600 or more realistic inputs, split by source the way M21 described, and keep 100 aside as the test dataset. Have the large model answer the rest with your best prompt. Filter its answers with checks in code, and fix 30 by hand to see what the filter missed.

Run SFT, and look at the tokens first

Format every example with the small model's chat template, print three of them with the masking shown, and confirm the end-of-turn token is counted. Train a LoRA adapter, keep the checkpoint with the lowest validation loss, and generate answers for 20 validation prompts at every checkpoint so you can read them.

Add preference pairs

Make 300 or more pairs, either by editing the SFT model's answers yourself or by picking the better of two answers sampled at a higher temperature. Compare the average length of chosen and rejected before training, then run DPO from the SFT adapter.

What to hand in

  • A screenshot or printout of one formatted training example, with its masked and counted tokens marked.
  • A table like Relist's, comparing the large model, the small model prompted, SFT and SFT plus DPO on your test dataset, with your code checks, average length, judge preference against the large model and cost per 1,000 calls.
  • How often your LLM judge agreed with you on 50 pairs you labeled yourself, with the order swapped for every pair.
  • One problem you found that a grader in code could catch, and whether you'd regenerate on failure or train against it, with the numbers behind the choice.
  • The command that serves both your base model and the adapter, and a request that uses the adapter by name.

Putting it together

Putting it together

Fine-tuning an LLM uses the same loop as any other model, with three kinds of training signal on top. Supervised fine-tuning learns from examples of the right answer, and preference tuning learns from pairs where one answer beats another. Training against a grader learns from scores that code gives the model's own answers. Most projects need only supervised fine-tuning and some add preference tuning, while a grader is for problems code can check, and only when regenerating on failure costs more than training.

The bugs that hurt first are usually in the format. Every example has to go through the same chat template the server uses, the loss should count the answer and nothing else, and the answer has to end with the end-of-turn token, or the model never learns to stop. Five minutes spent printing the tokens of a few examples catches all of it.

The data decides the rest. A small model trained on a large model's filtered answers can get close to it on one narrow job, and it learns that model's habits as well as its skills. Pairs from people's edits can then fix the habits, as long as you check what the chosen answers have in common, since longer answers tend to win when nobody's checking. The result is judged against the large model on the same test dataset, with code checks first and an LLM judge that's been checked against people, and it's served as an adapter beside the base model, with the same template the training used.

Relist's description writer, finished
PartWhat it is
Base modelThe same small open LLM that writes titles
Training data8,200 filtered large-model descriptions and 300 hand-fixed ones, split by seller
TrainingA LoRA adapter with SFT, then DPO on 3,400 merchandiser edits, beta 0.1
Guard in productionEvery number must appear in the seller's text or the title fields, with one regeneration on failure
EvaluationCode checks, plus an LLM judge that agrees with merchandisers on 85% of pairs, on 500 test uploads
ServingOne base model and two adapters in vLLM, picked by name per request
ResultPreferred over the large model 49% of the time, at $0.40 per 1,000 uploads against $11.00

M21 covers what's the same for every model, including dataset splits and forgetting. M42 covers fine-tuning embedding models, which is what Relist's similar-items search needs next, and M43 covers vision models. M15 covers LLM judges in depth, and M40 covers serving models like these.

Checkpoint · recall · 5 questions

What the module said

  1. 01

    In supervised fine-tuning of an LLM, which tokens should usually count toward the loss?

  2. 02

    What does a DPO training example contain?

  3. 03

    What does a higher beta do in DPO?

  4. 04

    What did Singhal et al. find about RLHF improvements?

  5. 05

    Why can one GPU serve both of Relist's adapters cheaply?

0 / 5 answered

Checkpoint · understanding · 5 questions

Reason it through

  1. 01

    Relist's first adapter barely changed the model's output at serving time. Its training examples used "Seller text:" and "Description:" labels. Why?

  2. 02

    Why did Relist's merchandisers' edits make good preference pairs?

  3. 03

    Relist could check every number in a description against the seller's text. Why didn't it train against that grader?

  4. 04

    A grader for invented numbers only fails descriptions with numbers not in the seller's text. What would training against it alone risk?

  5. 05

    Why does Relist count an LLM judge's win only when both orders agree?

0 / 5 answered

Checkpoint · debugging · 4 questions

Debug it

  1. 01

    A fine-tuned model writes a good description, then keeps going, repeating its last sentence until it hits the token limit. Training loss looked normal. What's going on?

  2. 02

    After DPO, a team's model wins far more judge comparisons, and its answers are twice as long as before. People reading them aren't impressed. What should they check first?

  3. 03

    Every adapter on Relist's server starts giving worse descriptions on the same day, with no new training. What most likely changed?

  4. 04

    A team's SFT model scores lower validation loss at every checkpoint, but its outputs at later checkpoints read worse. What should they do?

0 / 4 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.