MLGuerrillaStart with M1 →
Text and language·18 min read·Updated 23 September 2026

BERT and encoder models

These are small models that read text in both directions, which you can fine-tune on your own labels. They classify your text for a fraction of what an LLM costs, but they can't write a single word back to you.

BERT came out of Google in October 2018, and it set the pattern that encoder models still use today.

The pattern has two steps. The first one is pretraining. You take a model, show it billions of words of ordinary text, and mask some of the words, which means you hide them from the model and ask it to guess what they were.

That step takes a lot of compute, so almost nobody runs it themselves. For every model on this page, a lab has already done the pretraining and published the weights. Google did it for BERT, and Answer.AI and LightOn did it for ModernBERT. What you download is that finished model, and your own training starts from there.

The second step is fine-tuning. You take the pretrained model, put one small output layer on top of it, and train it again on your own labeled examples. The BERT paper makes the point that this is nearly all you have to do, saying the pretrained model "can be fine-tuned with just one additional output layer" without "substantial task-specific architecture modifications".

Eight years later, bert-base-uncased is still downloaded about 46 million times a month from Hugging Face. If the job is to decide which of your five teams a support ticket should go to, or whether a customer review is positive or negative, a 110M model fine-tuned on a few thousand of your own labeled examples can often do it, and it answers in milliseconds. A frontier model costs more because it can reason and write, and neither is needed here.

What it is

It turns text into vectors, and a head you train turns those into answers

An encoder takes a piece of text and gives you back one vector for every token in it. A token is a word or a piece of a word, and a vector is a list of numbers describing that token in the context of everything around it. Those vectors don't answer anything by themselves. To get an answer out of the model you attach a head, which is a small output layer with one slot for each of your labels, and you train that head on examples of the decision you want it to make.

BERT also puts a special [CLS] token in front of every input you hand it. The paper calls the final vector at that position "the aggregate sequence representation for classification tasks", which is a formal way of saying that one vector stands for the whole piece of text. If you are classifying, your head reads that single vector and picks a label. If you are pulling pieces out of a sentence, say every date in a contract, your head reads every token's vector and labels each one separately.

A five-step figure titled "One sentence becomes twelve vectors, and the head reads one of them", described as a real pass through the ModernBERT classifier fine-tuned further down the page, on one held-out message. Step one is the text, "My card was charged twice for the same order." Step two is tokens, showing twelve chips reading CLS, My, card, was, charged, twice, for, the, same, order, a full stop and SEP, with the CLS chip highlighted, and a note that there are 12 tokens with CLS added in front. Step three is the encoder, a box reading ModernBERT-base, 149M parameters, 22 layers, 768 wide, which reads all 12 tokens in one pass, with a note that nothing is generated here. Step four is one vector per token, a stack of rows labeled CLS, My, card, was, charged and twice, each marked 768 numbers, with a note that there are 6 more, one per token. The CLS row is highlighted and its first values are shown as +0.11, +0.12, +0.18 and so on, under a note saying the head reads this one. Step five is the head, a box reading one layer, 768 to 77 labels, which outputs transaction_charged_twice at 99.8% and getting_spare_card at 0.1%, with a note that the other 75 labels share 0.1% and that the head is trained on your labels. The caption reads: the encoder is the same for every task, and the head is the part you train and the only part that knows your labels.
Every step here is one forward pass. Nothing is written and nothing is sampled, so the same message always comes back with the same numbers.

Every position reads the whole input in both directions, so by the time the model judges a word it has already read the words that come after it. The word charge means one thing in "they charged my card twice" and a different thing in "the battery will not charge", and an encoder has read the rest of the sentence before it settles on either one. The encoder, decoder and encoder-decoder article draws that out with a figure, if you want to see the mechanism.

Where it shows up

Encoders do the small jobs that run on every record

Classifying text into labels you already have examples of

This is the job the model was built for. You have a few thousand support tickets that somebody already routed, or a pile of reviews somebody already scored, and you want the same decision made on every new one that arrives. The labels stay the same from week to week, and the model runs on every single record. The measured run further down this page is exactly this job.

Pulling the pieces out of a sentence

Named entity recognition, usually written as NER, puts a label on each token, so one pass through a contract marks every person and every date in it. GLiNER takes that further by letting you name the entity types you care about at request time, because it puts those type names into the input right next to the text.

Reranking search results

A cross-encoder reads the query and one document together, in the same input, and gives you a score for that pair. Reading them together is what makes it accurate, and it also means nothing can be computed ahead of time, so it only runs on the shortlist that a faster search step already returned. The design goes back to monoBERT in 2019.

Producing embeddings

Most text embedding models are encoders with a pooling step on the end, which usually means averaging the token vectors into a single vector for the whole document. Search and deduplication over your own content both start there.

Guardrails in front of an LLM

Prompt injection detection, which is checking whether the user's text is trying to override the instructions you gave the model, has to run on every single request, so the model doing the checking needs to be small and fast. protectai/deberta-v3-base-prompt-injection-v2 is a 184M encoder that returns the probability that an input is an injection attempt, and it gets about 850,000 downloads a month.

Turning a pile of text into columns you can chart

Running a fine-tuned classifier over a few million rows is cheap enough to put on a schedule and let it run overnight. Every review comes back with a topic and a sentiment attached, and free text you could only read becomes a column you can sort, filter and chart.

How it works

Mask words to pretrain it, then add a head to fine-tune it

Pretraining masks words and asks for them back

BERT masks 15% of the tokens in each sequence and tries to predict them, reading both sides of every gap. Hand ModernBERT the sentence "They [MASK] my card twice for the same order" and it answers charged at 33.5%, used at 8.7% and sent at 3.3%, working only from the words around the hole.

It can't be trained to predict the next word the way an LLM is, because a model that reads in both directions has already seen that word. The paper puts it as "bidirectional conditioning would allow each word to indirectly see itself". Masking gives it something to predict that it can't already see.

Later encoders changed parts of this. RoBERTa trained for longer on more data and dropped BERT's second training objective, which had asked the model whether two sentences followed each other in the original document. DeBERTaV3 swapped masked language modeling for replaced token detection, where the model looks at every token and decides whether it is the original word or a substitute that was slipped in. Every position in the sequence carries a training signal that way, which the paper describes as a more sample-efficient task.

Fine-tuning attaches your head and moves every weight

You load the pretrained encoder, attach a fresh output layer with one slot per label, and train on your labeled examples. The head starts out random, so at the first step it is guessing. Every other weight in the model starts pretrained and keeps moving while you train, which is why a few thousand examples are usually enough.

finetune.py — everything except the training loop
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "answerdotai/ModernBERT-base"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
    model_id, num_labels=77)                   # 77 output slots, randomly set

batch = tok(texts, padding="max_length", truncation=True,
            max_length=64, return_tensors="pt")
batch["labels"] = label_ids                    # one integer per example
loss = model(**batch).loss                     # backward, step, repeat

What a modern encoder changed

ModernBERT came out in December 2024, and it is the same design with eight years of newer engineering applied to it. It was trained on 2 trillion tokens and it reads up to 8,192 tokens at a time, where the original BERT stopped at 512. Its layers alternate between global attention, where every token reads every other one, and a 128-token sliding window, and the paper says "every third layer" does the global pass. It also strips out padding tokens, which are the filler added to short inputs so every row in a batch is the same length, so the model computes only on real words. The result is a 149M and a 395M model that process long sequences about twice as fast as the encoders before them.

A measured run on 13,083 real messages

Banking77 is a public dataset from Casanueva et al. (2020). It holds 13,083 real messages that customers sent to a bank, each one labeled with one of 77 intents, and it comes already split into 10,003 messages for training and 3,080 for testing. It stands in well for the kind of work people put a classifier on, because 77 labels is more than a prompt handles comfortably and a lot of those intents sit close to each other.

A figure titled "A fine-tuned 149M encoder against a prompted 3B model", on Banking77, 13,083 real customer messages over 77 intents, with 10,003 for training and 3,080 held out. The first card, outlined in terracotta, is ModernBERT-base, fine-tuned, described as encoder-only, 149M, with a head trained on the 77 labels. Its accuracy bar reads 93.4% on the 77-way intent task, and the facts beside it say it trained in 16 minutes on a laptop GPU, takes 77 ms per example on the CPU, produces zero answers outside the label list by construction, and was tested on all 3,080 held-out messages. The second card is Qwen2.5-3B, prompted, described as decoder-only, 3B, with all 77 label names in the system prompt. Its accuracy bar reads 47.3%, with no training, 150 ms per example median on the same machine, 4.3% of answers outside the label list, and 300 of the same held-out messages tested. A box below says 4.3% of the 3B model's answers are label names that do not exist, listing three of them, top_up_by_card, top_up_fee_charged and unsupported_transaction_type, none of which is one of the 77, and notes that the fine-tuned encoder has 77 output slots, so every answer is one of them, even for a message none of them fits. The caption reads: on this task, fine-tuning on 10,003 labeled examples beat prompting a 3B model, 93.4% to 47.3%.
Both numbers come from the same laptop on the same day. The runs differ in the model and its size as well as the labeled examples, so the gap can't be put down to any one of them.

Fine-tuning ModernBERT-base for three epochs took 15.6 minutes on the laptop's GPU, and it got 93.4% of the 3,080 held-out messages right. Prompting Qwen2.5-3B with all 77 label names in the system prompt, on 300 of those same messages, got 47.3%. Once it is trained, the encoder answers in 77 milliseconds on the CPU, with no GPU and no API call involved.

That figure is about this one task and these two models. A bigger prompted model would probably score higher than 47.3%, and this run didn't test one, so it can't say whether a larger model or a frontier model prompted with the same 77 labels would pass 93.4%. What the run does show is that 10,003 labeled examples took a 149M model to 93.4% on this task, and that the model runs on hardware you already own.

The two models also fail in different ways. 4.3% of the 3B model's answers were label names that do not exist, assembled out of the ones that do, like top_up_by_card when the real label is topping_up_by_card. A fine-tuned encoder has exactly 77 output slots, so there is nothing else it could return. Every mistake it makes on these test messages is a confusion between two real labels, which you can find in a confusion matrix and fix by labeling more examples of the pairs it keeps mixing up.

The same 77 slots mean it can't say "none of these". The softmax spreads all of its probability across the 77 intents, so a message about something no intent covers, like a question about the weather, still comes back as one of them. Banking77's test messages all belong to one of the 77 intents, so this run doesn't measure how often that happens. The usual fixes are an extra other label trained on real off-topic messages, and a threshold on the top probability that sends low answers to a person. Hendrycks and Gimpel (2017) found that "correctly classified examples tend to have greater maximum softmax probabilities than erroneously classified and out-of-distribution examples", which is why that threshold catches some of them. The tendency is an average, so some off-topic messages still get a high probability.

Versions

These are the checkpoints you'll see, as of September 2026

Encoder checkpoints worth knowing, with monthly downloads read 2026-09-20
ModelReleasedSizeLicenseDownloads a month
google-bert/bert-base-uncasedOctober 2018110MApache-2.046.6M
FacebookAI/roberta-baseJuly 2019125MMIT7.9M
distilbert/distilbert-base-uncasedOctober 201966MApache-2.07.6M
microsoft/deberta-v3-baseNovember 2021184MMIT2.9M
answerdotai/ModernBERT-baseDecember 2024149MApache-2.04.4M
answerdotai/ModernBERT-largeDecember 2024395MApache-2.00.6M
jhu-clsp/mmBERT-baseSeptember 2025307M, 110M without embeddingsMIT0.4M

ModernBERT is what to start from for new work in English, since it reads 8,192 tokens at a time and carries an Apache-2.0 license. mmBERT is the multilingual one, pretrained on 3 trillion tokens covering over 1,800 languages, so reach for it when the text you are working with is not in English.

The models running in production today are years old. A fine-tuned encoder keeps doing its job, and moving to a newer one costs a retraining run, so it's worth doing when a test on your own held-out data shows a gain.

Choosing

Choose an encoder when the labels are fixed and the volume is high

Which model to reach for
The jobReach forWhy
Classify into fixed labels you have examples ofA fine-tuned encoderThe accuracy comes from your labels, and the model is small enough to run anywhere
Classify with no labeled dataAn LLM with the categories in the prompt, or JevNothing to train, and you can change the categories in the request
Labels that change every weekAn LLM or JevA new label means a new request, where an encoder means retraining
High volume with stable labelsA fine-tuned encoderA 149M model on your own hardware costs a fraction of an API call
Data that can't leave your networkA fine-tuned encoderOpen weights that run on your own CPU, with nothing sent anywhere
Write a reply or a summaryAn LLMAn encoder produces vectors and can't write
Rank documents for a queryA cross-encoder rerankerIt reads the query and the document together
Find similar textAn embedding modelOne vector per document, compared by cosine similarity

If you are starting from nothing, the order that usually works is to prompt a model first and fine-tune an encoder later, once the volume is high enough for the cost to show up on a bill or the accuracy stops being good enough. The prompting stage is also how you get your labeled examples, because a good LLM's output, corrected by a person, is a labeled dataset.

Try it

How to try it

Before you point any of this at your own data, it is worth running it once on a public dataset so you know the pipeline works end to end. This script loads the dataset from the figure above and trains the classifier on it.

train_banking77.py — a 77-way classifier in one file
import torch
from datasets import load_dataset
from transformers import (AutoModelForSequenceClassification, AutoTokenizer,
                          Trainer, TrainingArguments)

ds = load_dataset("PolyAI/banking77")          # 10,003 train, 3,080 test
labels = ds["train"].features["label"].names   # 77 intents

model_id = "answerdotai/ModernBERT-base"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
    model_id, num_labels=len(labels))

enc = ds.map(lambda b: tok(b["text"], truncation=True, max_length=64),
             batched=True)

Trainer(model=model, train_dataset=enc["train"], eval_dataset=enc["test"],
        args=TrainingArguments(output_dir="out", num_train_epochs=3,
                               per_device_train_batch_size=32,
                               learning_rate=5e-5)).train()

On a laptop GPU this finishes in about 16 minutes. When you swap in your own CSV, nothing else in the file has to change.

Ask your AI coding tool

Fine-tune an encoder on my labeled text and tell me whether it's good enough to ship. Read data.csv with text and label columns, hold out 20% stratified by label, and fine-tune answerdotai/ModernBERT-base with AutoModelForSequenceClassification for 3 epochs. Report accuracy and macro F1 overall and per label, and save a confusion matrix as a PNG with the 10 most-confused label pairs listed underneath. Then run a learning-curve experiment that retrains on 10%, 25%, 50% and 100% of the training rows and plots accuracy against the number of examples, so I can see whether more labeling would help. Finally, measure median and p95 latency per example on CPU at batch size 1 and at batch size 32.

Limits

What it can't do

  • It can't write anything. There is no way to get a summary or a written reply out of it, because every answer it produces is a label or a score coming out of the head you trained.
  • It can't reject an input that fits none of its labels. Every input gets one of the labels you trained, so an off-topic message needs an other label of its own or a threshold on the probability to be caught.
  • Its labels are fixed the moment you train it. Adding a category means collecting new examples and running training again, where a prompted model takes the new category in the request.
  • It needs labeled data. A few dozen examples per label is the rough floor, and the accuracy in the figure came from about 130 examples per intent.
  • Its probabilities need checking before you put a threshold on them. The head's scores are turned into percentages by a softmax, and a 0.92 there does not mean the model is right 92% of the time. The usual fix is temperature scaling, which is one number, fitted on held-out examples, that flattens or sharpens every probability at once.
  • Long inputs are still a constraint. Older encoders stop at 512 tokens, so a long document has to be split into chunks first. ModernBERT's 8,192 is recent enough that most of the fine-tuning recipes you will find online still assume the old limit.
  • It carries whatever was in its pretraining data. The English checkpoints were trained on web text, and the biases and blind spots that came with that text survive fine-tuning.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.