What it is
An encoder reads a whole input at once, and a decoder writes one token at a time
The original transformer (Vaswani et al., 2017) was built for translation and had both halves. An encoder turned the French sentence into vectors, and a decoder wrote the English sentence one token at a time while reading them. Everything since is built from one half of it, or from both.

The three shapes differ in what they hand back.
- An encoder returns one vector per input token. Those vectors are a representation of the text, and on their own they answer nothing. Something has to turn them into an answer, usually a head, which is a small output layer trained on your task that turns the vectors into a label or a score.
- A decoder returns a probability for every token that could come next, at every position. Read that distribution at the last position and you have a one-step answer. Feed the chosen token back in and you have generation.
- An encoder-decoder runs an encoder over the input and a decoder over the output, joined by cross-attention, which lets every decoder position read the encoder's finished vectors. The input gets read in both directions, and the output still gets written one token at a time.
Which models use this
Which models are which shape
The shape is rarely in the model's name, so this table is the quick way to place a model you're about to use.
| Shape | Models | What people use them for |
|---|---|---|
| Encoder-only | BERT, RoBERTa, DeBERTa-v3, ModernBERT | Classification and entity tagging on your own labels |
| Encoder-only | BERT-family embedding models such as bge and e5, and BERT-family cross-encoder rerankers | Search and ranking |
| Decoder-only | Qwen3-Embedding and Qwen3-Reranker, which are built on Qwen3 language models | Search and ranking, from a causal backbone |
| Encoder-only | ViT, the CLIP and SigLIP image towers, DINOv3 | Turning an image into vectors |
| Decoder-only | GPT, Claude, Gemini, Llama, Qwen, DeepSeek | Chat, code and agent work |
| Encoder-decoder | T5, Flan-T5, BART, mT5 | Translation, summarization and older fine-tuned pipelines |
| Encoder-decoder | Whisper, most dedicated translation models | Audio in and text out, or one language into another |
| Encoder-decoder | T5Gemma and T5Gemma 2 | A current encoder-decoder built by converting Gemma |
Three entries in that table catch people out.
Whisper is an encoder-decoder where the two halves take different kinds of input. Its encoder reads a spectrogram of the audio, and its decoder writes text while attending to what the encoder produced. The same wiring covers vision-language models, where an image encoder feeds a text decoder.
Search and ranking show up under two shapes. A reranker reads the query and the document together and returns one score, which is a job, and the job doesn't fix the shape. Qwen3-Reranker-0.6B loads as Qwen3ForCausalLM and reads the query and the document in one prompt. Its model card computes the score from the logits of the tokens "yes" and "no" at the last position. Qwen3-Embedding is built on the same causal backbone, read on its Hugging Face config on 2026-09-24, so the shape of a search model has to be read off its card.
One layer of an LLM is a decoder layer. The diagram in the transformer article shows one, because that's what an LLM is made of. An encoder layer is the same block with the causal mask removed, and an encoder-decoder's decoder layer adds one cross-attention sublayer between self-attention and the feed-forward network.
How it works
The mask decides the shape, and the pretraining objective follows from it
The mask is one line of arithmetic
Attention scores every position against every other position, then mixes their values by those scores. A causal mask sets the scores for later positions to negative infinity before the softmax, so they come out as zero weight. Removing that line makes the same layer bidirectional. Nothing else about the block changes, which is why Encoder-Decoder Gemma (Zhang et al., 2025) can describe its encoder as having "exactly the same architecture as the decoder-only model, but self-attention is switched from causal to bidirectional".
Removing the mask breaks next-token training, so encoders are trained differently
A model that can see the whole sentence cannot be trained to predict the next word, because the answer is already in the input. BERT's authors put it plainly, that "bidirectional conditioning would allow each word to indirectly see itself". So an encoder is trained by hiding parts of the input and asking for them back. BERT masks 15% of the tokens and predicts those. T5 removes whole spans and asks the decoder to write them out. A decoder needs none of this, because the causal mask already hides the answer at every position.
| Shape | Pretraining objective | What that produces |
|---|---|---|
| Encoder-only | Masked language modeling, where 15% of tokens are hidden and predicted | A representation built to describe an input |
| Decoder-only | Next-token prediction over ordinary text | A model that can continue any text, and so can answer anything phrased as a continuation |
| Encoder-decoder | Span corruption, where removed spans are regenerated, or the UL2 mixture of both | A model that reads an input and writes a different output |
The same three tickets, run through all three
Running one small model of each shape on the same three support tickets shows what that difference costs you in practice. Every probability below is renormalized over billing, technical and sales.

The encoder answered a question whose evidence came later in the text. The blank sits in the first sentence and the ticket follows it, and the fill changes with the ticket. Hand the same prefix to the decoder and it returns one distribution for all three tickets, because at that position it has read nothing else. The decoder's answer is good once the ticket comes first, which is the same constraint in another form.
The decoder's probabilities are saturated. Two of its three answers round to 100%, from a model that gets 47.3% of a harder 77-label version of this task right. The Jev page covers why instruction tuning does this to a model's probabilities.
The oldest and smallest model got one wrong. Flan-T5-base dates from 2022 and puts 73.7% on billing for a ticket about a demo and a quote. The shape it was built with says nothing about that.
The shape and the task interface are separate choices
The attention mask is one decision, and what the model hands back to your code is a separate one. Whether a query and a document are read together is separate again. People often fold these into the words "encoder" and "decoder", and the table shows each of them going either way on either shape.
| Choice | One option | Another option | Examples on both shapes |
|---|---|---|---|
| Fixed or dynamic labels | A head trained on one label list | The labels go into the input with the text | ModernBERT with a head is fixed, GLiClass on an encoder is dynamic, and a decoder with AutoModelForSequenceClassification is fixed |
| Generating or scoring | Write tokens out | Read one score or distribution at a position | A decoder can do either, and Qwen3-Reranker reads the "yes" and "no" logits |
| Joint or separate encoding | Query and document read in one pass | Each one embedded alone and compared by a dot product | A BERT cross-encoder and Qwen3-Reranker are joint, and bge and Qwen3-Embedding are separate |
So a question like "can this model take a new category at request time" is about the interface, and "can this document be embedded before the query arrives" is about joint or separate encoding. The mask only tells you whether a position can read what comes after it.
Why decoder-only won, and what the research says
Wang et al. (2022) trained models over 5 billion parameters on more than 170 billion tokens to compare the shapes directly, and found "the best objective and architecture is the opposite" in two settings. A causal decoder trained on plain language modeling is strongest when you evaluate it straight after pretraining, which is the zero-shot prompting that made LLMs useful. An encoder-decoder trained with masked language modeling is strongest once you add multitask finetuning.
Practice followed the first result. Prompting was what people wanted, and decoder-only models had two more advantages in deployment. A KV cache lets a decoder reuse everything it has already read when it writes the next token, and one stack with one objective is easier to scale than two stacks with a corruption scheme.
The second result keeps getting rediscovered. Encoder-Decoder Gemma converts a pretrained Gemma 2 into an encoder-decoder and reports that under a similar inference budget these models reach "comparable (often better) pretraining performance but substantially better finetuning performance than their decoder-only counterpart". Google shipped that work as T5Gemma in July 2025, and T5Gemma 2 followed on December 18, 2025, built on Gemma 3 with 128K of context and support for over 140 languages. Both are released under the Gemma license, which has its own use restrictions.
Fine-tuning
Fine-tuning changes with the shape
Fine-tuning means continuing to train a pretrained model on your own data. What changes between the shapes is what you attach to the model and what a training example looks like.
An encoder is fine-tuned by putting a head on it
You load the pretrained encoder, attach an output layer sized to your labels, and train the whole thing on your labeled examples. BERT's abstract describes exactly this, that the pretrained model "can be fine-tuned with just one additional output layer" without "substantial task-specific architecture modifications". In transformers it's one class.
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "answerdotai/ModernBERT-base" # 149M parameters
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, num_labels=len(labels)) # the head starts random
batch = tok(texts, padding="max_length", truncation=True,
max_length=64, return_tensors="pt")
batch["labels"] = label_ids # one integer per example
loss = model(**batch).loss # then loss.backward() as usualThe head is the only new part, and every other weight moves too. A training example is a piece of text and a label, which is the cheapest kind of labeled data to produce. The BERT article walks a full run of this on 10,003 labeled messages.
A decoder is fine-tuned on examples of what to write
The default is supervised fine-tuning, where each example is an input and the response you want, and the loss is the model's next-token loss over the response. Nothing is attached, because the model already has a head over the vocabulary.
Full fine-tuning of a large decoder is expensive, so most people use LoRA. Hu et al. (2021) freeze the pretrained weights and train small rank-decomposition matrices inside each layer, which for GPT-3 175B "reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times", with "no additional inference latency" once the adapter is merged back in.
Two other routes come up once the output is a label.
- Attach a classification head to the decoder.
AutoModelForSequenceClassificationworks with Llama and Qwen as well, and it reads the hidden state at the last position. You get an encoder-style classifier out of a decoder backbone, and you pay for the size. - Attach nothing and read the logits. Give the model your options and read its probabilities for the option tokens at the answer position. No training, and the answer can only be one of your options.
An encoder-decoder is fine-tuned on pairs of sequences
Each example is an input sequence and a target sequence, and the loss is the decoder's next-token loss over the target. T5 made this the whole interface by writing every task, including classification, as text in and text out. AutoModelForSeq2SeqLM plus a labels field of tokenized targets is the whole setup.
Building your own Jev, which is what people did in the week after it launched
Jev takes text and a list of allowed answers, and returns a probability for each answer with no generated text. TypeSafe hasn't published its architecture, and within days of the September 15, 2026 launch there were public attempts to rebuild the shape from open models. They split along the line this page is about, and both routes work.
From a decoder, with no training at all. open-alternative-jev (Apache-2.0, published September 18, 2026) renders each question as a chat turn and reads the logits at each answer position, restricted to the option letters, out of one forward pass. Its README reports 92.9% on a 1,000-question RACE-H sample with Qwen3.6-27B, at 2.5 times fewer tokens than asking each question separately. It says plainly that it is "not a reproduction of Jev" and that it packages a capability "that ordinary open models already have and that chat APIs hide". The measured caveats are worth copying. Raw confidence came out about 5 points too high, one fitted temperature brought expected calibration error from 5.4% to 2.1%, and packing several questions into one sequence changed 6% to 9% of individual answers while aggregate accuracy held.
From a decoder, by fine-tuning it. kev (Apache-2.0, September 17, 2026) is a LoRA adapter and a readout head on a Qwen3 base model at 0.6B, 4B and 8B, with a block-causal mask so each question can read the document and never a sibling question. Its head is "trained with cross-entropy against labelled outcomes, so the probabilities are learned rather than generated as text", and on sources it never trained on its README scores kev-4b at 0.79 and kev-8b at 0.80 against 0.86 for Jev itself. Bespoke Nimble fine-tunes Qwen3.5-9B with LoRA on the answer tokens alone, and reports matching 90.1% of reference labels on 324 held-out examples, against 66.4% for the base model and 93.2% for jev-1.13.0. Its authors say they did not distill from Jev and that they built it in one day.
From an encoder, by training heads on it. OpenJev (Verdict) post-trains a 151M ModernBERT with a GLiClass backbone to return choices, ordinal scores and yes/no probabilities in "a single non-autoregressive forward pass", which its README times at under 35 milliseconds. A 151M encoder is two orders of magnitude smaller than a 27B decoder. Its README describes the candidate labels going into the encoder's input next to the document, separated by [TEXT] and [LABEL] markers, so a new option list arrives with the request the same way it does on the decoder route.
The encoder route has a longer history than the launch suggests. GLiNER (2023) and GLiClass (2025) both put the candidate labels into the encoder's input alongside the text, score them in one bidirectional pass, and accept a new label list at inference with no retraining.
Compare the three routes to a typed decision on my own data. I have tickets.csv with a text column and a label column over a fixed list of categories. For route one, fine-tune answerdotai/ModernBERT-base with AutoModelForSequenceClassification on 80% of the rows. For route two, prompt Qwen/Qwen2.5-3B-Instruct with the category list and parse its answer. For route three, run the same model with no fine-tuning, read the logits at the answer position restricted to the first token of each category, and softmax those. Evaluate all three on the held-out 20% and report accuracy, the rate of answers that fall outside my category list, median latency per example and peak memory. Then add a reliability diagram for routes one and three, binning predictions by top probability in 0.1-wide bins and plotting accuracy per bin, so I can see which one's probabilities I can threshold.
Choosing
Choosing a shape
Most of the time you're choosing a model and inheriting its shape. The choice becomes real when you're building something that runs often enough for size to matter.
| The job | Reach for | Why |
|---|---|---|
| Classify text into labels you have examples of | An encoder with a head | The smallest model that does the job, and the head is trained on your labels |
| Rank or score pairs | An embedding model for recall, then a reranker | Either shape works here, since BERT-family and Qwen3-based models both do this job |
| Write anything, or work as an agent | A decoder | Generation is what the causal mask is for |
| Decide between options you can list, at low volume | A decoder with its option logits read | No training, and the answer stays inside your list |
| Decide between options you can list, at high volume | An encoder with a head, or Jev | A 150M encoder runs for a fraction of a 3B decoder's cost |
| Translate or summarize with a fixed input and output | An encoder-decoder | The input is read bidirectionally and the output is short |
| Turn audio or images into text | An encoder-decoder | The encoder takes the signal, and the decoder writes |
Across the table, the shape decides which positions each token can read, and the task interface and the training data decide what the model returns and how well.
Limits
What the shape does not fix
- An encoder can't write. It returns vectors, so every answer has to come from a head you trained or a lookup you built.
- A decoder can't read ahead. Evidence has to come before the position where the answer is read, which is why prompts put the document first and the question last.
- Neither shape gives you calibrated probabilities for free. A decoder that has been instruction-tuned is usually overconfident, and both of the open Jev reproductions measured above had to fit a temperature on held-out labels to fix it.
- A classification head fixes its labels at training time. Adding a category means retraining that head, whether it sits on an encoder or on a decoder. GLiClass and OpenJev show an encoder taking a new label list in its input, so the fixed list belongs to the head.
- Converting between shapes is real work. Wang et al. and Encoder-Decoder Gemma both show it can be done from a pretrained checkpoint, and both spend a substantial pretraining budget doing it.
- The shape says nothing about the license. T5Gemma 2 is Gemma-licensed, ModernBERT is Apache-2.0, and several of the Jev reproductions inherit whatever their base model carries.
Go deeper
Vaswani et al. (2017): Attention Is All You Need · Devlin et al. (2018): BERT, Pre-training of Deep Bidirectional Transformers · Raffel et al. (2019): Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) · Lewis et al. (2019): BART, Denoising Sequence-to-Sequence Pre-training · Wang et al. (2022): What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization? · Tay et al. (2022): UL2, Unifying Language Learning Paradigms · Zhang et al. (2025): Encoder-Decoder Gemma, Improving the Quality-Efficiency Trade-Off via Adaptation · Hu et al. (2021): LoRA, Low-Rank Adaptation of Large Language Models · Warner et al. (2024): ModernBERT
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
