MLGuerrillaStart with M1 →
Architectures·18 min read·Updated 24 September 2026

Encoder-only, decoder-only and encoder-decoder

These are the three ways a transformer can be wired. An encoder reads a whole input at once, a decoder writes one token at a time, and an encoder-decoder does one of each.

Encoder and decoder name two ways of reading a sequence. In an encoder, every position can read every other position, including the ones that come later. In a decoder, each position can read only what came before it. The BERT paper puts the naming in a footnote, where the bidirectional version "is often referred to as a Transformer encoder while the left-context-only version is referred to as a Transformer decoder since it can be used for text generation".

That one rule about what each position can see decides what the model is good for, how it's pretrained, what comes out of it and how you fine-tune it. This page covers all four, and it ends on how people are building their own version of Jev out of open models, which is the same question in a current form.

What it is

An encoder reads a whole input at once, and a decoder writes one token at a time

The original transformer (Vaswani et al., 2017) was built for translation and had both halves. An encoder turned the French sentence into vectors, and a decoder wrote the English sentence one token at a time while reading them. Everything since is built from one half of it, or from both.

A figure titled "The three shapes differ in what each position is allowed to read", with three stacked cards. The first card, Encoder-only, says every position reads every other position in both directions, and shows six token chips reading the, export, button, throws, a, 500, with the third chip highlighted and arcs connecting all six to it, captioned "the highlighted position reads all six". Its attention mask is a 6 by 6 grid with every cell filled, labeled "all open". What comes out is one vector per token, with a head on top that turns them into your labels, and the models built this way are BERT, RoBERTa, DeBERTa, ModernBERT, embedding models and rerankers, and ViT, the CLIP and SigLIP towers and DINOv3. The second card, Decoder-only, says each position reads only the positions before it, and shows the same six chips with the fifth highlighted and arcs reaching it only from the four on its left, captioned "position 5 reads 1 to 4, and nothing after it". Its attention mask is a lower triangle. What comes out is a probability for every next token, read at the last position or used to keep writing, and the models are GPT, Claude, Gemini, Llama, Qwen, reasoning models and agents, and almost every model sold as an LLM. The third card, Encoder-decoder, shows four input chips labeled "encoder, bidirectional" with an arrow into three output chips labeled "decoder, causal", noting that cross-attention carries the encoder's output into every decoder position, with one full mask and one triangular mask side by side. What comes out is a generated sequence conditioned on everything the encoder read, and the models are T5, Flan-T5, BART, mT5, Whisper, translation models and T5Gemma 2. The caption reads: the original transformer had both halves, and every model since is one half, the other half, or both.
The layers themselves are identical in all three. What changes is the attention mask and what you read off the end.

The three shapes differ in what they hand back.

  • An encoder returns one vector per input token. Those vectors are a representation of the text, and on their own they answer nothing. Something has to turn them into an answer, usually a head, which is a small output layer trained on your task that turns the vectors into a label or a score.
  • A decoder returns a probability for every token that could come next, at every position. Read that distribution at the last position and you have a one-step answer. Feed the chosen token back in and you have generation.
  • An encoder-decoder runs an encoder over the input and a decoder over the output, joined by cross-attention, which lets every decoder position read the encoder's finished vectors. The input gets read in both directions, and the output still gets written one token at a time.

Which models use this

Which models are which shape

The shape is rarely in the model's name, so this table is the quick way to place a model you're about to use.

Common models by shape, as of September 2026
ShapeModelsWhat people use them for
Encoder-onlyBERT, RoBERTa, DeBERTa-v3, ModernBERTClassification and entity tagging on your own labels
Encoder-onlyBERT-family embedding models such as bge and e5, and BERT-family cross-encoder rerankersSearch and ranking
Decoder-onlyQwen3-Embedding and Qwen3-Reranker, which are built on Qwen3 language modelsSearch and ranking, from a causal backbone
Encoder-onlyViT, the CLIP and SigLIP image towers, DINOv3Turning an image into vectors
Decoder-onlyGPT, Claude, Gemini, Llama, Qwen, DeepSeekChat, code and agent work
Encoder-decoderT5, Flan-T5, BART, mT5Translation, summarization and older fine-tuned pipelines
Encoder-decoderWhisper, most dedicated translation modelsAudio in and text out, or one language into another
Encoder-decoderT5Gemma and T5Gemma 2A current encoder-decoder built by converting Gemma

Three entries in that table catch people out.

Whisper is an encoder-decoder where the two halves take different kinds of input. Its encoder reads a spectrogram of the audio, and its decoder writes text while attending to what the encoder produced. The same wiring covers vision-language models, where an image encoder feeds a text decoder.

Search and ranking show up under two shapes. A reranker reads the query and the document together and returns one score, which is a job, and the job doesn't fix the shape. Qwen3-Reranker-0.6B loads as Qwen3ForCausalLM and reads the query and the document in one prompt. Its model card computes the score from the logits of the tokens "yes" and "no" at the last position. Qwen3-Embedding is built on the same causal backbone, read on its Hugging Face config on 2026-09-24, so the shape of a search model has to be read off its card.

One layer of an LLM is a decoder layer. The diagram in the transformer article shows one, because that's what an LLM is made of. An encoder layer is the same block with the causal mask removed, and an encoder-decoder's decoder layer adds one cross-attention sublayer between self-attention and the feed-forward network.

How it works

The mask decides the shape, and the pretraining objective follows from it

The mask is one line of arithmetic

Attention scores every position against every other position, then mixes their values by those scores. A causal mask sets the scores for later positions to negative infinity before the softmax, so they come out as zero weight. Removing that line makes the same layer bidirectional. Nothing else about the block changes, which is why Encoder-Decoder Gemma (Zhang et al., 2025) can describe its encoder as having "exactly the same architecture as the decoder-only model, but self-attention is switched from causal to bidirectional".

Removing the mask breaks next-token training, so encoders are trained differently

A model that can see the whole sentence cannot be trained to predict the next word, because the answer is already in the input. BERT's authors put it plainly, that "bidirectional conditioning would allow each word to indirectly see itself". So an encoder is trained by hiding parts of the input and asking for them back. BERT masks 15% of the tokens and predicts those. T5 removes whole spans and asks the decoder to write them out. A decoder needs none of this, because the causal mask already hides the answer at every position.

What each shape is trained to do
ShapePretraining objectiveWhat that produces
Encoder-onlyMasked language modeling, where 15% of tokens are hidden and predictedA representation built to describe an input
Decoder-onlyNext-token prediction over ordinary textA model that can continue any text, and so can answer anything phrased as a continuation
Encoder-decoderSpan corruption, where removed spans are regenerated, or the UL2 mixture of bothA model that reads an input and writes a different output

The same three tickets, run through all three

Running one small model of each shape on the same three support tickets shows what that difference costs you in practice. Every probability below is renormalized over billing, technical and sales.

A figure titled "Each shape arrives at an answer by reading the input a different way", comparing three models on three support tickets. The columns are ModernBERT-base, encoder-only at 149M, which fills a blank in "routed to the ___ team" with the ticket sitting after the blank; Qwen2.5-3B, decoder-only at 3B, which answers a chat turn asking for billing, technical or sales, with the ticket coming first; and Flan-T5-base, encoder-decoder at 248M, scored on the first token of its output, where the encoder sees the whole prompt. For the ticket "The customer was charged twice for the same invoice", where billing is right, ModernBERT gives billing 90.2%, technical 4.4%, sales 5.4%; Qwen gives billing 100.0%; Flan-T5 gives billing 58.1%, technical 22.1%, sales 19.8%. All three are correct. For "The export button throws a 500 error since yesterday's release", where technical is right, ModernBERT gives technical 70.9%, sales 23.6%, billing 5.5%; Qwen gives technical 100.0%; Flan-T5 gives technical 89.7%. All three are correct. For "The customer wants a demo and a quote for 50 seats", where sales is right, ModernBERT gives sales 88.7%; Qwen gives sales 97.6%; Flan-T5 gives billing 73.7% and sales only 9.5%, which is wrong. A boxed note at the bottom says that if you put the ticket after the blank and ask a decoder to fill it, every ticket gets the same answer, because Qwen2.5-3B at that position returns I 48.2%, ticket 24.3% and It 13.5% for all three tickets, having not read them yet. The caption says all three answer, and what changes is where the evidence has to sit and what the number attached to the answer means.
Flan-T5-base is from 2022 and has 248M parameters, and it misroutes the sales ticket. Shape sets what a model can do, and size and training data still decide how well it does it.

The encoder answered a question whose evidence came later in the text. The blank sits in the first sentence and the ticket follows it, and the fill changes with the ticket. Hand the same prefix to the decoder and it returns one distribution for all three tickets, because at that position it has read nothing else. The decoder's answer is good once the ticket comes first, which is the same constraint in another form.

The decoder's probabilities are saturated. Two of its three answers round to 100%, from a model that gets 47.3% of a harder 77-label version of this task right. The Jev page covers why instruction tuning does this to a model's probabilities.

The oldest and smallest model got one wrong. Flan-T5-base dates from 2022 and puts 73.7% on billing for a ticket about a demo and a quote. The shape it was built with says nothing about that.

The shape and the task interface are separate choices

The attention mask is one decision, and what the model hands back to your code is a separate one. Whether a query and a document are read together is separate again. People often fold these into the words "encoder" and "decoder", and the table shows each of them going either way on either shape.

Three choices that are often confused with the shape
ChoiceOne optionAnother optionExamples on both shapes
Fixed or dynamic labelsA head trained on one label listThe labels go into the input with the textModernBERT with a head is fixed, GLiClass on an encoder is dynamic, and a decoder with AutoModelForSequenceClassification is fixed
Generating or scoringWrite tokens outRead one score or distribution at a positionA decoder can do either, and Qwen3-Reranker reads the "yes" and "no" logits
Joint or separate encodingQuery and document read in one passEach one embedded alone and compared by a dot productA BERT cross-encoder and Qwen3-Reranker are joint, and bge and Qwen3-Embedding are separate

So a question like "can this model take a new category at request time" is about the interface, and "can this document be embedded before the query arrives" is about joint or separate encoding. The mask only tells you whether a position can read what comes after it.

Why decoder-only won, and what the research says

Wang et al. (2022) trained models over 5 billion parameters on more than 170 billion tokens to compare the shapes directly, and found "the best objective and architecture is the opposite" in two settings. A causal decoder trained on plain language modeling is strongest when you evaluate it straight after pretraining, which is the zero-shot prompting that made LLMs useful. An encoder-decoder trained with masked language modeling is strongest once you add multitask finetuning.

Practice followed the first result. Prompting was what people wanted, and decoder-only models had two more advantages in deployment. A KV cache lets a decoder reuse everything it has already read when it writes the next token, and one stack with one objective is easier to scale than two stacks with a corruption scheme.

The second result keeps getting rediscovered. Encoder-Decoder Gemma converts a pretrained Gemma 2 into an encoder-decoder and reports that under a similar inference budget these models reach "comparable (often better) pretraining performance but substantially better finetuning performance than their decoder-only counterpart". Google shipped that work as T5Gemma in July 2025, and T5Gemma 2 followed on December 18, 2025, built on Gemma 3 with 128K of context and support for over 140 languages. Both are released under the Gemma license, which has its own use restrictions.

Fine-tuning

Fine-tuning changes with the shape

Fine-tuning means continuing to train a pretrained model on your own data. What changes between the shapes is what you attach to the model and what a training example looks like.

An encoder is fine-tuned by putting a head on it

You load the pretrained encoder, attach an output layer sized to your labels, and train the whole thing on your labeled examples. BERT's abstract describes exactly this, that the pretrained model "can be fine-tuned with just one additional output layer" without "substantial task-specific architecture modifications". In transformers it's one class.

finetune_encoder.py — the parts that matter
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "answerdotai/ModernBERT-base"          # 149M parameters
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
    model_id, num_labels=len(labels))             # the head starts random

batch = tok(texts, padding="max_length", truncation=True,
            max_length=64, return_tensors="pt")
batch["labels"] = label_ids                       # one integer per example
loss = model(**batch).loss                        # then loss.backward() as usual

The head is the only new part, and every other weight moves too. A training example is a piece of text and a label, which is the cheapest kind of labeled data to produce. The BERT article walks a full run of this on 10,003 labeled messages.

A decoder is fine-tuned on examples of what to write

The default is supervised fine-tuning, where each example is an input and the response you want, and the loss is the model's next-token loss over the response. Nothing is attached, because the model already has a head over the vocabulary.

Full fine-tuning of a large decoder is expensive, so most people use LoRA. Hu et al. (2021) freeze the pretrained weights and train small rank-decomposition matrices inside each layer, which for GPT-3 175B "reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times", with "no additional inference latency" once the adapter is merged back in.

Two other routes come up once the output is a label.

  • Attach a classification head to the decoder. AutoModelForSequenceClassification works with Llama and Qwen as well, and it reads the hidden state at the last position. You get an encoder-style classifier out of a decoder backbone, and you pay for the size.
  • Attach nothing and read the logits. Give the model your options and read its probabilities for the option tokens at the answer position. No training, and the answer can only be one of your options.

An encoder-decoder is fine-tuned on pairs of sequences

Each example is an input sequence and a target sequence, and the loss is the decoder's next-token loss over the target. T5 made this the whole interface by writing every task, including classification, as text in and text out. AutoModelForSeq2SeqLM plus a labels field of tokenized targets is the whole setup.

Building your own Jev, which is what people did in the week after it launched

Jev takes text and a list of allowed answers, and returns a probability for each answer with no generated text. TypeSafe hasn't published its architecture, and within days of the September 15, 2026 launch there were public attempts to rebuild the shape from open models. They split along the line this page is about, and both routes work.

From a decoder, with no training at all. open-alternative-jev (Apache-2.0, published September 18, 2026) renders each question as a chat turn and reads the logits at each answer position, restricted to the option letters, out of one forward pass. Its README reports 92.9% on a 1,000-question RACE-H sample with Qwen3.6-27B, at 2.5 times fewer tokens than asking each question separately. It says plainly that it is "not a reproduction of Jev" and that it packages a capability "that ordinary open models already have and that chat APIs hide". The measured caveats are worth copying. Raw confidence came out about 5 points too high, one fitted temperature brought expected calibration error from 5.4% to 2.1%, and packing several questions into one sequence changed 6% to 9% of individual answers while aggregate accuracy held.

From a decoder, by fine-tuning it. kev (Apache-2.0, September 17, 2026) is a LoRA adapter and a readout head on a Qwen3 base model at 0.6B, 4B and 8B, with a block-causal mask so each question can read the document and never a sibling question. Its head is "trained with cross-entropy against labelled outcomes, so the probabilities are learned rather than generated as text", and on sources it never trained on its README scores kev-4b at 0.79 and kev-8b at 0.80 against 0.86 for Jev itself. Bespoke Nimble fine-tunes Qwen3.5-9B with LoRA on the answer tokens alone, and reports matching 90.1% of reference labels on 324 held-out examples, against 66.4% for the base model and 93.2% for jev-1.13.0. Its authors say they did not distill from Jev and that they built it in one day.

From an encoder, by training heads on it. OpenJev (Verdict) post-trains a 151M ModernBERT with a GLiClass backbone to return choices, ordinal scores and yes/no probabilities in "a single non-autoregressive forward pass", which its README times at under 35 milliseconds. A 151M encoder is two orders of magnitude smaller than a 27B decoder. Its README describes the candidate labels going into the encoder's input next to the document, separated by [TEXT] and [LABEL] markers, so a new option list arrives with the request the same way it does on the decoder route.

The encoder route has a longer history than the launch suggests. GLiNER (2023) and GLiClass (2025) both put the candidate labels into the encoder's input alongside the text, score them in one bidirectional pass, and accept a new label list at inference with no retraining.

Ask your AI coding tool

Compare the three routes to a typed decision on my own data. I have tickets.csv with a text column and a label column over a fixed list of categories. For route one, fine-tune answerdotai/ModernBERT-base with AutoModelForSequenceClassification on 80% of the rows. For route two, prompt Qwen/Qwen2.5-3B-Instruct with the category list and parse its answer. For route three, run the same model with no fine-tuning, read the logits at the answer position restricted to the first token of each category, and softmax those. Evaluate all three on the held-out 20% and report accuracy, the rate of answers that fall outside my category list, median latency per example and peak memory. Then add a reliability diagram for routes one and three, binning predictions by top probability in 0.1-wide bins and plotting accuracy per bin, so I can see which one's probabilities I can threshold.

Choosing

Choosing a shape

Most of the time you're choosing a model and inheriting its shape. The choice becomes real when you're building something that runs often enough for size to matter.

Which shape fits the job
The jobReach forWhy
Classify text into labels you have examples ofAn encoder with a headThe smallest model that does the job, and the head is trained on your labels
Rank or score pairsAn embedding model for recall, then a rerankerEither shape works here, since BERT-family and Qwen3-based models both do this job
Write anything, or work as an agentA decoderGeneration is what the causal mask is for
Decide between options you can list, at low volumeA decoder with its option logits readNo training, and the answer stays inside your list
Decide between options you can list, at high volumeAn encoder with a head, or JevA 150M encoder runs for a fraction of a 3B decoder's cost
Translate or summarize with a fixed input and outputAn encoder-decoderThe input is read bidirectionally and the output is short
Turn audio or images into textAn encoder-decoderThe encoder takes the signal, and the decoder writes

Across the table, the shape decides which positions each token can read, and the task interface and the training data decide what the model returns and how well.

Limits

What the shape does not fix

  • An encoder can't write. It returns vectors, so every answer has to come from a head you trained or a lookup you built.
  • A decoder can't read ahead. Evidence has to come before the position where the answer is read, which is why prompts put the document first and the question last.
  • Neither shape gives you calibrated probabilities for free. A decoder that has been instruction-tuned is usually overconfident, and both of the open Jev reproductions measured above had to fit a temperature on held-out labels to fix it.
  • A classification head fixes its labels at training time. Adding a category means retraining that head, whether it sits on an encoder or on a decoder. GLiClass and OpenJev show an encoder taking a new label list in its input, so the fixed list belongs to the head.
  • Converting between shapes is real work. Wang et al. and Encoder-Decoder Gemma both show it can be done from a pretrained checkpoint, and both spend a substantial pretraining budget doing it.
  • The shape says nothing about the license. T5Gemma 2 is Gemma-licensed, ModernBERT is Apache-2.0, and several of the Jev reproductions inherit whatever their base model carries.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.