MLGuerrillaStart with M1 →
Text and language·15 min read·Updated 23 September 2026

Encoder-decoder models

These models pair an encoder that reads the input with a decoder that writes the output. T5 made every task look like a string in and a string out, and the shape still runs translation and speech-to-text.

An encoder-decoder model has two stacks. The encoder reads the input in one pass with every position able to see every other, and the decoder writes the output one token at a time while reading the encoder's vectors through cross-attention. It's the shape the original transformer was built with in 2017, for translation.

T5 (Raffel et al., 2019) made the shape general by writing every task as text in and text out, including classification, where the model writes the label as a word. That framing is where the name comes from, the Text-to-Text Transfer Transformer.

What it is

The input goes in whole, and the output comes out one token at a time

Both stacks are built from the same transformer blocks. The decoder differs in two places, a mask on its own attention and a second attention layer pointed at the encoder's output.

A three-part figure titled "The encoder reads all 12 tokens at once, the decoder writes one at a time", built from google/flan-t5-base translating one sentence, with 12 encoder blocks, 12 decoder blocks, 768 wide and 12 heads. The top band lays the architecture out left to right. The input, "translate English to French: The export button is broken.", is 12 tokens. It feeds an ENCODER card marked x 12 blocks, whose four stacked layers are self-attention with no mask, add and norm, feed forward 768 to 2048, and add and norm, with a note that 12 vectors come out and are computed once. A dashed terracotta line leaves that output and enters the DECODER card, also marked x 12 blocks, at its cross-attention layer. The decoder's six stacked layers are masked self-attention, add and norm, cross-attention from the encoder, add and norm, feed forward, and add and norm, with a note that it runs one position at a time. The decoder feeds an OUTPUT box, linear plus softmax, giving a probability for each of 32,128 tokens per step. The middle band, labeled ZOOM inside the encoder's self-attention, shows the 12 input tokens as chips, translate, English, to, French, colon, The, export, button, is, broken, period and end-of-sequence, with the chip for export highlighted and arcs connecting it to all eleven others, noting that the highlighted token reads all 12 including the six that come after it, and that every other token does the same in the same pass. The bottom band, labeled ZOOM the decoder, step by step, shows all sixteen decoding steps as chips with the token chosen and its probability: Le 32%, bouton 63%, de 36%, a space at 50%, l 72%, an apostrophe at 80%, ex 97%, port 100%, ation 90%, est 89%, a space at 35%, ren 7%, for 19%, c 92%, e-acute 94% and a period at 99%. The finished sentence reads "Le bouton de l'exportation est renforce", with an accent on the last e. A note says sixteen steps means sixteen forward passes through the decoder against one pass through the encoder, that the last word is wrong, and that step 12 is where it hesitated at 7%. The caption reads: same blocks on both sides, and the decoder adds a mask on its own attention and a second attention pointed at the encoder.
One pass through the encoder against sixteen through the decoder, for this one sentence. A decoder-only model with a KV cache would make the same one pass over its prompt and then sixteen steps, each through its whole stack.

Splitting the work in half has two consequences.

The input is read once, in both directions, and the result is reused for every output token. A decoder-only model with a KV cache also reads its prompt once, in one pass called prefill, and reuses the stored attention keys and values at every step after that, so both shapes pass over the input once. The difference is in how the input is read and what each output step costs. The encoder lets every input token see every other one, where a decoder-only model reads its prompt left to right. And each output step in an encoder-decoder runs only the decoder stack, which can be much smaller than the encoder.

That second point is the reason the halves can be sized independently. Google's T5Gemma release pairs a 9B encoder with a 2B decoder for exactly this reason, because summarization needs a deep reading of a long input and writes a short output.

How it works

Span corruption is what trains both halves

T5's pretraining removes random spans from a sentence, replaces each with a numbered sentinel token, and asks the decoder to write the removed spans back in order. The encoder sees the damaged sentence, the decoder writes the missing pieces.

A figure titled "T5 removes spans from the input and has the decoder write them back", in two panels. The encoder panel, at 110M parameters and bidirectional, reads once, taking the input "Paris is the <extra_id_0> of <extra_id_1>." and noting that the sentinels mark where text was taken out, that nothing is generated on this side, and that the whole sentence is available at every position. An arrow labeled cross-attention leads to the decoder panel, at 138M parameters and causal, which writes one token at a time and produces "<extra_id_0> capital <extra_id_1> France", with a note that each step reads what it has written so far plus everything the encoder produced. Below, two more real decodes are shown: the input "The export button <extra_id_0> since yesterday's <extra_id_1>." produces "<extra_id_0> has been disabled <extra_id_1> update", and the input "The customer <extra_id_0> twice for the same invoice." produces "<extra_id_0> pays". A box at the bottom says every downstream task gets written in the same shape, because T5 turns classification, translation and summarization into one format, a string in and a string out. The caption reads: one objective trains the encoder to read and the decoder to write, which is why the two halves arrive together.
Real greedy decodes from t5-base. The model was never told these were questions. It's filling in text that was taken out, which is the only thing it was pretrained to do.

BART (Lewis et al., 2019) does the same job with a different damage function, shuffling sentences and replacing spans with a single mask token. Its paper describes the architecture as one that "can be seen as generalizing BERT (due to the bidirectional encoder), GPT (with the left-to-right decoder), and many other more recent pretraining schemes", which is the clearest one-line summary of where the shape sits.

Fine-tuning follows the same interface. Each example is an input string and a target string, the loss is the decoder's next-token loss over the target, and the class in transformers is AutoModelForSeq2SeqLM. Classification becomes a target of one word, summarization a target of one paragraph, and the training loop doesn't change between them.

Where it shows up

Where the shape is still the default

Speech to text

Whisper is an encoder-decoder whose encoder reads a spectrogram and whose decoder writes text. openai/whisper-large-v3 is downloaded about 4.7 million times a month, which makes it one of the most used encoder-decoders in the world.

Translation

Dedicated translation models stayed with the shape. facebook/nllb-200-distilled-600M covers 200 languages and gets about 1 million downloads a month, under a CC-BY-NC-4.0 license that rules out commercial use. Google's MADLAD-400 models are Apache-2.0 and cover more than 400 languages.

Summarization pipelines that already work

facebook/bart-large-cnn still pulls about 1.3 million downloads a month. It's a 2019 model fine-tuned on news summarization, and for the narrow job of turning an article into three sentences it's cheaper than an API call and needs no prompt.

Fine-tuned text-to-text jobs

Flan-T5 is the common starting point when the output is short text and you have examples. Extraction into a fixed format, normalization of messy strings, query rewriting for search. google/flan-t5-base at 248M parameters runs on a CPU.

Encoder-decoder LLMs

T5Gemma and T5Gemma 2 are the current attempt to bring the shape back for general-purpose work, and the versions section below covers what they are and what they cost you in licensing.

Versions

The models you'll see, as of September 2026

Encoder-decoder checkpoints, with monthly downloads read 2026-09-20
ModelReleasedWhat it doesLicenseDownloads a month
google-t5/t5-baseOctober 2019The original text-to-text modelApache-2.02.4M
facebook/bart-large-cnn2019News summarizationMIT1.3M
google/flan-t5-baseOctober 2022T5 after instruction tuningApache-2.01.4M
openai/whisper-large-v3November 2023Speech to textApache-2.04.7M
facebook/nllb-200-distilled-600MJuly 2022Translation across 200 languagesCC-BY-NC-4.01.0M
google/t5gemma-2-270m-270mDecember 2025A current encoder-decoder LLM, multimodalGemma45K

What T5Gemma changed

Google published Encoder-Decoder Gemma (Zhang et al., April 2025), which builds an encoder-decoder by copying the weights of a pretrained decoder-only model into both halves, switching the encoder's self-attention from causal to bidirectional, and continuing training with a span-corruption objective. The paper reports that under a similar inference budget the result reaches "comparable (often better) pretraining performance but substantially better finetuning performance than their decoder-only counterpart".

T5Gemma shipped on July 9, 2025, built from Gemma 2. T5Gemma 2 followed on December 18, 2025, built from Gemma 3, with three sizes from 270M-270M up to 4B-4B, a 128K context window, support for over 140 languages and image input. Its two architecture changes both save parameters. The embeddings are tied across the encoder and the decoder, and the decoder's self-attention and cross-attention are merged into one layer. Both releases carry the Gemma license, which has use restrictions that Apache-2.0 models don't.

Choosing

Choosing between this and a decoder

Which shape fits the job
The jobReach forWhy
Speech to textWhisper, or a hosted speech modelThe audio encoder is the whole point
Translation you run yourselfAn encoder-decoder translation modelSmall, fast, and trained on nothing else besides translation
Summarizing at high volume with a fixed output shapeA fine-tuned encoder-decoderA small model runs on your own hardware, and a small decoder keeps each output step cheap
Chat, agents, tool calls, long generationA decoder-only LLMEvery ecosystem, serving stack and technique assumes it
Classifying into fixed labelsAn encoder with a headA label is not text worth generating
Quality with no engineering effortA frontier decoder-only model through an APIThe largest general-purpose models are decoder-only

The size of the model and what it was trained on usually count for more than the shape, and the summary in the next section shows what a 248M model does with an instruction it wasn't fine-tuned for.

Try it

How to try it

AutoModelForSeq2SeqLM loads every text model in the table above the same way. Whisper takes audio, so it has its own class. This one turns a support ticket into a one-sentence summary.

seq2seq.py — an instruction in, an answer out
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

model_id = "google/flan-t5-base"                  # 248M, Apache-2.0
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

prompt = ("Summarize this support ticket in one short sentence: "
          "Hi, since yesterday's update the export button on the reports "
          "page does nothing when I click it. I need the CSV for a meeting "
          "on Friday.")
ids = tok(prompt, return_tensors="pt")            # 45 tokens in
out = model.generate(**ids, max_new_tokens=40)
print(tok.decode(out[0], skip_special_tokens=True))
# -> Exporting reports does not work on the CSV page.

The summary gets the problem right, and it also invents a "CSV page" that the ticket never mentions. That's what a 248M model does with an instruction it was never fine-tuned on, and fine-tuning on pairs of tickets and their summaries is the usual next step, and the prompt below runs that comparison on your own input and target pairs.

Ask your AI coding tool

Fine-tune an encoder-decoder on my extraction task and compare it with a prompted LLM. I have pairs.jsonl with an input string and a target string. Fine-tune google/flan-t5-base with AutoModelForSeq2SeqLM and Seq2SeqTrainer on 80% of the pairs for 3 epochs, holding out the rest. Report exact-match accuracy and character-level edit distance on the held-out pairs, plus median latency per example on CPU. Then run the same held-out inputs through a local qwen2.5:3b via Ollama with a prompt containing three examples, and report the same three numbers. Print the ten held-out pairs where the two disagree most so I can read them.

Limits

What it can't do

  • It can't take a new instruction as well as a large decoder can. Flan-T5 follows short instructions it was tuned for, and anything outside that needs fine-tuning.
  • The input length is the encoder's limit. T5 and BART stop around 512 to 1,024 tokens, so long documents get chunked. T5Gemma 2's 128K is the exception.
  • The ecosystem moved on. Serving stacks, quantization formats and agent frameworks are built around decoder-only models, so an encoder-decoder deployment does more of its own plumbing.
  • Two stacks are more to fine-tune. You're training an encoder, a decoder and the cross-attention between them, where a decoder-only fine-tune touches one stack and usually only a LoRA adapter inside it.
  • Licenses vary more than they look. NLLB's checkpoints are non-commercial, and T5Gemma carries the Gemma terms, where T5 and Flan-T5 are Apache-2.0.
  • Small old checkpoints are small old models. A 2019 or 2022 model at a few hundred million parameters gets facts wrong and has no knowledge past its training data.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.