What it is
The input goes in whole, and the output comes out one token at a time
Both stacks are built from the same transformer blocks. The decoder differs in two places, a mask on its own attention and a second attention layer pointed at the encoder's output.

Splitting the work in half has two consequences.
The input is read once, in both directions, and the result is reused for every output token. A decoder-only model with a KV cache also reads its prompt once, in one pass called prefill, and reuses the stored attention keys and values at every step after that, so both shapes pass over the input once. The difference is in how the input is read and what each output step costs. The encoder lets every input token see every other one, where a decoder-only model reads its prompt left to right. And each output step in an encoder-decoder runs only the decoder stack, which can be much smaller than the encoder.
That second point is the reason the halves can be sized independently. Google's T5Gemma release pairs a 9B encoder with a 2B decoder for exactly this reason, because summarization needs a deep reading of a long input and writes a short output.
How it works
Span corruption is what trains both halves
T5's pretraining removes random spans from a sentence, replaces each with a numbered sentinel token, and asks the decoder to write the removed spans back in order. The encoder sees the damaged sentence, the decoder writes the missing pieces.

BART (Lewis et al., 2019) does the same job with a different damage function, shuffling sentences and replacing spans with a single mask token. Its paper describes the architecture as one that "can be seen as generalizing BERT (due to the bidirectional encoder), GPT (with the left-to-right decoder), and many other more recent pretraining schemes", which is the clearest one-line summary of where the shape sits.
Fine-tuning follows the same interface. Each example is an input string and a target string, the loss is the decoder's next-token loss over the target, and the class in transformers is AutoModelForSeq2SeqLM. Classification becomes a target of one word, summarization a target of one paragraph, and the training loop doesn't change between them.
Where it shows up
Where the shape is still the default
Speech to text
Whisper is an encoder-decoder whose encoder reads a spectrogram and whose decoder writes text. openai/whisper-large-v3 is downloaded about 4.7 million times a month, which makes it one of the most used encoder-decoders in the world.
Translation
Dedicated translation models stayed with the shape. facebook/nllb-200-distilled-600M covers 200 languages and gets about 1 million downloads a month, under a CC-BY-NC-4.0 license that rules out commercial use. Google's MADLAD-400 models are Apache-2.0 and cover more than 400 languages.
Summarization pipelines that already work
facebook/bart-large-cnn still pulls about 1.3 million downloads a month. It's a 2019 model fine-tuned on news summarization, and for the narrow job of turning an article into three sentences it's cheaper than an API call and needs no prompt.
Fine-tuned text-to-text jobs
Flan-T5 is the common starting point when the output is short text and you have examples. Extraction into a fixed format, normalization of messy strings, query rewriting for search. google/flan-t5-base at 248M parameters runs on a CPU.
Encoder-decoder LLMs
T5Gemma and T5Gemma 2 are the current attempt to bring the shape back for general-purpose work, and the versions section below covers what they are and what they cost you in licensing.
Versions
The models you'll see, as of September 2026
| Model | Released | What it does | License | Downloads a month |
|---|---|---|---|---|
google-t5/t5-base | October 2019 | The original text-to-text model | Apache-2.0 | 2.4M |
facebook/bart-large-cnn | 2019 | News summarization | MIT | 1.3M |
google/flan-t5-base | October 2022 | T5 after instruction tuning | Apache-2.0 | 1.4M |
openai/whisper-large-v3 | November 2023 | Speech to text | Apache-2.0 | 4.7M |
facebook/nllb-200-distilled-600M | July 2022 | Translation across 200 languages | CC-BY-NC-4.0 | 1.0M |
google/t5gemma-2-270m-270m | December 2025 | A current encoder-decoder LLM, multimodal | Gemma | 45K |
What T5Gemma changed
Google published Encoder-Decoder Gemma (Zhang et al., April 2025), which builds an encoder-decoder by copying the weights of a pretrained decoder-only model into both halves, switching the encoder's self-attention from causal to bidirectional, and continuing training with a span-corruption objective. The paper reports that under a similar inference budget the result reaches "comparable (often better) pretraining performance but substantially better finetuning performance than their decoder-only counterpart".
T5Gemma shipped on July 9, 2025, built from Gemma 2. T5Gemma 2 followed on December 18, 2025, built from Gemma 3, with three sizes from 270M-270M up to 4B-4B, a 128K context window, support for over 140 languages and image input. Its two architecture changes both save parameters. The embeddings are tied across the encoder and the decoder, and the decoder's self-attention and cross-attention are merged into one layer. Both releases carry the Gemma license, which has use restrictions that Apache-2.0 models don't.
Choosing
Choosing between this and a decoder
| The job | Reach for | Why |
|---|---|---|
| Speech to text | Whisper, or a hosted speech model | The audio encoder is the whole point |
| Translation you run yourself | An encoder-decoder translation model | Small, fast, and trained on nothing else besides translation |
| Summarizing at high volume with a fixed output shape | A fine-tuned encoder-decoder | A small model runs on your own hardware, and a small decoder keeps each output step cheap |
| Chat, agents, tool calls, long generation | A decoder-only LLM | Every ecosystem, serving stack and technique assumes it |
| Classifying into fixed labels | An encoder with a head | A label is not text worth generating |
| Quality with no engineering effort | A frontier decoder-only model through an API | The largest general-purpose models are decoder-only |
The size of the model and what it was trained on usually count for more than the shape, and the summary in the next section shows what a 248M model does with an instruction it wasn't fine-tuned for.
Try it
How to try it
AutoModelForSeq2SeqLM loads every text model in the table above the same way. Whisper takes audio, so it has its own class. This one turns a support ticket into a one-sentence summary.
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_id = "google/flan-t5-base" # 248M, Apache-2.0
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
prompt = ("Summarize this support ticket in one short sentence: "
"Hi, since yesterday's update the export button on the reports "
"page does nothing when I click it. I need the CSV for a meeting "
"on Friday.")
ids = tok(prompt, return_tensors="pt") # 45 tokens in
out = model.generate(**ids, max_new_tokens=40)
print(tok.decode(out[0], skip_special_tokens=True))
# -> Exporting reports does not work on the CSV page.The summary gets the problem right, and it also invents a "CSV page" that the ticket never mentions. That's what a 248M model does with an instruction it was never fine-tuned on, and fine-tuning on pairs of tickets and their summaries is the usual next step, and the prompt below runs that comparison on your own input and target pairs.
Fine-tune an encoder-decoder on my extraction task and compare it with a prompted LLM. I have pairs.jsonl with an input string and a target string. Fine-tune google/flan-t5-base with AutoModelForSeq2SeqLM and Seq2SeqTrainer on 80% of the pairs for 3 epochs, holding out the rest. Report exact-match accuracy and character-level edit distance on the held-out pairs, plus median latency per example on CPU. Then run the same held-out inputs through a local qwen2.5:3b via Ollama with a prompt containing three examples, and report the same three numbers. Print the ten held-out pairs where the two disagree most so I can read them.
Limits
What it can't do
- It can't take a new instruction as well as a large decoder can. Flan-T5 follows short instructions it was tuned for, and anything outside that needs fine-tuning.
- The input length is the encoder's limit. T5 and BART stop around 512 to 1,024 tokens, so long documents get chunked. T5Gemma 2's 128K is the exception.
- The ecosystem moved on. Serving stacks, quantization formats and agent frameworks are built around decoder-only models, so an encoder-decoder deployment does more of its own plumbing.
- Two stacks are more to fine-tune. You're training an encoder, a decoder and the cross-attention between them, where a decoder-only fine-tune touches one stack and usually only a LoRA adapter inside it.
- Licenses vary more than they look. NLLB's checkpoints are non-commercial, and T5Gemma carries the Gemma terms, where T5 and Flan-T5 are Apache-2.0.
- Small old checkpoints are small old models. A 2019 or 2022 model at a few hundred million parameters gets facts wrong and has no knowledge past its training data.
Go deeper
Raffel et al. (2019): Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) · Lewis et al. (2019): BART, Denoising Sequence-to-Sequence Pre-training · Xue et al. (2020): mT5, a massively multilingual pre-trained text-to-text transformer · Chung et al. (2022): Scaling Instruction-Finetuned Language Models (Flan-T5) · Radford et al. (2022): Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) · Zhang et al. (2025): Encoder-Decoder Gemma, Improving the Quality-Efficiency Trade-Off via Adaptation · Google: T5Gemma, a new collection of encoder-decoder Gemma models
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
