What it is
A transformer updates every token using the tokens around it, all at once
Before 2017, the standard model for text was a recurrent neural network, which reads one token at a time and carries a running summary forward. For translation, an encoder network read the whole source sentence into one fixed-length vector, and a decoder network wrote the translation from that vector. Bahdanau et al. (2014) argued that "the use of a fixed-length vector is a bottleneck", and added attention, a way for the decoder to look back at every source word and weight the ones it needs for the word it's about to write.
The transformer kept attention and dropped the recurrent part. In the paper's words, it's "based solely on attention mechanisms, dispensing with recurrence and convolutions entirely". With no token-by-token reading, every position can be computed at the same time, which suits GPUs. The paper's base model trained in 12 hours on 8 GPUs.
The input doesn't have to be text. Anything you can cut into a sequence of pieces and turn into vectors works.
- Text goes in as tokens, which are words or pieces of words.
- An image goes in as square patches, which is the idea in An Image is Worth 16x16 Words (Dosovitskiy et al., 2020), the paper that introduced the Vision Transformer.
- Audio goes in as short slices of a spectrogram, which is what Whisper (Radford et al., 2022) does.
- The compressed image that a diffusion model denoises goes in as patches, which is what diffusion transformers (Peebles and Xie, 2022) do.
How attention works
Attention lets each token pull in information from the tokens it needs
Each token gets a query, a key and a value
Every token starts as a vector. Attention multiplies that vector by three learned matrices to get three new vectors for the token:
- The query describes what this token is looking for in the other tokens.
- The key describes what this token offers, so other tokens' queries can match against it.
- The value is the information this token passes on when another token attends to it.
To update one token, attention takes its query and computes a dot product with the key of every token it's allowed to see. A high dot product means the query and key point the same way, which is a good match. The scores are divided by the square root of the key length, and softmax turns them into weights that add up to 1. The token's new vector is the weighted sum of all the value vectors. The paper writes the whole thing as one line, softmax(QKᵀ / √d_k) V, where Q, K and V hold the queries, keys and values of every token as rows.
The division by √d_k is there because dot products of long vectors get large. The authors suspected that large scores push softmax into a region "where it has extremely small gradients", which makes training stall, and scaling the scores down keeps it out of that region.

The scores in the figure are made up, and the weights are their exact softmax. In a trained model the scores come from the learned query and key matrices.
A decoder masks the tokens that come later
In an LLM, a token can only attend to itself and the tokens before it, because when the model is generating text the later tokens haven't been written yet. The paper implements this by "masking out (setting to −∞)" the scores for later positions before softmax, so they get a weight of exactly 0. That's why "twice" is greyed out in the figure. An encoder, like the one inside BERT, has no mask, so every token can attend to every other token.
Attention for all tokens is a few matrix multiplications
The figure shows one token, and a real layer does every token at the same time. Stacking all the queries into a matrix Q and all the keys into K, the product QKᵀ gives every query-key score in one multiplication, and GPUs are built for exactly that kind of work. The code at the end of this page is the whole computation in a dozen lines of numpy.
Several heads give each token more than one weighted average
One head, meaning one trio of query, key and value matrices, produces one set of weights for each token, so everything that token pulls in gets blended into a single weighted average. If a token needs one fact from one earlier token and a different fact from another, one head has to split its weight between them and mix the two values together. The paper describes the problem in one line, "With a single attention head, averaging inhibits this." So a transformer runs several attention heads side by side, each with its own smaller matrices, and joins their outputs. The 2017 base model used 8 heads, each working with vectors of 64 numbers, which adds up to its model width of 512.
The heads aren't given different jobs. Training decides what each one matches on, and heads often end up doing overlapping work. Michel et al. (2019) found that "a large percentage of attention heads can be removed at test time without significantly impacting performance", and that some layers could be cut down to one head, while a few specific heads could not be removed without a large drop. So a single head in a trained model can't be relied on to track one clean relationship, even when a heatmap of it looks tidy.
Newer models change how heads share work. Qwen2.5-3B has 16 query heads and only 2 sets of keys and values, so groups of 8 query heads share one key and one value. This is grouped-query attention, from Ainslie et al. (2023), and it makes generation faster, because the model has fewer keys and values to store and read for every token. That store is the KV cache, which the LLM article covers.
Word order has to be added in
Attention on its own ignores order. The weighted sum comes out the same if you shuffle the tokens, so "the customer charged the store" and "the store charged the customer" would look alike. The transformer has to be told where each token sits.
The 2017 paper added "positional encodings" to the token embeddings, a fixed pattern of sine and cosine waves that's different for every position. Many open LLMs, Qwen2.5 among them, now use rotary position embedding, or RoPE, from Su et al. (2021). It rotates each query and key by an angle that depends on its position, so the dot product of two tokens depends on how far apart they are.
The layer
A transformer layer is attention plus a feed-forward network, with a shortcut around each
Attention is one of two parts in each layer. The second is a feed-forward network, two or three matrix multiplications applied to each token on its own. It holds about two-thirds of a transformer's parameters, and Geva et al. (2020) found that these layers work like key-value memories, matching patterns in the text and pushing up the words likely to come next. Around each part is a residual connection, a shortcut that adds the part's input to its output, so each part only has to learn a change to the token's vector. Normalization keeps the numbers in a stable range.

The details of the design have changed a good deal since 2017.
| Transformer base (2017) | Qwen2.5-3B (2024) | |
|---|---|---|
| Shape | Encoder and decoder, 6 layers each | Decoder only, 36 layers |
| Vector size per token | 512 | 2,048 |
| Attention heads | 8, each with its own keys and values | 16 query heads sharing 2 key-value heads |
| Feed-forward network | 512 → 2,048 → 512, with ReLU | 2,048 → 11,008 → 2,048, with a gated SiLU |
| Position | Sine and cosine encodings added to embeddings | RoPE applied to queries and keys |
| Normalization | LayerNorm after each part, LayerNorm(x + Sublayer(x)) | RMSNorm before each part |
Xiong et al. (2020) showed that with normalization placed inside the residual branch, "the gradients are well-behaved at initialization", and that these models could train without the slow warm-up the original design needed, which is one reason newer models normalize before each part.
Three ways to arrange the stack
The same layer gets arranged in three ways, and the arrangement decides what a model is good at:
- Encoder only models, like BERT, let every token see every other token. They read and classify text, and they can't generate any.
- Decoder only models, like GPT and today's LLMs, keep the causal mask. They generate text one token at a time.
- Encoder-decoder models, like T5 and Whisper, let the decoder attend to the encoder's output through a second attention block, called cross-attention. They fit jobs that turn one sequence into another, like audio into a transcript.
The encoder, decoder and encoder-decoder article goes through all three in detail.
The encoder-only, decoder-only and encoder-decoder entry on this index compares them in detail.
Cost
Attention's cost grows with the square of the input length
Every token's query is compared with every token's key, so a sequence of n tokens needs n × n scores in every head of every layer. The paper's Table 1 lists the cost of a self-attention layer as O(n² · d), where d is the vector size. Doubling the input quadruples the number of scores.
Two pieces of engineering keep this manageable in practice:
- FlashAttention, from Dao et al. (2022), computes exactly the same result while moving far less data between the GPU's memory levels, which makes attention much faster on long inputs. It still computes all n × n scores, just faster.
- The KV cache stores each token's keys and values during generation, so a new token only computes its own query against the stored keys.
Even so, filling a million-token context window on every call is expensive, and that cost is why other architectures that grow linearly with length, like state-space models, are an active area.
Choosing
Pick a transformer for almost anything with enough data, and know the neighbors
For language, and for most vision and audio work at scale, the transformer is the default, and the question is usually which transformer-based model to use. The neighbors on this index matter in a few situations.
| Situation | Consider | Why |
|---|---|---|
| Long sequences where attention's n² cost dominates | A state-space model or a hybrid | Mamba (Gu and Dao, 2023) reports "linear scaling in sequence length" |
| Images on a small device, or a small labeled image dataset | A convolutional neural network | Convolutions build in the assumption that nearby pixels belong together, so they need less data |
| A small streaming time series on limited hardware | An RNN or LSTM | It keeps a fixed-size state and processes one step at a time |
| A large model that has to stay cheap per token | A transformer with mixture-of-experts layers | Each token runs through only a few of the feed-forward networks |
The data point about images comes from the ViT paper itself. Trained on a mid-sized dataset like ImageNet, the authors found that ViT scored "a few percentage points below ResNets of comparable size", because "Transformers lack some of the inductive biases inherent to CNNs". It pulled ahead once it was pretrained on much larger datasets. So the less data you have, the more a convolutional network's built-in assumptions help. Each of these architectures has its own entry on this index.
Try it
See attention in code
This is scaled dot-product attention in numpy, with the causal mask, run on five random token vectors. Every line maps to a step above.
import numpy as np
def attention(Q, K, V, causal=True):
"""Scaled dot-product attention from Vaswani et al. (2017), equation 1."""
d_k = Q.shape[-1]
scores = Q @ K.T / np.sqrt(d_k) # how well each query matches each key
if causal: # a token may only look backwards
n = scores.shape[0]
scores = np.where(np.tril(np.ones((n, n))) == 1, scores, -np.inf)
weights = np.exp(scores - scores.max(axis=-1, keepdims=True))
weights /= weights.sum(axis=-1, keepdims=True) # softmax, row by row
return weights @ V, weights
tokens = ["The", "customer", "was", "charged", "twice"]
rng = np.random.default_rng(0)
X = rng.normal(size=(5, 8)) # 5 tokens, 8 numbers each
Wq, Wk, Wv = (rng.normal(size=(8, 4)) for _ in range(3))
out, w = attention(X @ Wq, X @ Wk, X @ Wv)
print(np.round(w, 2))The printed matrix has one row per token and one column per token it attends to. Everything above the diagonal is 0 because of the mask, and each row adds up to 1. The weights here are random, since the matrices haven't been trained.
Write a Python script that loads a small pretrained transformer with Hugging Face transformers (use Qwen/Qwen2.5-0.5B-Instruct), runs the sentence "The customer was charged twice for the same order" through it with output_attentions=True, and plots the attention weights as heatmaps with matplotlib. Show one heatmap per head for layers 1, 12 and 24, label both axes with the tokens, and save the plots as PNG files. Print the model's number of layers, query heads and key-value heads from its config first, so I can match them against what I see.
Limits
What it can't do
- It can't take in more than its context window. The cost of attention grows with the square of the input length, and every model is trained and served with a fixed limit on how many tokens it reads at once.
- It doesn't know word order unless position is added. Attention treats its input as an unordered collection, and the model depends entirely on the positional encoding or RoPE to know which token came first.
- Its attention weights aren't an explanation of its decisions. Jain and Wallace (2019) found that "standard attention modules do not provide meaningful explanations and should not be treated as though they do", partly because different attention patterns can give the same prediction.
- It keeps no memory between inputs. Each run starts from only the tokens it's given, so anything it should remember has to be put back into the input.
- It needs a lot of data to beat models with built-in assumptions. The ViT result above is the clearest example, where a transformer trained on ImageNet alone trailed comparable convolutional networks.
Go deeper
Vaswani et al. (2017): Attention Is All You Need · Bahdanau et al. (2014): Neural Machine Translation by Jointly Learning to Align and Translate · Dosovitskiy et al. (2020): An Image is Worth 16x16 Words (ViT) · Su et al. (2021): RoFormer, Rotary Position Embedding · Ainslie et al. (2023): GQA, grouped-query attention · Dao et al. (2022): FlashAttention · Michel et al. (2019): Are Sixteen Heads Really Better than One? · Jain and Wallace (2019): Attention is not Explanation
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
