MLGuerrillaStart with M1 →
Architectures·16 min read·Updated 24 September 2026

Vision Transformer (ViT)

A Vision Transformer cuts an image into squares, treats each square as a token, and runs the same transformer a language model runs. It is the backbone inside CLIP, SigLIP, DINOv2 and every vision-language model on this index.

A Vision Transformer takes an image as input and returns a vector for every square it cut the image into, plus one extra vector standing for the whole thing. The architecture doing the work is the transformer, unchanged from the one built for translation.

Getting there took one idea. Cut the image into a grid of fixed squares, flatten each square, multiply it by one learned matrix to get a vector, and hand the resulting sequence to a transformer as though the squares were words. Dosovitskiy et al. (2020) put it in their title as an image being worth 16 by 16 words, and showed in the abstract that "a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks".

What it is

The picture becomes a sequence, and each patch's position is a learned vector

Running one photo through the original checkpoint shows every step of that, with the numbers each one produces.

A four-step figure titled "An image becomes 196 tokens, and a transformer reads them", from one photo through google/vit-base-patch16-224 at 86.6M parameters. Step one shows a 224 by 224 street photo overlaid with a 14 by 14 grid, giving 196 patches of 16 pixels, with one square outlined in terracotta. It explains that each patch is 768 pixel values flattened and multiplied by one learned matrix to give 768 numbers, that nothing convolutional happens and no patch sees any other yet. Three real rows are printed: patch 1 begins -0.17, -0.08, +0.06, +0.07, -0.01; patch 106, the terracotta square, begins -0.32, -0.03, -0.19, -0.08, +0.14; and patch 196 begins -0.19, -0.03, -0.09, +0.09, -0.01, each with 768 numbers. Step two, headed a learned position vector is added and it says where each patch sits, shows a 14 by 14 heatmap of how similar patch 106's position embedding is to every other patch's, with a bright centre and a visible cross along its row and column. Four measurements are given on a track running from -0.70 to +1.00: the patch itself at +1.00, the patch beside it at +0.80, ten patches away at +0.22, and the far corner at -0.16. It notes that the model taught itself the row and the column, which is what that cross is. Step three shows 197 tokens, a CLS chip outlined in terracotta followed by p1 through p5 and 191 more, passing through 12 layers in 266 milliseconds on a Mac CPU and producing traffic light at 0.156 of 1,000 classes. It explains that the CLS token belongs to no patch, is read out at the end, and is the only one the classifier looks at, and that the next four guesses were mailbag 0.061, street sign 0.051, trench coat 0.050 and cloak 0.034, because ImageNet has no class for a street. Step four, outlined in terracotta, shuffles the 196 patches and shows the scrambled image beside three bars: traffic light before at 0.156, traffic light after at 0.011, and umbrella after at 0.324. It notes that every patch is still there and every pixel unchanged, only the order moved, and that the model is now twice as confident about an umbrella as it ever was about the street. The caption reads: the patches carry the picture, and the position vectors carry where the picture was.
The three printed vectors are the real output of the patch embedding layer on this photo. Patch 106 is the square outlined on the image, so you can see which part of the street it came from.

The patch embedding is one matrix multiplication and nothing else, with no convolution and no pooling. The layout inside a patch survives, because every patch is flattened in the same pixel order, so the matrix can learn edge and texture patterns within those 16 by 16 pixels. The paper found its learned filters "resemble plausible basis functions for a low-dimensional representation of the fine structure within each patch". What the embedding can't tell is where a patch sits among the others, for example that patch 92 is directly above patch 106. That arrangement comes from the position vectors added in step two.

Those position vectors are learned, and the paper is explicit that they start out knowing nothing, saying "the position embeddings at initialization time carry no information about the 2D positions of the patches and all spatial relations between the patches have to be learned from scratch". What they look like afterwards is the heatmap: a patch is most similar to itself, then to its neighbours, with a cross along its own row and column. The model worked out the shape of a grid from photographs.

The shuffle in step four shows the model uses that arrangement. Every pixel is still present and only the order of the patches changed, and the answer went from a traffic light at 0.156 to an umbrella at 0.324. The drop shows that the model reads the arrangement of the patches, and a CNN given the same shuffled photo would likely lose its answer too.

Where it shows up

Where Vision Transformers show up

As the eyes of a vision-language model

Every VLM on this index has a ViT in front of the language model. The image-to-tokens step that article describes is this architecture, and the patches it produces become tokens in the prompt.

As the image tower of CLIP and SigLIP

CLIP and SigLIP are two towers, and the image one is a ViT. The most-downloaded Vision Transformer on Hugging Face is not sold as a Vision Transformer at all, it is openai/clip-vit-base-patch32.

As a general-purpose feature extractor

DINOv2 trains a ViT without labels and produces patch features that transfer to segmentation, depth and retrieval. It is the default backbone when you want image features and have no labels to fine-tune with.

Image classification, where it replaced the CNN

Give it a labelled dataset and a classification head and it does the job ResNets did, which is what the original paper set out to show.

Document retrieval over page images

ColPali runs a ViT over a rendered page and keeps one vector per patch. The patches are parts of a document rather than parts of a photograph, and the architecture does not care.

How it works

Patches, positions, and heads that behave like convolutions

Three steps before any attention happens

The image is resized to a fixed size, 224 by 224 in the checkpoint measured here. It is cut into a grid of squares, 14 by 14 of them at 16 pixels each. Each square's pixels are flattened into a list of 768 numbers and multiplied by one learned matrix, giving a vector of 768 numbers per patch.

A learned position vector is then added to each one, and a single extra vector called the CLS token is put at the front. That token belongs to no patch and exists so the model has a slot to accumulate a summary in, the same way BERT uses one.

From there it is an ordinary transformer. Every token attends to every other token at every layer, and after the last layer the CLS token is what the classifier reads.

Attention distance, which is where the difference from a CNN shows

A convolution in the first layer can only see a small square around each pixel, and its view grows one layer at a time. A transformer can look anywhere from the first layer. Whether it actually does is a question you can measure, and the paper measured it with something it calls attention distance, described as "analogous to receptive field size in CNNs".

A figure titled "Some heads look at one patch, and some look at the whole picture", plotting mean attention distance in pixels against layer for google/vit-base-patch16-224 on one photo, with one dot for each of the 12 heads in each of the 12 layers, 144 dots in total. A dashed terracotta line marks 112 pixels, half the width of the image. In layer 1 the dots are spread all the way from 0.0 pixels, circled and labelled as a head that reads only its own patch, up to 116.2 pixels, circled and labelled as being in the same layer. Layers 2 through 9 show a wide spread, with the lowest dots rising steadily and the highest staying near the dashed line. From layer 10 onward the dots cluster tightly above the dashed line, and by layer 12 the closest-reading head is at 114.6 pixels, so every head is looking across most of the photo. A panel below, headed why that is the whole point, says a convolution can only see a small square in its first layer and its view grows one layer at a time, that the paper calls attention distance analogous to receptive field size in CNNs, and that in layer 1 this model has a head that never leaves its own patch and a head that reaches 116 pixels across the image. It closes by noting that both of those were learned from data and neither was built into the architecture. The caption reads: locality is something a Vision Transformer can learn, and it is not something it is given.
This is the paper's own Figure 7 measurement, reproduced on one photograph. The spread in the left-hand column is the finding.

The first layer is the interesting one. One head in it has a mean attention distance of 0.0 pixels, meaning it reads its own patch and nothing else, which is the most local thing a head can do. Another head in the same layer reaches 116.2 pixels, which is more than half the width of the image. The paper found the same split and wrote that "some heads attend to most of the image already in the lowest layers" while others "have consistently small attention distances in the low layers".

By layer 12 the spread is gone and every head sits between 114.6 and 130.5 pixels. Depth pushes everything global.

It needs more data than a CNN, and that was the original problem

A convolution has locality and translation equivariance built into every layer, so it starts out already assuming useful things about images. A ViT assumes far less. Its feed-forward layers work on one patch at a time, and the paper says the two-dimensional layout "is used very sparingly", once when the image is cut into patches and again when position embeddings are resized for a new resolution. The attention layers have no notion of neighbours at all, which is why the paper says transformers "lack some of the inductive biases inherent to CNNs, such as translation equivariance and locality, and therefore do not generalize well when trained on insufficient amounts of data".

Trained on ImageNet alone, the first ViTs came in below ResNets of the same size. Pre-trained on the 14 million images of ImageNet-21k, ViT-L/16 reached 85.30% on ImageNet. The clear win came with Google's in-house JFT-300M dataset, where ViT-H/14 reached 88.55% against 87.54% for BiT-L, a ResNet pre-trained on the same data. The paper sums it up as "large scale training trumps inductive bias".

A later paper narrowed that gap from the training side. DeiT (Touvron et al., 2020) trained an 86M-parameter ViT on ImageNet alone with heavy augmentation and regularization, and reported 83.1% top-1 "with no external data". So a ViT on a mid-sized dataset can work, as long as the training recipe supplies what the architecture doesn't.

That is why almost nobody trains one from scratch. You start from a checkpoint someone pre-trained at that scale, and the BERT article makes the same argument for text.

Resolution and cost are the same dial

Doubling the side of the image quadruples the number of patches, and attention cost grows faster than linearly in the number of tokens. A 16-pixel patch on a 224-pixel image gives 196 tokens, and the same patch size on a 448-pixel image gives 784. The VLM article measures what that costs in a prompt, where a single photo through Qwen2.5-VL arrived as 3,588 tokens.

Fine-tuning at a higher resolution than the checkpoint was trained at means the position embeddings no longer line up, and they get interpolated onto the new grid. That step is standard and it is worth knowing it happened when accuracy moves for no visible reason.

Versions

The checkpoints you'll see, as of September 2026

Vision Transformers, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
openai/clip-vit-base-patch32March 2022Check the card22.07M
google/vit-base-patch16-224March 2022Apache-2.07.51M
facebook/dinov2-baseJuly 2023Apache-2.03.00M
google/vit-base-patch16-224-in21kMarch 2022Apache-2.01.85M
timm/vit_base_patch16_224.augreg2_in21k_ft_in1kDecember 2022Apache-2.00.45M
microsoft/resnet-50, the CNN it replacedMarch 2022Apache-2.00.56M

The top row is the point of the table. The most-downloaded Vision Transformer is the image tower of CLIP, pulled three times as often as the original ViT checkpoint, because most people reach for this architecture inside something else.

vit-base-patch16-224 is the checkpoint to read the code of and the one this article measured. The -in21k variant is the same model without the ImageNet-1k classification head, which is what you want when you are fine-tuning on your own labels.

dinov2-base is the one to start from for features. It is the same architecture trained without labels, and the DINO article covers what that changes.

Choosing

Choosing, and when a CNN is still the answer

Which vision backbone to reach for
The jobReach forWhy
Image features, with no labels of your ownDINOv2Self-supervised, and its patch features transfer without fine-tuning
Classifying into your own labelsvit-base-patch16-224-in21k plus a headPre-trained at scale, and the head is the only new part
Matching images against textCLIP or SigLIPTwo towers trained together, which is a different objective from this one
Feeding images to a language modelA VLMThe ViT is already inside it, with a projector on top
Searching documents as page imagesColPaliSame backbone, one vector kept per patch
A small labelled dataset and no pre-training budgetA CNN, or a pre-trained ViTA from-scratch ViT underperforms here, which the paper measured
Real-time detection on a deviceYOLOBuilt for latency, where this is built for transfer
Very high resolution inputsA hierarchical backbone such as Swin, or a CNNPatch count and attention cost both grow fast

Start from a pre-trained checkpoint in every case. In the original paper ViT pulled clearly ahead of a ResNet only with JFT-300M pre-training, and a checkpoint lets you inherit that pre-training without paying for it.

Try it

How to try it

Classifying one image is four lines, and the interesting part is what you can read out of the model while you do it.

classify.py — one image, one label
import torch
from PIL import Image
from transformers import ViTForImageClassification, ViTImageProcessor

mid = "google/vit-base-patch16-224"                  # 86.6M, Apache-2.0
proc = ViTImageProcessor.from_pretrained(mid)
model = ViTForImageClassification.from_pretrained(mid).eval()

px = proc(images=Image.open("photo.jpg"), return_tensors="pt").pixel_values
with torch.no_grad():
    probs = model(px).logits.softmax(-1)[0]
top = probs.topk(5)
for p, i in zip(top.values, top.indices):
    print(round(float(p), 3), model.config.id2label[int(i)])

Reading the position embeddings takes one more line, and it is the fastest way to convince yourself the grid was learned.

positions.py — what the model knows about where things are
pe = model.vit.embeddings.position_embeddings[0, 1:]   # drop CLS, 196 x 768
pe = torch.nn.functional.normalize(pe, dim=-1)
sim = (pe @ pe.T).detach()

probe = 7 * 14 + 7                        # the patch at row 7, column 7
grid = sim[probe].reshape(14, 14)
print(round(float(grid[7, 7]), 2),        # itself
      round(float(grid[7, 8]), 2),        # the patch beside it
      round(float(grid[0, 0]), 2))        # the far corner
Ask your AI coding tool

Show me what a Vision Transformer has actually learned about space, on my own images. Load google/vit-base-patch16-224 and run it on a folder of photos with output_attentions turned on. For each photo, report the top five predicted labels with probabilities, and compute the mean attention distance for every head in every layer by weighting each patch pair's attention by the distance between their centres in pixels, then plot those 144 numbers as a scatter of distance against layer. Separately, take the learned position embeddings, drop the CLS row, and print the cosine between one patch and every other as a 14 by 14 grid so I can see the row and column structure. Finally, shuffle the 196 patches of each photo in pixel space with a fixed seed, run it again, and report how far the original top label's probability fell.

Limits

What it can't do

  • It needs a lot of pre-training, or a heavy training recipe. With the original recipe on ImageNet alone it lands below a comparable ResNet, which the paper states plainly. DeiT closed that gap with augmentation and regularization, and most of what's good about it still arrives with the checkpoint.
  • Where each patch sits comes only from learned position vectors. Those vectors are tied to the grid the checkpoint was trained on, so changing input resolution means interpolating them onto a new grid, and accuracy can move when you do.
  • Cost grows fast with resolution. Patches scale with area and attention scales worse than linearly with tokens, which is why high-resolution work usually reaches for a hierarchical backbone.
  • The output is a label or a set of vectors, and not a location. Boxes and masks need a detection or segmentation head on top, and YOLO exists for the case where latency matters.
  • It only sees what fits in the input square. A 224-pixel resize throws away detail, so small text and fine structure are gone before the first layer. That is one of the reasons ColPali runs at page resolution.
  • This run is one photo on one checkpoint. It shows the mechanism and the attention spread, and the paper's own figure is the version measured across a dataset.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.