MLGuerrillaStart with M1 →
Vision-language·14 min read·Updated 24 September 2026

CLIP

CLIP is a pair of encoders that put images and text into one vector space, so you can compare a photo to a sentence. It is what lets you classify or search images with words you type.

CLIP, short for Contrastive Language-Image Pre-training, came out of OpenAI in 2021. It runs two encoders, one for images and one for text, and trains them so that a photo and its caption land close together as vectors. Once both live in the same space, comparing them is a cosine similarity, which is how a model that was never trained on your categories can still sort your photos into them.

The paper's headline result was that this matched "the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on". This page covers what the scores mean and how the training works. The last section covers where the approach falls over.

What it is

Images and captions become vectors in one space, and you compare them

Feeding an image through the image encoder gives a vector. Feeding a sentence through the text encoder gives a vector of the same size. The cosine similarity between them is a single number, and the whole method follows from that.

To classify with no training, you write one short caption per class and embed them once, then take the caption whose vector is closest to the image's. To search a library, you embed every image once, then embed the query text and return the nearest images.

Scoring five images against five captions with the smallest CLIP model puts real numbers on that.

A figure titled "CLIP scores every image against every caption", showing a 5 by 5 grid of cosine similarities. The rows are thumbnails: the whole street photo, a crop of a stroller, a crop of a bicycle, a crop of a car and a crop of a striped hat. The columns are the captions: a baby stroller, a bicycle, a small black car, a striped knitted hat, and a busy city crosswalk. Cells are shaded terracotta by value. The highest value in each row is the matching pair: the whole photo scores 0.294 for the crosswalk caption, the stroller crop 0.312 for the stroller caption, the bicycle crop 0.289 for the bicycle caption, the car crop 0.281 for the car caption and the hat crop 0.298 for the hat caption. Every value in the grid sits between 0.185 and 0.312. A note says the right pairing is only the highest value in its row, and the caption says softmax turns the bicycle row's 0.026 gap into 93% against 7%, so treat these scores as a ranking.
The whole photo scores highest for the crosswalk caption and only 0.213 for the stroller, even though the stroller is in it. One vector has to stand for the whole image.

The right answer is always the highest score in its row, so ranking works. The numbers themselves sit in a narrow band, between 0.185 and 0.312, with no natural cutoff between "matches" and "doesn't". Libraries hide this by multiplying the scores by a scale learned during training and running a softmax over them, which turns the bicycle row's gap of 0.026 into 93% against 7%. Those percentages only compare the captions you supplied.

That narrow band has a cause, and you can see it by looking at where the vectors sit.

A figure titled "Images and captions share a space, and they sit in different parts of it", built from the same five images and five captions embedded with CLIP ViT-B/32 and placed on the one axis that separates them. On the left, ten labeled rows each carry a dot at its position along the first principal component, which accounts for 47.7% of the variance. The five images, the street photo, the stroller crop, the bicycle crop, the car crop and the hat crop, all sit as terracotta dots on the left side. The five captions, a busy city crosswalk, a baby stroller, a bicycle, a small black car and a striped knitted hat, all sit as dark dots on the right side. A shaded band between the two groups is labeled "the gap", and a note says only the dot's position is data and the rows are spread out so the labels fit. On the right, four measured averages over the same ten vectors are shown as bars: between two images 0.657, between two captions 0.655, image to its own caption 0.295 highlighted in terracotta, and any image to any caption 0.229. A note says two images look far more alike to CLIP than an image and its own caption do. The caption reads: this is why a raw image-caption score never reaches 1, and why 0.3 is a good match rather than a weak one.
The two encoders were never trained to put an image on top of its caption. They were trained to make the right pair score higher than the wrong ones, and that is a weaker requirement.

Two images average 0.657 with each other and two captions average 0.655, while an image and its own caption average 0.295. The image vectors and the text vectors occupy separate regions, and every cross-modal score you ever read is a comparison across that divide. Liang et al. (2022) named this the modality gap and traced it to how these models are initialized and trained, showing that the two modalities end up "embedded at arm's length in their shared representation".

Where it shows up

It shows up wherever text and images have to meet

Searching images with words

Embed a photo library once and keep the vectors in a vector database. A search for "a red bicycle leaning on a wall" embeds the query and returns the nearest images, with no tags or filenames involved.

Classifying without training data

Moderation triage, product categories, "is this a screenshot or a photo". You write the candidate captions, and you can change them tomorrow without retraining anything. Accuracy lands below a model fine-tuned on your labels, which is the trade you're making.

Filtering and deduplicating training data

CLIP similarity is a cheap way to check whether an image and its caption belong together. The LAION-5B dataset was built this way, dropping English pairs whose CLIP cosine similarity fell below 0.28.

Inside image generators

The text encoder that turns a prompt into vectors for an image generator often is CLIP's. Stable Diffusion 1.5's model card describes it as using "a fixed, pretrained text encoder (CLIP ViT-L/14)". The diffusion models article covers what happens next with those vectors.

Scoring generated images

CLIPScore (Hessel et al., 2021) uses the same similarity to judge how well a caption matches an image, and image generators are often compared this way. It inherits every weakness below, so it's a rough signal.

How it works

Training pulls matching pairs together and pushes the rest apart

CLIP was trained on 400 million image-text pairs collected from the internet. Each training step takes a batch of N pairs and computes all N × N similarities between the images and the texts in that batch. The paper describes the objective as training "to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the incorrect pairings".

So the model never learns class labels. It learns to rank the right caption above the other captions that happened to be in the batch, which is why bigger batches make the task harder and the features better. A learned temperature scales the similarities before the loss. In the code it's stored as a multiplier, logit_scale, and the paper capped it at 100 to keep training stable. That same multiplier later makes those confident-looking percentages.

The two encoders are ordinary architectures. The image side is a ResNet or a vision transformer, and the text side is a transformer over the caption's tokens, which makes both of them encoders. Each ends in a projection to the shared space, and the vectors are normalized to unit length, which is what makes their dot product a cosine.

Zero-shot classification then reuses the text encoder as a classifier builder. You embed "A photo of a {label}." for every label, and those vectors act as the weights of a classifier you never trained. The wording changes the result more than you'd expect. The paper found that adding "A photo of a" to a bare label improved ImageNet accuracy by 1.3%, and that prompt engineering with ensembles of several phrasings "boost zero-shot classification performance by almost 5 points on average across 36 datasets".

Versions

The versions you'll see, as of September 2026

CLIP and its open re-implementations
ModelFromNotes
CLIPOpenAI, January 2021The original, MIT-licensed code and weights, from ViT-B/32 up to ViT-L/14
OpenCLIPLAION and collaboratorsAn open re-implementation with many checkpoints trained on open datasets
MetaCLIPMetaPublished on Hugging Face as facebook/metaclip-*, trained with a data-curation recipe of its own
SigLIPGoogleA CLIP-style model trained with a sigmoid loss in place of the batch softmax, covered in its own entry on this index

The interface is the same across all of them, so swapping checkpoints is a one-line change, and the one to use is whichever scores best on your own images.

Choosing

Pick CLIP when words have to meet images, and a neighbor when they don't

Which model to reach for
The jobReach forWhy
Search or classify images using words you typeCLIP or SigLIPText and images share one space, so a caption is a query
Find visually similar images, with no text involvedDINOv2 or DINOv3Self-supervised features capture appearance more finely than caption-trained ones
Draw boxes around what a phrase namesGrounding DINOCLIP scores whole images, and detection needs a detector
Answer a question about an image in wordsA vision-language modelCLIP has no text decoder, so it can't write anything
Classify into fixed classes you have labels forA fine-tuned classifier, which can start from CLIP featuresTraining on your labels beats zero-shot on a fixed task

The CLIP-versus-DINO choice comes up often, and the figure above shows why it isn't obvious. CLIP's features are shaped by what captions mention, so two images described the same way look similar to it, even if they don't look alike. DINO's are shaped by appearance alone. Text search needs the first, and near-duplicate detection is usually better with the second.

Try it

How to try it

This is the code behind the figure, cut down to one image and a few captions.

zero_shot.py — score one image against candidate captions
import torch
from PIL import Image
from transformers import CLIPModel, CLIPProcessor

model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32").eval()
proc = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")

captions = ["a baby stroller", "a bicycle", "a small black car"]
inputs = proc(text=captions, images=Image.open("crop.jpg").convert("RGB"),
              return_tensors="pt", padding=True)

with torch.no_grad():
    out = model(**inputs)

cosine = (out.image_embeds @ out.text_embeds.T)[0]  # unit vectors, so cosines
scale = model.logit_scale.exp().item()              # 100.0 for this checkpoint
probs = out.logits_per_image.softmax(dim=-1)[0]     # cosine * scale, softmaxed
for caption, c, p in zip(captions, cosine, probs):
    print(f"{caption:20} cos {c:.3f}   {p * 100:5.1f}%")

The raw cosine comes from the two embeddings directly, and the scale is read from the checkpoint. OpenAI's weights sit at the cap of exactly 100, so dividing logits_per_image by 100 happens to work for them. Two fine-tuned checkpoints checked for this page read 99.79 and 99.96, and SigLIP adds a learned bias after the scale, so a hard-coded 100 gives slightly wrong cosines on the first two and wrong ones on SigLIP.

Ask your AI coding tool

Build a text search over a folder of images with CLIP. Use openai/clip-vit-base-patch32 through transformers to embed every image once and store the normalized vectors with their paths in a FAISS inner-product index that persists to disk. Then give me a search function that takes a text query, embeds it with the same model, and returns the top 10 paths with their cosine similarities. Add an evaluation mode that takes a CSV of queries and the filename that should rank first, and reports top-1 and top-5 accuracy, and print the distribution of cosine similarities for correct and incorrect matches so I can see how much they overlap.

Limits

What it can't do

  • Its scores only rank candidates. Every similarity in the figure sits between 0.185 and 0.312, so a fixed threshold like "0.3 means a match" doesn't transfer between image types or caption styles. Calibrate against your own labeled examples, or compare candidates and take the best.
  • It needs a list of candidates. With no captions to compare against, there's nothing to rank, so it can't tell you what's in an image the way a captioning model would.
  • One vector stands for the whole image. In the figure, the full street photo scores 0.213 for "a baby stroller" even though a stroller is in it, because the vector describes the scene. Small objects need crops or a detector.
  • Word order and relations get lost. Yuksekgonul et al. (2022) built the ARO benchmark and found CLIP-style models behave "like bags-of-words", so "a person on a horse" and "a horse on a person" can score almost the same.
  • Text printed in the image pulls the score. In a test for this page, writing the word "bicycle" across the stroller crop moved the "a bicycle" caption from 0.0% to 25.9% and its cosine from 0.219 to 0.273. Goh et al. (2021) showed stronger versions of this effect, which they called a typographic attack.
  • It's weak at counting and fine-grained distinctions, like the exact model of a car or the number of objects in a photo, which the paper's own limitations section covers.
  • It carries the biases of its training data. The original weights come from web images and captions collected up to 2021, with the paper devoting a section to the biases that survive in the model.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.