What it is
Each pair gets its own score, so the numbers stand alone
CLIP produces two different numbers, and it helps to keep them apart. The raw one is a cosine similarity between the image vector and the caption vector, and it depends only on that image and that caption. The percentages most code prints come from a softmax over the captions you supplied, and those do depend on the list, because the softmax makes them add up to 100%. SigLIP passes each pair's scaled similarity through a sigmoid, so its number also depends only on the pair, and it lands between 0 and 1 without any normalization across candidates.
What SigLIP's number doesn't do is tell you how likely the caption is to be right. The paper supports the difference in the loss, and says nothing about the scores being calibrated on your images, so treat a sigmoid score like any other raw score and check it against labeled examples before you put a threshold on it.
Running the same five images and five captions through both models shows the difference in the numbers.

Look at the stroller row. CLIP reports 99.9% for "a baby stroller" because that caption beat the four others. SigLIP reports 10.4% for the same pair, and the pair is correct. The 10.4% is where this checkpoint's learned scale and bias happened to put a correct stroller crop, while the correct car and hat pairs came out at 96.5% and 97.9%, so the gap between them says nothing about which match the model is more sure of.
Removing the right caption from the list shows which of the numbers depend on the list. These come from the same run's recorded similarities, with the stroller crop scored against the four remaining captions.
| Number | Five captions | "a baby stroller" removed |
|---|---|---|
| CLIP raw cosine for "a busy city crosswalk" | 0.236 | 0.236 |
| CLIP softmax for "a busy city crosswalk" | 0.0% | 82.3% |
| SigLIP sigmoid for "a busy city crosswalk" | 0.0% | 0.0% |
CLIP's softmax puts 82.3% on a crosswalk caption for a photo of a stroller, because it has to spread 100% across whatever is left. Its raw cosine doesn't move, and neither does SigLIP's score, since both are functions of the one pair.
That independence is what makes SigLIP convenient for filtering. You can score one image against one caption with no competitors, and put a threshold on the score once you have tuned it on labeled pairs. CLIP's raw cosine supports the same one-pair question, and the difference is that SigLIP was trained on exactly that yes-or-no question, where CLIP's training only ever compared a caption against the rest of its batch.
Where it shows up
It shows up in search, in filtering, and inside other models
Deciding whether a caption and an image go together
Content checks and dataset filtering both come down to scoring one pair, and so does validating alt text. The sigmoid output means the same thing when there's only one candidate, and the threshold still has to come from your own labeled examples, because a correct pair can score 10% and another correct pair 98%.
Search and zero-shot classification
Everything the CLIP article covers works the same way here, with the same code shape. SigLIP 2 also trains on many languages, so queries don't have to be in English.
As the eyes of a vision-language model
Many open vision-language models use a SigLIP encoder to turn an image into tokens that a language model can read. PaliGemma (Beyer et al., 2024) is built on "the SigLIP-So400m vision encoder and the Gemma-2B language model", and So400m is the shape-optimized 400-million-parameter vision tower that usually comes with it.
How it works
The loss is the whole difference
Both models embed images and captions into one space and score a pair by a dot product of normalized vectors, scaled by a learned temperature. The training objective is where they part.
- CLIP takes a batch of N pairs, builds the N × N matrix of similarities, and runs a softmax over each row and column. The right pair has to beat every other caption in that batch, so every score depends on which other examples were drawn.
- SigLIP applies a sigmoid to each of those N × N similarities on its own and asks a yes/no question of each, with the true pairs labeled yes and every other combination labeled no. The paper describes the loss as operating "solely on image-text pairs", with no need for "a global view of the pairwise similarities for normalization".
Because most pairs in a batch are negatives, the loss would drift toward "no" for everything, so SigLIP learns a bias term that shifts the scores back. Both the temperature and the bias are ordinary learned parameters you can read out of the checkpoint. In the base SigLIP model used for the figure they came out at 117.3 and −12.9, and in SigLIP 2 at 112.7 and −16.8, which is part of why the same pairs score lower in the newer model.
The practical payoff is in training. Removing the need for every device to see the whole batch's similarities makes large-batch training cheaper, and the paper reports the method "performing better at smaller batch sizes" as well. Its headline example trained a model to "84.5% ImageNet zero-shot accuracy in two days" on four TPU chips, using a frozen pretrained image tower.
SigLIP 2 keeps the loss and adds more training signal, including captioning-based pretraining and self-supervised losses, along with multilingual data. Its paper reports that it beats the original "at all model scales in core capabilities, including zero-shot classification, image-text retrieval".
Versions
The versions you'll see, as of September 2026
| Family | Released | Typical checkpoints | Notes |
|---|---|---|---|
| SigLIP | 2023 | google/siglip-base-patch16-224, google/siglip-so400m-patch14-384 | The original sigmoid-loss models, English |
| SigLIP 2 | February 2025 | google/siglip2-base-patch16-256, google/siglip2-so400m-patch14-384, google/siglip2-giant-opt-patch16-384 | Multilingual, better dense features, the default choice today |
| SigLIP 2 NaFlex | February 2025 | google/siglip2-base-patch16-naflex | Handles varying input sizes and aspect ratios, where the others take one fixed square |
The so400m models are the ones VLMs like PaliGemma build on, and the base models are small enough to run on a laptop CPU, which is where the numbers in the figure came from.
Choosing
Pick SigLIP when a score has to stand alone
| The job | Reach for | Why |
|---|---|---|
| Score one image against one caption, with a threshold | SigLIP or SigLIP 2 | Each pair is scored on its own, and the threshold comes from your labeled pairs |
| Rank a fixed list of captions for an image | Either model | Both rank the same way, and CLIP's softmax is convenient for a forced choice |
| A vision encoder to feed a language model | SigLIP 2, usually so400m | It's the encoder PaliGemma and other open VLMs are built on |
| Non-English queries | SigLIP 2 | Its training data is multilingual |
| Visual similarity with no text involved | DINOv2 or DINOv3 | Self-supervised features capture appearance without caption bias |
| Boxes around what a phrase names | Grounding DINO | These models score whole images and return no boxes |
Try it
How to try it
The API matches CLIP's, with two differences worth copying exactly. The processor needs padding="max_length", and the scores come from a sigmoid.
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
model_id = "google/siglip2-base-patch16-224"
model = AutoModel.from_pretrained(model_id).eval()
proc = AutoProcessor.from_pretrained(model_id)
captions = ["a baby stroller", "a bicycle", "a small black car"]
inputs = proc(text=captions, images=Image.open("crop.jpg").convert("RGB"),
padding="max_length", return_tensors="pt")
with torch.no_grad():
out = model(**inputs)
for caption, p in zip(captions, torch.sigmoid(out.logits_per_image)[0]):
print(f"{caption:20} {p * 100:5.2f}%") # each score stands on its ownHelp me pick a SigLIP threshold for filtering image-caption pairs. I have pairs.csv with an image path and a caption, plus a human label of match or no match. Write a script that scores every pair with google/siglip2-base-patch16-224, then reports precision and recall at thresholds from 0.05 to 0.95 in steps of 0.05, and prints the threshold with the best F1. Plot the score distributions for matches and non-matches on one chart so I can see the overlap. Then rerun the whole thing with google/siglip-base-patch16-224 and show both results side by side, since the two models use different learned biases.
Limits
What it can't do
- Its scores aren't comparable across models. The same pairs scored 10.4% with SigLIP and 15.5% with SigLIP 2 for the stroller, because each model learns its own bias, so a threshold has to be retuned whenever you change checkpoints.
- A low score on its own doesn't establish a mismatch. Until a threshold has been fit on labeled pairs that look like your data, a low score can belong to a correct caption. Even past a validated threshold, the score gives no clue which part of the caption failed or whether a different wording would fit.
- One vector still stands for the whole image, so small objects in a busy photo go missing the same way they do with CLIP.
- It inherits CLIP's blind spots, including weak handling of word order and relations, and sensitivity to text printed inside the image.
- Its scores aren't probabilities of a match. The figure's correct pairs range from 10% to 98%, so a fixed cut-off like 50% would have rejected two of the five. Fit the threshold, or a calibration step, on labeled pairs from your own data.
Go deeper
Zhai et al. (2023): Sigmoid Loss for Language Image Pre-Training (SigLIP) · Tschannen et al. (2025): SigLIP 2 · Beyer et al. (2024): PaliGemma, a versatile 3B VLM for transfer · Radford et al. (2021): CLIP, the model SigLIP changes one loss in
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
