MLGuerrillaStart with M1 →
Vision-language·12 min read·Updated 24 September 2026

SigLIP

SigLIP is a CLIP-style pair of image and text encoders trained with a sigmoid loss, so every image-caption pair gets a score that stands on its own. It is also the vision encoder inside many open VLMs.

SigLIP is CLIP with one part swapped. It has the same two encoders and the same shared space, and it changes the training loss from a softmax over the whole batch to a sigmoid on each image-caption pair. Google released it in 2023, and SigLIP 2 followed in February 2025.

That one change has two effects. Training needs less coordination across the batch, which makes it cheaper at a given quality. And the score for one image and one caption no longer depends on which other captions you scored alongside it. That second effect is easy to overread, because a score that doesn't depend on the other candidates is still an uncalibrated number until you have checked it against labels from your own data.

What it is

Each pair gets its own score, so the numbers stand alone

CLIP produces two different numbers, and it helps to keep them apart. The raw one is a cosine similarity between the image vector and the caption vector, and it depends only on that image and that caption. The percentages most code prints come from a softmax over the captions you supplied, and those do depend on the list, because the softmax makes them add up to 100%. SigLIP passes each pair's scaled similarity through a sigmoid, so its number also depends only on the pair, and it lands between 0 and 1 without any normalization across candidates.

What SigLIP's number doesn't do is tell you how likely the caption is to be right. The paper supports the difference in the loss, and says nothing about the scores being calibrated on your images, so treat a sigmoid score like any other raw score and check it against labeled examples before you put a threshold on it.

Running the same five images and five captions through both models shows the difference in the numbers.

A figure titled "SigLIP scores each pair on its own, so a row doesn't have to add up", with two 5 by 5 grids of percentages side by side and thumbnails down the left: the whole street photo, a stroller, a bicycle, a car and a striped hat. The left grid, labeled CLIP, softmax across the row, has one large value per row and a row total of 100% for every row: 99.9 for the crosswalk caption on the photo, 99.9 stroller, 93.1 bicycle with 6.6 on stroller, 99.6 car and 99.1 hat. The right grid, labeled SigLIP, a sigmoid on each pair, has the same winners with much smaller values and row totals that vary: 65.8 for crosswalk on the photo with a 66% row total, 10.4 for the stroller with 10%, 35.5 for the bicycle with 35%, 96.5 for the car with 96% and 97.9 for the hat with 98%. A note says CLIP's rows always add to 100% whatever the captions are, and SigLIP's rows add to whatever fits. A closing line says every winner is the correct caption, SigLIP still gave the stroller 10%, and its threshold should be fitted on labels.
The stroller crop is a correct match in both models. CLIP's softmax reports 99.9% because it beat four other captions, and SigLIP reports 10.4% for the same correct pair, so a low sigmoid score here says nothing about the match being wrong.

Look at the stroller row. CLIP reports 99.9% for "a baby stroller" because that caption beat the four others. SigLIP reports 10.4% for the same pair, and the pair is correct. The 10.4% is where this checkpoint's learned scale and bias happened to put a correct stroller crop, while the correct car and hat pairs came out at 96.5% and 97.9%, so the gap between them says nothing about which match the model is more sure of.

Removing the right caption from the list shows which of the numbers depend on the list. These come from the same run's recorded similarities, with the stroller crop scored against the four remaining captions.

The stroller crop, with and without "a baby stroller" in the candidate list
NumberFive captions"a baby stroller" removed
CLIP raw cosine for "a busy city crosswalk"0.2360.236
CLIP softmax for "a busy city crosswalk"0.0%82.3%
SigLIP sigmoid for "a busy city crosswalk"0.0%0.0%

CLIP's softmax puts 82.3% on a crosswalk caption for a photo of a stroller, because it has to spread 100% across whatever is left. Its raw cosine doesn't move, and neither does SigLIP's score, since both are functions of the one pair.

That independence is what makes SigLIP convenient for filtering. You can score one image against one caption with no competitors, and put a threshold on the score once you have tuned it on labeled pairs. CLIP's raw cosine supports the same one-pair question, and the difference is that SigLIP was trained on exactly that yes-or-no question, where CLIP's training only ever compared a caption against the rest of its batch.

Where it shows up

It shows up in search, in filtering, and inside other models

Deciding whether a caption and an image go together

Content checks and dataset filtering both come down to scoring one pair, and so does validating alt text. The sigmoid output means the same thing when there's only one candidate, and the threshold still has to come from your own labeled examples, because a correct pair can score 10% and another correct pair 98%.

Search and zero-shot classification

Everything the CLIP article covers works the same way here, with the same code shape. SigLIP 2 also trains on many languages, so queries don't have to be in English.

As the eyes of a vision-language model

Many open vision-language models use a SigLIP encoder to turn an image into tokens that a language model can read. PaliGemma (Beyer et al., 2024) is built on "the SigLIP-So400m vision encoder and the Gemma-2B language model", and So400m is the shape-optimized 400-million-parameter vision tower that usually comes with it.

How it works

The loss is the whole difference

Both models embed images and captions into one space and score a pair by a dot product of normalized vectors, scaled by a learned temperature. The training objective is where they part.

  • CLIP takes a batch of N pairs, builds the N × N matrix of similarities, and runs a softmax over each row and column. The right pair has to beat every other caption in that batch, so every score depends on which other examples were drawn.
  • SigLIP applies a sigmoid to each of those N × N similarities on its own and asks a yes/no question of each, with the true pairs labeled yes and every other combination labeled no. The paper describes the loss as operating "solely on image-text pairs", with no need for "a global view of the pairwise similarities for normalization".

Because most pairs in a batch are negatives, the loss would drift toward "no" for everything, so SigLIP learns a bias term that shifts the scores back. Both the temperature and the bias are ordinary learned parameters you can read out of the checkpoint. In the base SigLIP model used for the figure they came out at 117.3 and −12.9, and in SigLIP 2 at 112.7 and −16.8, which is part of why the same pairs score lower in the newer model.

The practical payoff is in training. Removing the need for every device to see the whole batch's similarities makes large-batch training cheaper, and the paper reports the method "performing better at smaller batch sizes" as well. Its headline example trained a model to "84.5% ImageNet zero-shot accuracy in two days" on four TPU chips, using a frozen pretrained image tower.

SigLIP 2 keeps the loss and adds more training signal, including captioning-based pretraining and self-supervised losses, along with multilingual data. Its paper reports that it beats the original "at all model scales in core capabilities, including zero-shot classification, image-text retrieval".

Versions

The versions you'll see, as of September 2026

SigLIP checkpoints on Hugging Face
FamilyReleasedTypical checkpointsNotes
SigLIP2023google/siglip-base-patch16-224, google/siglip-so400m-patch14-384The original sigmoid-loss models, English
SigLIP 2February 2025google/siglip2-base-patch16-256, google/siglip2-so400m-patch14-384, google/siglip2-giant-opt-patch16-384Multilingual, better dense features, the default choice today
SigLIP 2 NaFlexFebruary 2025google/siglip2-base-patch16-naflexHandles varying input sizes and aspect ratios, where the others take one fixed square

The so400m models are the ones VLMs like PaliGemma build on, and the base models are small enough to run on a laptop CPU, which is where the numbers in the figure came from.

Choosing

Pick SigLIP when a score has to stand alone

Which model to reach for
The jobReach forWhy
Score one image against one caption, with a thresholdSigLIP or SigLIP 2Each pair is scored on its own, and the threshold comes from your labeled pairs
Rank a fixed list of captions for an imageEither modelBoth rank the same way, and CLIP's softmax is convenient for a forced choice
A vision encoder to feed a language modelSigLIP 2, usually so400mIt's the encoder PaliGemma and other open VLMs are built on
Non-English queriesSigLIP 2Its training data is multilingual
Visual similarity with no text involvedDINOv2 or DINOv3Self-supervised features capture appearance without caption bias
Boxes around what a phrase namesGrounding DINOThese models score whole images and return no boxes

Try it

How to try it

The API matches CLIP's, with two differences worth copying exactly. The processor needs padding="max_length", and the scores come from a sigmoid.

pair_score.py — score each caption against an image on its own
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor

model_id = "google/siglip2-base-patch16-224"
model = AutoModel.from_pretrained(model_id).eval()
proc = AutoProcessor.from_pretrained(model_id)

captions = ["a baby stroller", "a bicycle", "a small black car"]
inputs = proc(text=captions, images=Image.open("crop.jpg").convert("RGB"),
              padding="max_length", return_tensors="pt")

with torch.no_grad():
    out = model(**inputs)

for caption, p in zip(captions, torch.sigmoid(out.logits_per_image)[0]):
    print(f"{caption:20} {p * 100:5.2f}%")     # each score stands on its own
Ask your AI coding tool

Help me pick a SigLIP threshold for filtering image-caption pairs. I have pairs.csv with an image path and a caption, plus a human label of match or no match. Write a script that scores every pair with google/siglip2-base-patch16-224, then reports precision and recall at thresholds from 0.05 to 0.95 in steps of 0.05, and prints the threshold with the best F1. Plot the score distributions for matches and non-matches on one chart so I can see the overlap. Then rerun the whole thing with google/siglip-base-patch16-224 and show both results side by side, since the two models use different learned biases.

Limits

What it can't do

  • Its scores aren't comparable across models. The same pairs scored 10.4% with SigLIP and 15.5% with SigLIP 2 for the stroller, because each model learns its own bias, so a threshold has to be retuned whenever you change checkpoints.
  • A low score on its own doesn't establish a mismatch. Until a threshold has been fit on labeled pairs that look like your data, a low score can belong to a correct caption. Even past a validated threshold, the score gives no clue which part of the caption failed or whether a different wording would fit.
  • One vector still stands for the whole image, so small objects in a busy photo go missing the same way they do with CLIP.
  • It inherits CLIP's blind spots, including weak handling of word order and relations, and sensitivity to text printed inside the image.
  • Its scores aren't probabilities of a match. The figure's correct pairs range from 10% to 98%, so a fixed cut-off like 50% would have rejected two of the five. Fit the threshold, or a calibration step, on labeled pairs from your own data.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.