What it is
One space is real, and one ranked list is not
The natural thing to build once you hear "one space" is a single index holding everything you own, searched with one query. The figure is that index, built from this site with one model, google/siglip-base-patch16-224, over 287 text passages and 13 images. It is one checkpoint on one small library, so it shows what can go wrong with a single ranked list, and it can't tell you how large the problem is for a different model.

Text sits close to text at an average cosine of 0.621. Images sit close to images at 0.512. Text sits at 0.006 from images, which is as close to unrelated as two directions get.
So for this model the index works as built and the ranked list is useless. Across the eight questions, all 48 of the top-six slots went to a text passage, and the best-scoring image on any question reached 0.105 against that question's best text hit of 0.740. No image was ever going to appear.
This separation has a name, the modality gap. Liang et al. (2022) show that images and text sit "at arm's length" in the shared space of models such as CLIP, and trace it to two causes. A freshly initialized encoder already puts its outputs in a narrow cone, so the two towers start apart, and contrastive training then keeps them "separate by a certain distance, which is influenced by the temperature parameter in the loss function". CLIP-style training pulls a photo towards its own caption and pushes it away from other captions, and nothing in that objective asks a photo to land near text in general.
How wide the gap is varies by model and by training recipe, and the hosted models later on this page are trained to make mixed lists work. So the SigLIP run is a reason to measure your own model, and it doesn't settle what every multimodal model does.
Where it shows up
Where multimodal embeddings show up
Searching an image library with words
The original use and still the best one. Embed a million photos once, type a sentence, and the nearest photos come back. The query and the library are different kinds of thing, so the modality gap never gets in the way.
Product search over photographs
A shopper types "striped linen shirt" and the catalogue photos rank against it, with no one having tagged them. Marqo/marqo-fashionSigLIP is a SigLIP fine-tuned on fashion data, and its card reports an average recall of 0.231 against FashionCLIP2.0's 0.163 on text-to-image retrieval across six public fashion datasets.
Finding duplicate or near-duplicate images
Image against image averaged 0.512 here, over a set that includes one photo and five crops of it. Deduplication and reverse image search are comparisons within one modality, which is the case these models handle best.
One index over documents that contain both
This is where the newest hosted models are aimed. Cohere's docs describe embed-v4.0 as taking "Text, Images, Mixed texts/images (i.e. PDFs)" at 256, 512, 1,024 or 1,536 dimensions with a 128k-token context, and Google's docs say gemini-embedding-2 "is the first multimodal embedding model in the Gemini API" and that it "maps text, images, video, audio, and documents into a unified embedding space".
Zero-shot classification
Embed your label names as text, embed the image, and take the nearest label. The CLIP article covers what that gets you and where the scores stop meaning anything.
How it works
Two towers, the gap between them, and what to do about it
The training puts pairs together and nothing else
Both CLIP and SigLIP train two encoders at once, one for images and one for text, on hundreds of millions of image-and-caption pairs. The loss raises the similarity of each real pair and lowers it for mismatched ones. The bi-encoders article covers why two separate towers is what makes a searchable index possible at all.
Only the pairs are ever compared. The objective never says a photograph should be near the average sentence, or that two sentences about the same subject should be near each other, so neither happens reliably.
The text tower is not a text embedding model
That last point has a practical edge. On the eight questions, SigLIP's text tower found the right article 4 of 8 times, which is the same score bge-small got, though on a different set of questions. It is not useless and it is also not what it was built for, and its input caps at 64 tokens, so anything longer than a caption gets truncated.
The ColPali paper measured this properly at scale. On ViDoRe, its document retrieval benchmark, Jina-CLIP scored 17.7 average nDCG@5 and Nomic-vision scored 12.9, against 65.1 for plain BM25 and 67.0 for a text embedding model. Contrastive image-text models are far behind ordinary text retrieval on documents, and that gap is the reason the ColPali design exists.
The scores are not on a scale you can read
The best image hit for "a busy city crosswalk with people walking" was the street photo at 0.123, and it was the right answer. That number is not a 12 percent confidence. SigLIP trains with a sigmoid loss that has its own learned temperature and bias, so the raw cosine sits in a narrow band with no calibration behind it, which the CLIP article measures in detail.
Asking four caption-like questions of the 13 images got two right. "Blue sky above buildings" found the sky crop, "a busy city crosswalk with people walking" found the street photo, and "a bicycle on the road" returned a diagram from the diffusion article instead of the crop with a bicycle in it.
Two indexes, and what rank fusion does to them
The usual fix is one index per modality, each searched on its own, which gives you a text list and an image list that are each internally consistent. Merging them is where people reach for reciprocal rank fusion, which gives every item 1 / (k + rank) from each list it appears in and sorts by the total, with k usually 60.
RRF was built for lists that rank the same documents, like BM25 and a dense retriever over one corpus, where an item both lists agree on collects two votes and rises. A text list and an image list never share an item, so every item gets exactly one vote, and the vote depends only on its position. Take "train a classifier on my own labels" from the figure.
| Item | Its own score | Rank in its list | RRF score | Merged position |
|---|---|---|---|---|
| Best text passage | 0.761 | 1 | 1/61 = 0.0164 | 1 or 2 |
| Best image | 0.069 | 1 | 1/61 = 0.0164 | 1 or 2 |
| Second text passage | 0.756 | 2 | 1/62 = 0.0161 | 3 or 4 |
| Second image | below 0.069 | 2 | 1/62 = 0.0161 | 3 or 4 |
The best image ties the best passage, even though it was the best of thirteen weak image matches for a question about text classification. The merge alternates text and image all the way down, and which one goes first in a tie is decided by the sort. So RRF over disjoint lists is an ordering policy you chose, one from each list in turn, and it says nothing about whether the image is relevant.
Once you see it as a policy, you can choose one on purpose.
- Show the two lists as separate result groups. Most image-and-text search interfaces do this, and it makes no claim about which modality fits better.
- Gate each list with its own threshold. Fit a cut-off for image scores and another for text scores on labeled queries, and only show an image group when an image clears its own cut-off. Check first that your model separates matches from non-matches at all. In this run the correct images for the caption-like queries scored from 0.123 for the street photo down to below 0.026 for the bicycle crop, which missed the top three for "a bicycle on the road". The irrelevant best image for the classifier question scored 0.069, inside that range, so no image threshold would have separated them for SigLIP base on this library.
- Weight the lists by the query. A query that describes something visual gets more image slots, and a routing rule or a small classifier decides which queries those are.
The other route is a model trained to make one list work. That is what the hosted multimodal models below describe, and it is worth measuring on your own data, with the same three numbers as the figure, before you trust one list.
Versions
The models you'll see, as of September 2026
| Model | Released | License | Downloads a month |
|---|---|---|---|
openai/clip-vit-large-patch14 | March 2022 | Check the card | 8.05M |
google/siglip2-so400m-patch14-384 | February 2025 | Apache-2.0 | 1.02M |
Marqo/marqo-fashionSigLIP | August 2024 | Apache-2.0 | 402K |
nomic-ai/nomic-embed-vision-v1.5 | June 2024 | Apache-2.0 | 191K |
jinaai/jina-clip-v2 | October 2024 | CC-BY-NC-4.0 | 107K |
CLIP from 2022 is still pulled eight times as often as anything newer, which is the pattern the BERT article found in encoders and the ColBERT article found in retrievers. SigLIP 2 is the one to start from on quality. jina-clip-v2 is worth knowing about because its card claims the thing the rest of this category gives up, saying its text encoder "can serve as an effective multilingual long-context dense retriever" on a par with the company's own text embedding model. Its licence is CC-BY-NC-4.0, which rules out commercial use.
nomic-embed-vision-v1.5 pairs with nomic-embed-text-v1.5 and shares its space, which is the open version of the one-index idea. Both models need trust_remote_code, and on the machine behind this article the vision half would not load under transformers 5.17 at all, which is worth budgeting time for.
| Model | Inputs the docs list | Dimensions |
|---|---|---|
Cohere embed-v4.0 | Text, images, and mixed text and images such as PDFs | 256, 512, 1024 or 1536 |
Google gemini-embedding-2 | Text, images, video, audio and documents | 3072 by default, truncatable |
Voyage voyage-multimodal-3.5 | Interleaved text and visual data, including screenshots of PDFs, slides, tables, figures and videos | 1024 by default, plus 256, 512 and 2048 |
Those are vendor descriptions of what the models accept, and none of them is a measurement of how well one ranked list behaves across two modalities. Ask for that number, or measure it.
Choosing
Choosing, and when one index is the wrong shape
| The job | Reach for | Why |
|---|---|---|
| Searching photos with typed sentences | SigLIP 2, or CLIP | The query and the library are different kinds of thing, which is the case this is built for |
| Photos in one specific domain | A fine-tuned model such as marqo-fashionSigLIP | A domain model beats a general one by a wide margin on its own catalogue |
| Finding duplicate images | Any of them | Image against image is the strongest comparison these models make |
| Searching text properly | A text embedding model | A contrastive text tower caps at caption length and was never trained for it |
| One search box over text and images | Two indexes, shown as groups or gated by per-modality thresholds | The score scales do not meet, and fusing by rank only interleaves the two lists |
| Searching PDFs and slides as they look | ColPali | 81.3 against 17.7 for a contrastive model on the same benchmark |
| Deciding whether one image matches one caption | SigLIP with a fitted threshold | A raw cosine here has no threshold you can carry between datasets |
| Text, images, audio and video in one call | A hosted multimodal model | Nothing open covers that spread today |
Build the two-index version first, with separate result groups or a threshold per modality. It takes an afternoon, the failure modes are visible, and it gives you the baseline to measure a hosted multimodal model against when you decide whether one index is worth paying for.
Try it
How to try it
Embedding both kinds of input is one call each, and the point of running it is to see the two score ranges for yourself.
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
mid = "google/siglip-base-patch16-224"
proc = AutoProcessor.from_pretrained(mid)
model = AutoModel.from_pretrained(mid).eval()
norm = torch.nn.functional.normalize
with torch.no_grad():
px = proc(images=[Image.open(p) for p in paths], return_tensors="pt")
img = norm(model.get_image_features(**px).pooler_output, dim=-1)
tx = proc(text=captions, padding="max_length", truncation=True,
max_length=64, return_tensors="pt")
txt = norm(model.get_text_features(**tx).pooler_output, dim=-1)
print("text to image:", float((txt @ img.T).mean()))
print("image to image:", float((img @ img.T).mean()))The RRF function is the same one the BM25 article uses. Run it on your two lists and look at the output, because with no shared ids it will alternate text and image whatever the scores are.
def rrf(*rankings, k=60):
votes = {}
for ranking in rankings:
for rank, item in enumerate(ranking, start=1):
votes[item] = votes.get(item, 0) + 1 / (k + rank)
return sorted(votes, key=votes.get, reverse=True)
results = rrf(text_hits, image_hits) # no shared ids, so this alternatesShow me whether one index over text and images actually works on my own data. Take a folder of images and a folder of Markdown files, split the text into passages of at least 25 words, and embed both with google/siglip-base-patch16-224. Report the average cosine between every pair of text vectors, every pair of image vectors, and every text against every image, so I can see the modality gap as three numbers. Then take a list of my own questions, search the combined index once per question, and print the top ten results with the modality of each and its score, plus the best score any text item reached and the best any image reached. Finally build the two-index version, where text and images are searched separately. Label which images are relevant for each question, then report the lowest score of a relevant image and the highest score of an irrelevant one, so I can see whether a per-modality threshold can separate them. Show the two lists as separate groups, and print what reciprocal rank fusion at k of 60 would produce so I can see it alternate the two lists.
Limits
What it can't do
- With SigLIP base, text and images did not compete in one ranked list. The measured gap in this run was 0.621 within text, 0.512 within images and 0.006 between them, and no image reached the top three for any of the eight questions. Other models have gaps of different widths, so measure yours.
- Rank fusion over two modalities is an ordering policy. With no shared items, RRF alternates the lists by position and ties the best image with the best passage whatever their scores.
- The text tower is not a text search engine. It caps at 64 tokens on SigLIP, and on document retrieval the ColPali paper puts contrastive models at 17.7 and 12.9 nDCG@5 against 65.1 for BM25.
- The scores have no calibration. A correct image match scored 0.123 in this run. There is no threshold that carries from one dataset to the next, which the SigLIP article goes into.
- General models are weak on specific catalogues. A fine-tuned domain model is usually the difference between a demo and a product, and that means labelled pairs of your own.
- Open coverage stops at text and images. Video, audio and full documents in one space are hosted-only today, and those are vendor claims until you measure them.
- The run here is 287 passages and 13 images. It is enough to show the gap, which is large and consistent, and nowhere near enough to rank models. Run it on your own library.
Go deeper
Radford et al. (2021): CLIP, learning transferable visual models from natural language supervision · Zhai et al. (2023): SigLIP, sigmoid loss for language image pre-training · Liang et al. (2022): Mind the Gap, understanding the modality gap in multi-modal contrastive representation learning · Faysse et al. (2024): ColPali, whose benchmark table measures these models on document retrieval · Cohere's embedding docs, where embed-v4.0's inputs and dimensions are listed
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
