What it is
Every passage becomes a vector once, and a query becomes one more
The two halves of the work happen at different times. Embedding your documents is a batch job you run once and repeat only when the content changes. Embedding the query happens on every search, and it is one model call over a short piece of text.
Comparing is the cheap part. Both vectors get scaled to length 1, and the comparison is a dot product, which for unit-length vectors is the cosine of the angle between them. A modern vector database does that against millions of stored vectors in milliseconds, using an index that skips most of the comparisons.

Look at what the top hit is. The model put a YOLO passage about fine-tuning on labelled images above the BERT passage that answers the question, because both are about training on your own labels and the model has no way to know which kind of labels were meant. That near miss is the normal failure of this stage, and the rerankers article measures how much of it a second stage recovers.
Where it shows up
Where embeddings show up
Retrieval for RAG
Retrieval-augmented generation puts passages from your own documents into an LLM's input so it answers from them. Embeddings are how the right passages get found. The LLM article covers what the model does with them once they arrive.
Search that understands a phrasing you didn't anticipate
Keyword search needs the user to type a word that appears in the document. Embedding search needs them to mean the same thing, which covers synonyms, paraphrases and questions asked from the wrong angle. Most production search runs both and merges the results, because keyword search is still better at names, error codes and part numbers.
Finding near-duplicates and clustering
Two support tickets describing the same bug land close together, so a threshold over pairwise similarity groups them. That threshold has to be chosen on pairs from your own tickets that someone has marked as duplicate or not, with the exact model you run, and the section on scores further down shows why a number borrowed from elsewhere won't hold. The same operation across a whole document collection gives you clusters you never defined.
Classification without a trained classifier
Embed your labelled examples once, then classify a new text by the label of its nearest neighbours. This works well enough on small label sets to be worth trying before you fine-tune anything, and the BERT article covers what fine-tuning buys you over it.
Deduplicating training data
Dropping near-identical examples before training is a standard cleanup step, and embeddings are how you find them at scale.
How it works
How the model produces one vector, and what the number of dimensions costs
The architecture is an encoder with a pooling step
Most embedding models are encoders, used in the bi-encoder arrangement. An encoder gives back one vector for every token, so a 25-token passage comes out as 25 vectors, and a search index has room for one. Pooling is the step that closes that gap.

Two poolings cover most of the models you'll meet. Some take the vector at the [CLS] position, which is a slot the tokenizer adds at the front of every input and which belongs to no word, and some average all the token vectors together. bge-small-en-v1.5 is a CLS model, E5 and GTE are averaging models, and the card for each one says so.
The two give different vectors for the same text. On the passage in the figure they sit at a cosine of 0.94, and running the eight queries behind this article both ways brought back a different top passage for three of them, with accuracy at four of eight either way. A test that small cannot tell you which pooling is better. The model card can, because the model was trained with one of them and the other is a readout nobody optimised. Use the one the card names, for your documents and your queries alike.
The training is what makes the vectors useful. Reimers and Gurevych (2019) built Sentence-BERT by running the same encoder over two sentences and training so that related pairs come out close, which is the contrastive recipe CLIP uses across images and text. Every model on this page is trained some version of that way.
Many models need a prefix on the query
Instruction-tuned embedding models are trained to treat a query and a document differently, and they expect you to say which one you are handing them. bge-small-en-v1.5 wants queries prefixed with "Represent this sentence for searching relevant passages:", and Qwen3-Embedding wants an instruction line before the query. Skipping the prefix costs accuracy without any error message, and it is the most common way a first embedding pipeline underperforms.
The dimension number is a storage and speed decision
A vector of 1,024 numbers in 32-bit floats is 4 KB, so a million passages is 4 GB of index before any overhead. Cutting the vector in half halves that and roughly halves the comparison cost.
You can cut it because of Matryoshka Representation Learning (Kusupati et al., 2022), which trains a model so that the first half of a vector is itself a usable vector, and the first quarter after that. Google's docs say gemini-embedding-2 is trained this way, defaults to 3,072 numbers, and can be truncated to 768 or 1,536 with little loss. Cohere's embed-v4.0 offers 256, 512, 1,024 and 1,536, and Voyage's models offer 256 through 2,048.
Truncation only works on a model trained for it. Qwen3-Embedding's card lists output sizes from 32 up to the full length, and the hosted models above document their sizes, while a model trained without Matryoshka, such as bge-small-en-v1.5, has no reason to keep its meaning in the first numbers of the vector. Check the model card before cutting, and scale each shortened vector back to length 1 before comparing, since dropping numbers changes its length.
On the corpus behind the figures, truncating Qwen3-Embedding-0.6B from its full 1,024 numbers down to 128 returned the same top result for all eight queries. At 64 numbers, one of the eight changed. Eight queries can show a large loss and cannot show a small one, and the same top result does not mean the same ordering below it, so treat this as a reason to test 128 on your own questions and not as proof that 128 is free.
Long passages are cut off without a warning
The other kind of truncation happens on the way in. Every embedding model has a maximum input length, 512 tokens for bge-small-en-v1.5, and the tokenizer call in the code further down passes truncation=True, so anything past that point is dropped before the model sees it. No error is raised and the vector looks normal. It simply stands for the first 512 tokens of the passage.
The passages behind the figures were short enough that nothing was cut. On real documents, count tokens per chunk once, and make the chunk size smaller than the model's limit, including the query prefix on the query side.
A cosine is a ranking, and a threshold has to be fitted
The scores an embedding model returns look like confidence and behave nothing like it. The figure below scores every pair of the 287 passages on this site.

Two unrelated passages average 0.600, and the highest unrelated pair on this corpus scored 0.866. The eight real queries found their best answer somewhere between 0.654 and 0.824. A cutoff at 0.8 would have thrown away most of the correct answers and kept an unrelated pair.
So use the score to sort, take the top few, and decide relevance with something else. That something else is usually a reranker, which reads the query and one passage together and scores that pair.
A cosine cutoff can still work in a narrower setting. With one model, one kind of document and a few hundred pairs labelled as related or not, you can pick the cutoff that gives the precision you need and measure how much recall it costs. That number belongs to that model and that data. Unrelated passages average 0.600 under bge-small-en-v1.5 here, and another model puts its unrelated pairs somewhere else entirely, so the cutoff has to be fitted again whenever the model, the prefix or the kind of document changes.
Versions
The models you'll see, as of September 2026
| Model | Released | Dimensions | License | Downloads a month |
|---|---|---|---|---|
BAAI/bge-small-en-v1.5 | September 2023 | 384 | MIT | 64.3M |
BAAI/bge-m3 | January 2024 | 1024 | MIT | 37.3M |
BAAI/bge-large-en-v1.5 | September 2023 | 1024 | MIT | 10.9M |
Qwen/Qwen3-Embedding-0.6B | June 2025 | 1024 | Apache-2.0 | 8.7M |
intfloat/multilingual-e5-large | June 2023 | 1024 | MIT | 7.2M |
google/embeddinggemma-300m | July 2025 | 768 | Gemma | 2.7M |
Qwen/Qwen3-Embedding-8B | June 2025 | 4096 | Apache-2.0 | 2.6M |
Qwen/Qwen3-VL-Embedding-8B | January 2026 | check the card | check the card | 1.2M |
The download column tells the same story the BERT article tells. A model from September 2023, bge-small-en-v1.5, is pulled 64.3 million times a month, which is more than every model released in 2026 on that list combined. A 33M-parameter encoder that runs on a CPU keeps being the right answer.
| Model | Input | Dimensions | Context |
|---|---|---|---|
Cohere embed-v4.0 | Text, images, and mixed documents such as PDFs | 256, 512, 1024 or 1536 | 128k tokens |
Google gemini-embedding-2 | Text, images, video, audio and documents | 3072 by default, truncatable | See the API docs |
Google gemini-embedding-001 | Text | 3072 by default, truncatable | See the API docs |
Voyage voyage-3.5 and voyage-3.5-lite | Text | 1024 by default, plus 256, 512, 2048 | 32,000 tokens |
Two of those take more than text. Google describes gemini-embedding-2 as "the first multimodal embedding model in the Gemini API", putting text, images, video, audio and documents into one space, and Cohere's embed-v4.0 takes PDFs with their layout intact. One index across several kinds of content is the thing that changed in this category over the past year.
Watch for behaviour changes when you upgrade. Google's docs note that gemini-embedding-001 returns one embedding per string, while Embeddings 2 "produces a single aggregated embedding for multiple inputs", which is a different function with the same name.
Choosing
Which model to pick, and when an embedding is the wrong tool
| The job | Reach for | Why |
|---|---|---|
| A first retrieval system in English | bge-small-en-v1.5 or bge-large-en-v1.5 | Small, free, CPU-friendly, and the most tested thing in the category |
| Retrieval across many languages | bge-m3 or a hosted multilingual model | Trained for it, where an English model degrades quietly |
| The best accuracy you can self-host | A current Qwen3-Embedding size | It measured 6 of 8 against 4 of 8 for bge-small on the corpus behind these figures |
| Searching PDFs and screenshots without OCR | Cohere embed-v4.0, or the colpali approach | They embed the page as an image, so layout survives |
| One index over text and images | gemini-embedding-2 or a multimodal embedding model | One space means one query reaches every kind of content |
| Deciding whether one text matches one other text | A reranker, or SigLIP for image pairs | A cosine cutoff holds only for the model and task it was fitted on with labelled pairs, and a reranker scores the pair directly |
| Exact matching on names, codes or IDs | Keyword search such as BM25 | An embedding blurs exactly the detail those depend on |
| Writing an answer from what you found | An LLM | Embeddings retrieve, and they produce no text |
On the corpus behind these figures, bge-small got 4 of 8 queries right and Qwen3-Embedding-0.6B got 6 of 8, while taking 1,022 milliseconds per passage against 14. An 18-times bigger model bought two queries, and it made indexing 71 times slower. The reranker on the same corpus also bought two queries, from 4 of 8 to 6 of 8, without touching the index.
Those timings come from one setup. Both models ran in 32-bit floats on a Mac CPU, bge-small in batches of 32 passages and Qwen3 in batches of 16, over passages that were all well under 512 tokens. A GPU, a different batch size or longer passages change both numbers, and the ratio between them is the part that carries over. Two queries out of eight is also a small enough difference that a different set of eight questions could shrink it or widen it.
Try it
How to try it
Embedding a folder of text takes about fifteen lines. This is the code behind the first figure, cut down.
import torch
from transformers import AutoModel, AutoTokenizer
model_id = "BAAI/bge-small-en-v1.5" # 33M, MIT, runs on a CPU
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id).eval()
def embed(texts, prefix=""):
batch = tok([prefix + t for t in texts], padding=True,
truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
out = model(**batch).last_hidden_state[:, 0] # bge pools the CLS token
return torch.nn.functional.normalize(out, dim=-1)
index = embed(passages) # once, then store it
q = embed(["how do I train a classifier on my own labels"],
prefix="Represent this sentence for searching relevant passages: ")
scores = (q @ index.T)[0]
for i in scores.topk(3).indices:
print(round(float(scores[i]), 3), passages[i][:90])Keep the vectors in a vector database once the corpus outgrows memory. The interface is the same, and what you gain is an index that finds the nearest vectors without comparing against all of them.
Build a retrieval evaluation over my own documents so I can choose an embedding model with numbers. Read every Markdown file under a folder I pass in, split it into passages of at least 25 words, and embed them with BAAI/bge-small-en-v1.5, BAAI/bge-m3 and Qwen/Qwen3-Embedding-0.6B, applying each model's own query prefix. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy per model, plus indexing time and query time. Count the tokens in each passage and report how many go past each model's input limit, since those get cut before embedding. Then check each model's card for the shorter output sizes it was trained to support. For a model that lists them, such as Qwen3-Embedding, truncate its vectors to 512, 256, 128 and 64 dimensions, scale each one back to length 1, and report the same accuracy at each size, so I can see what shrinking the index costs. For a model that lists none, such as bge-small-en-v1.5, keep the full size and report it as the baseline a truncated vector has to beat. Finally, score every pair of passages and plot the distribution of unrelated pairs against correct query hits, one chart per model, since each model puts its scores in a different range.
Limits
What it can't do
- The score is a ranking unless you calibrate it. The distributions in the figure overlap, so a borrowed cutoff like 0.8 both keeps unrelated text and throws away correct answers. A cutoff fitted on your own labelled pairs can work for one model and one kind of document, and it has to be refitted when either changes.
- One vector has to stand for the whole passage. A long passage covering several subjects gets averaged into something that represents none of them well, which is why chunking your documents is a real decision and not a formality.
- Changing the model means rebuilding the index. Vectors from two models are not comparable, so an upgrade is a full re-embedding run over everything you have.
- It is weak on the things keyword search is strong on. Product codes, names and identifiers get blurred into the general meaning, which is why production systems run keyword search alongside it.
- Long passages are silently cut. Anything past the model's input limit, 512 tokens for
bge-small-en-v1.5, is dropped before embedding, and the vector gives no sign of it. - The query prefix is easy to get wrong. Instruction-tuned models expect one, and omitting it costs accuracy silently.
- Benchmarks move faster than they transfer. MTEB rankings change week to week and are measured on public datasets, so the only comparison that settles a choice is the one you run on your own documents and your own questions.
Go deeper
Reimers and Gurevych (2019): Sentence-BERT, Sentence Embeddings using Siamese BERT-Networks · Muennighoff et al. (2022): MTEB, Massive Text Embedding Benchmark · Kusupati et al. (2022): Matryoshka Representation Learning · Khattab and Zaharia (2020): ColBERT, late interaction over BERT · The MTEB leaderboard, which moves week to week
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
