MLGuerrillaStart with M1 →
Embeddings and retrieval·17 min read·Updated 24 September 2026

Text embeddings

These models turn a piece of text into one vector, so that two texts about the same thing end up close together. They are what makes search by meaning work, and they sit under every RAG system.

An embedding model takes a piece of text and returns one vector, which is a list of a few hundred numbers. The useful part is what happens to two of those vectors. Text about the same subject lands in the same region, so measuring the angle between two vectors tells you how close the two pieces of text are in meaning.

That one property is what search by meaning is built on. You embed every document you have once and keep the vectors. When a question arrives you embed it too, compare it against everything you stored, and return the closest passages. Nothing has to match word for word, so "how do I train a classifier on my own labelled examples" can find a paragraph that never uses any of those words.

What it is

Every passage becomes a vector once, and a query becomes one more

The two halves of the work happen at different times. Embedding your documents is a batch job you run once and repeat only when the content changes. Embedding the query happens on every search, and it is one model call over a short piece of text.

Comparing is the cheap part. Both vectors get scaled to length 1, and the comparison is a dot product, which for unit-length vectors is the cosine of the angle between them. A modern vector database does that against millions of stored vectors in milliseconds, using an index that skips most of the comparisons.

A figure titled "Every passage becomes a vector once, and a query becomes one more", built from 287 passages from the 13 articles on this site, embedded with bge-small-en-v1.5 and searched with one real question. The top card, labeled BUILT ONCE, shows three boxes joined by arrows: 287 passages from 13 articles and 16,831 words of this site, then bge-small-en-v1.5 at 33M parameters taking 14 ms per passage, then 287 vectors of 384 numbers each, stored and reused. Below it, labeled ASKED EVERY TIME, is the query "how do I train a classifier on my own labelled examples", with an arrow to a note saying one more vector, then 287 dot products. The second card lists the three nearest passages. The first scores 0.762 and comes from /models/yolo, reading "To detect your own classes, like hard hats or a particular product, you fine-tune on labeled ima...". The second scores 0.746, comes from /models/bert, reads "You load the pretrained encoder, attach a fresh output layer with one slot per label, and train ...", and is marked "answers it". The third scores 0.727, also from /models/bert, and is also marked "answers it". A note says the top hit is a YOLO passage about fine-tuning on labelled images, which is close but the wrong page. The caption reads: nothing here was trained on this site, and the vectors put the question near an answer it has never seen.
The query and the passage that answers it share almost no words. What they share is the subject, which is the only thing the vector encodes.

Look at what the top hit is. The model put a YOLO passage about fine-tuning on labelled images above the BERT passage that answers the question, because both are about training on your own labels and the model has no way to know which kind of labels were meant. That near miss is the normal failure of this stage, and the rerankers article measures how much of it a second stage recovers.

Where it shows up

Where embeddings show up

Retrieval for RAG

Retrieval-augmented generation puts passages from your own documents into an LLM's input so it answers from them. Embeddings are how the right passages get found. The LLM article covers what the model does with them once they arrive.

Search that understands a phrasing you didn't anticipate

Keyword search needs the user to type a word that appears in the document. Embedding search needs them to mean the same thing, which covers synonyms, paraphrases and questions asked from the wrong angle. Most production search runs both and merges the results, because keyword search is still better at names, error codes and part numbers.

Finding near-duplicates and clustering

Two support tickets describing the same bug land close together, so a threshold over pairwise similarity groups them. That threshold has to be chosen on pairs from your own tickets that someone has marked as duplicate or not, with the exact model you run, and the section on scores further down shows why a number borrowed from elsewhere won't hold. The same operation across a whole document collection gives you clusters you never defined.

Classification without a trained classifier

Embed your labelled examples once, then classify a new text by the label of its nearest neighbours. This works well enough on small label sets to be worth trying before you fine-tune anything, and the BERT article covers what fine-tuning buys you over it.

Deduplicating training data

Dropping near-identical examples before training is a standard cleanup step, and embeddings are how you find them at scale.

How it works

How the model produces one vector, and what the number of dimensions costs

The architecture is an encoder with a pooling step

Most embedding models are encoders, used in the bi-encoder arrangement. An encoder gives back one vector for every token, so a 25-token passage comes out as 25 vectors, and a search index has room for one. Pooling is the step that closes that gap.

A four-step figure titled "One vector per token goes in, one vector for the passage comes out", from a real run of bge-small-en-v1.5 on one 25-token passage from this site. Step one shows the passage cut into 25 token chips, starting with a CLS chip outlined in terracotta, then you, load, the, pre, and so on through to a separator chip, with the note that the CLS token is added by the tokenizer, is not part of your text, and is there so the model has a slot that belongs to no word. Step two, the encoder returns one vector per token, shows a table of 25 rows by 384 numbers with five real rows printed: CLS begins -0.62, -0.16, -0.01, -0.44, -0.15, +0.34; you begins -0.37, -0.31, +0.06, -0.43, -0.13, +0.17; load begins -0.82, +0.02, +0.35, -0.36, -0.06, +0.80; output begins -0.18, -0.08, +0.22, -0.09, +0.25, +0.15; train begins -0.47, -0.39, +0.15, -0.05, -0.26, +0.65; each with 378 more, and a line for the other 20 rows. It notes that every row is that token in this sentence, and that a search index cannot hold 25 rows per passage, so the 25 have to become 1. Step three splits into two cards. The left one, outlined in terracotta, is pool by taking the CLS row, which keeps that one row and drops the other 24, giving one vector of 384 numbers beginning -0.068, -0.018, -0.001, -0.048, and is what bge-small does. The right one is pool by averaging the 25 rows, which adds all 25 up, divides by 25 and scales to length 1, giving -0.067, -0.019, +0.007, -0.038, and is what E5, GTE and many others do. Both are labelled as the vector that goes in the index. Step four, the two poolings are not the same vector, reports that on this passage the two sit at a cosine of 0.94, that running the eight queries from this article both ways brings back a different top passage for 3 of the 8 at the same 4-of-8 accuracy, and that the model was trained with one of them, so that is the one to use for documents and queries alike. The caption reads: the model card tells you which row or average it was trained to use, and that is the one to use.
The token vectors in step two are the encoder's real output on that passage. The two pooled vectors below them are what you would actually store, and they are not the same list of numbers.

Two poolings cover most of the models you'll meet. Some take the vector at the [CLS] position, which is a slot the tokenizer adds at the front of every input and which belongs to no word, and some average all the token vectors together. bge-small-en-v1.5 is a CLS model, E5 and GTE are averaging models, and the card for each one says so.

The two give different vectors for the same text. On the passage in the figure they sit at a cosine of 0.94, and running the eight queries behind this article both ways brought back a different top passage for three of them, with accuracy at four of eight either way. A test that small cannot tell you which pooling is better. The model card can, because the model was trained with one of them and the other is a readout nobody optimised. Use the one the card names, for your documents and your queries alike.

The training is what makes the vectors useful. Reimers and Gurevych (2019) built Sentence-BERT by running the same encoder over two sentences and training so that related pairs come out close, which is the contrastive recipe CLIP uses across images and text. Every model on this page is trained some version of that way.

Many models need a prefix on the query

Instruction-tuned embedding models are trained to treat a query and a document differently, and they expect you to say which one you are handing them. bge-small-en-v1.5 wants queries prefixed with "Represent this sentence for searching relevant passages:", and Qwen3-Embedding wants an instruction line before the query. Skipping the prefix costs accuracy without any error message, and it is the most common way a first embedding pipeline underperforms.

The dimension number is a storage and speed decision

A vector of 1,024 numbers in 32-bit floats is 4 KB, so a million passages is 4 GB of index before any overhead. Cutting the vector in half halves that and roughly halves the comparison cost.

You can cut it because of Matryoshka Representation Learning (Kusupati et al., 2022), which trains a model so that the first half of a vector is itself a usable vector, and the first quarter after that. Google's docs say gemini-embedding-2 is trained this way, defaults to 3,072 numbers, and can be truncated to 768 or 1,536 with little loss. Cohere's embed-v4.0 offers 256, 512, 1,024 and 1,536, and Voyage's models offer 256 through 2,048.

Truncation only works on a model trained for it. Qwen3-Embedding's card lists output sizes from 32 up to the full length, and the hosted models above document their sizes, while a model trained without Matryoshka, such as bge-small-en-v1.5, has no reason to keep its meaning in the first numbers of the vector. Check the model card before cutting, and scale each shortened vector back to length 1 before comparing, since dropping numbers changes its length.

On the corpus behind the figures, truncating Qwen3-Embedding-0.6B from its full 1,024 numbers down to 128 returned the same top result for all eight queries. At 64 numbers, one of the eight changed. Eight queries can show a large loss and cannot show a small one, and the same top result does not mean the same ordering below it, so treat this as a reason to test 128 on your own questions and not as proof that 128 is free.

Long passages are cut off without a warning

The other kind of truncation happens on the way in. Every embedding model has a maximum input length, 512 tokens for bge-small-en-v1.5, and the tokenizer call in the code further down passes truncation=True, so anything past that point is dropped before the model sees it. No error is raised and the vector looks normal. It simply stands for the first 512 tokens of the passage.

The passages behind the figures were short enough that nothing was cut. On real documents, count tokens per chunk once, and make the chunk size smaller than the model's limit, including the query prefix on the query side.

A cosine is a ranking, and a threshold has to be fitted

The scores an embedding model returns look like confidence and behave nothing like it. The figure below scores every pair of the 287 passages on this site.

A figure titled "A cosine of 0.7 can be a perfect answer or two unrelated paragraphs", built from every pair of the 287 passages on this site scored with bge-small-en-v1.5 and counted into bands. On the left is a histogram with cosine similarity from 0.3 to 1.0 along the bottom. Grey bars show 20,000 pairs of passages from different articles, peaking between 0.55 and 0.65. Terracotta bars show 3,217 pairs from the same article, peaking slightly to the right between 0.6 and 0.7. The two distributions overlap almost completely. On the right, four measurements are listed: two unrelated passages average 0.600, the highest unrelated pair found scored 0.866, the eight real queries scored between 0.654 and 0.824 on their best hit, and the worst of those correct hits was 0.654. A note says a wrong pair scored higher than every correct answer found. The caption reads: sort by this number and it works, and put a cutoff on it and you are guessing.
The same model, the same corpus, one pass. The distributions are not separated, so no single cutoff divides them.

Two unrelated passages average 0.600, and the highest unrelated pair on this corpus scored 0.866. The eight real queries found their best answer somewhere between 0.654 and 0.824. A cutoff at 0.8 would have thrown away most of the correct answers and kept an unrelated pair.

So use the score to sort, take the top few, and decide relevance with something else. That something else is usually a reranker, which reads the query and one passage together and scores that pair.

A cosine cutoff can still work in a narrower setting. With one model, one kind of document and a few hundred pairs labelled as related or not, you can pick the cutoff that gives the precision you need and measure how much recall it costs. That number belongs to that model and that data. Unrelated passages average 0.600 under bge-small-en-v1.5 here, and another model puts its unrelated pairs somewhere else entirely, so the cutoff has to be fitted again whenever the model, the prefix or the kind of document changes.

Versions

The models you'll see, as of September 2026

Open models, with monthly downloads read 2026-09-21
ModelReleasedDimensionsLicenseDownloads a month
BAAI/bge-small-en-v1.5September 2023384MIT64.3M
BAAI/bge-m3January 20241024MIT37.3M
BAAI/bge-large-en-v1.5September 20231024MIT10.9M
Qwen/Qwen3-Embedding-0.6BJune 20251024Apache-2.08.7M
intfloat/multilingual-e5-largeJune 20231024MIT7.2M
google/embeddinggemma-300mJuly 2025768Gemma2.7M
Qwen/Qwen3-Embedding-8BJune 20254096Apache-2.02.6M
Qwen/Qwen3-VL-Embedding-8BJanuary 2026check the cardcheck the card1.2M

The download column tells the same story the BERT article tells. A model from September 2023, bge-small-en-v1.5, is pulled 64.3 million times a month, which is more than every model released in 2026 on that list combined. A 33M-parameter encoder that runs on a CPU keeps being the right answer.

Hosted models, from vendor docs read 2026-09-21
ModelInputDimensionsContext
Cohere embed-v4.0Text, images, and mixed documents such as PDFs256, 512, 1024 or 1536128k tokens
Google gemini-embedding-2Text, images, video, audio and documents3072 by default, truncatableSee the API docs
Google gemini-embedding-001Text3072 by default, truncatableSee the API docs
Voyage voyage-3.5 and voyage-3.5-liteText1024 by default, plus 256, 512, 204832,000 tokens

Two of those take more than text. Google describes gemini-embedding-2 as "the first multimodal embedding model in the Gemini API", putting text, images, video, audio and documents into one space, and Cohere's embed-v4.0 takes PDFs with their layout intact. One index across several kinds of content is the thing that changed in this category over the past year.

Watch for behaviour changes when you upgrade. Google's docs note that gemini-embedding-001 returns one embedding per string, while Embeddings 2 "produces a single aggregated embedding for multiple inputs", which is a different function with the same name.

Choosing

Which model to pick, and when an embedding is the wrong tool

Which model to reach for
The jobReach forWhy
A first retrieval system in Englishbge-small-en-v1.5 or bge-large-en-v1.5Small, free, CPU-friendly, and the most tested thing in the category
Retrieval across many languagesbge-m3 or a hosted multilingual modelTrained for it, where an English model degrades quietly
The best accuracy you can self-hostA current Qwen3-Embedding sizeIt measured 6 of 8 against 4 of 8 for bge-small on the corpus behind these figures
Searching PDFs and screenshots without OCRCohere embed-v4.0, or the colpali approachThey embed the page as an image, so layout survives
One index over text and imagesgemini-embedding-2 or a multimodal embedding modelOne space means one query reaches every kind of content
Deciding whether one text matches one other textA reranker, or SigLIP for image pairsA cosine cutoff holds only for the model and task it was fitted on with labelled pairs, and a reranker scores the pair directly
Exact matching on names, codes or IDsKeyword search such as BM25An embedding blurs exactly the detail those depend on
Writing an answer from what you foundAn LLMEmbeddings retrieve, and they produce no text

On the corpus behind these figures, bge-small got 4 of 8 queries right and Qwen3-Embedding-0.6B got 6 of 8, while taking 1,022 milliseconds per passage against 14. An 18-times bigger model bought two queries, and it made indexing 71 times slower. The reranker on the same corpus also bought two queries, from 4 of 8 to 6 of 8, without touching the index.

Those timings come from one setup. Both models ran in 32-bit floats on a Mac CPU, bge-small in batches of 32 passages and Qwen3 in batches of 16, over passages that were all well under 512 tokens. A GPU, a different batch size or longer passages change both numbers, and the ratio between them is the part that carries over. Two queries out of eight is also a small enough difference that a different set of eight questions could shrink it or widen it.

Try it

How to try it

Embedding a folder of text takes about fifteen lines. This is the code behind the first figure, cut down.

search.py — embed a corpus and search it
import torch
from transformers import AutoModel, AutoTokenizer

model_id = "BAAI/bge-small-en-v1.5"            # 33M, MIT, runs on a CPU
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id).eval()

def embed(texts, prefix=""):
    batch = tok([prefix + t for t in texts], padding=True,
                truncation=True, max_length=512, return_tensors="pt")
    with torch.no_grad():
        out = model(**batch).last_hidden_state[:, 0]   # bge pools the CLS token
    return torch.nn.functional.normalize(out, dim=-1)

index = embed(passages)                                # once, then store it
q = embed(["how do I train a classifier on my own labels"],
          prefix="Represent this sentence for searching relevant passages: ")
scores = (q @ index.T)[0]
for i in scores.topk(3).indices:
    print(round(float(scores[i]), 3), passages[i][:90])

Keep the vectors in a vector database once the corpus outgrows memory. The interface is the same, and what you gain is an index that finds the nearest vectors without comparing against all of them.

Ask your AI coding tool

Build a retrieval evaluation over my own documents so I can choose an embedding model with numbers. Read every Markdown file under a folder I pass in, split it into passages of at least 25 words, and embed them with BAAI/bge-small-en-v1.5, BAAI/bge-m3 and Qwen/Qwen3-Embedding-0.6B, applying each model's own query prefix. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy per model, plus indexing time and query time. Count the tokens in each passage and report how many go past each model's input limit, since those get cut before embedding. Then check each model's card for the shorter output sizes it was trained to support. For a model that lists them, such as Qwen3-Embedding, truncate its vectors to 512, 256, 128 and 64 dimensions, scale each one back to length 1, and report the same accuracy at each size, so I can see what shrinking the index costs. For a model that lists none, such as bge-small-en-v1.5, keep the full size and report it as the baseline a truncated vector has to beat. Finally, score every pair of passages and plot the distribution of unrelated pairs against correct query hits, one chart per model, since each model puts its scores in a different range.

Limits

What it can't do

  • The score is a ranking unless you calibrate it. The distributions in the figure overlap, so a borrowed cutoff like 0.8 both keeps unrelated text and throws away correct answers. A cutoff fitted on your own labelled pairs can work for one model and one kind of document, and it has to be refitted when either changes.
  • One vector has to stand for the whole passage. A long passage covering several subjects gets averaged into something that represents none of them well, which is why chunking your documents is a real decision and not a formality.
  • Changing the model means rebuilding the index. Vectors from two models are not comparable, so an upgrade is a full re-embedding run over everything you have.
  • It is weak on the things keyword search is strong on. Product codes, names and identifiers get blurred into the general meaning, which is why production systems run keyword search alongside it.
  • Long passages are silently cut. Anything past the model's input limit, 512 tokens for bge-small-en-v1.5, is dropped before embedding, and the vector gives no sign of it.
  • The query prefix is easy to get wrong. Instruction-tuned models expect one, and omitting it costs accuracy silently.
  • Benchmarks move faster than they transfer. MTEB rankings change week to week and are measured on public datasets, so the only comparison that settles a choice is the one you run on your own documents and your own questions.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.