MLGuerrillaStart with M1 →
Embeddings and retrieval·15 min read·Updated 24 September 2026

ColBERT and late interaction

ColBERT is a retriever that stores one vector for every token instead of one for the whole passage, and compares the two sets at query time. It was the most accurate retriever on the corpus behind this article, and its index was 27.7 times the size.

ColBERT is a retrieval model. It takes a query and a collection of passages as input, and it returns the passages most likely to answer the query, which is the same job an embedding model does.

What makes it different is how much it keeps. An embedding model runs an encoder over a passage, gets one vector per token, and then pools those into a single vector for the whole passage. ColBERT skips the pooling. It stores all of the token vectors, does the same for the query, and compares the two sets token by token when someone searches. Khattab and Zaharia (2020) called that comparison late interaction, because the two texts are encoded separately and only meet at the end.

What it is

It is the most accurate retriever here, and the most expensive to store

The argument for late interaction is easiest to see next to the alternatives. The figure runs five ways of searching over the same 287 passages, taken from this site's own articles, with the same eight questions asked of each.

A grid figure titled "Five ways to search the same 287 passages", from this site's own 13 articles cut into passages and asked the same eight questions, where a filled terracotta chip means the right article came back first. The five columns are BM25 keyword search, bge-small with one vector each, those two fused with RRF, bge-small followed by a reranker over the top 10, and ColBERT with one vector per token. Row by row: for "train a classifier on my own labels", BM25 returns bert correctly while bge-small returns yolo, and the fusion, the reranker and ColBERT all return bert. For "why can't it predict the next word", every column returns bert correctly. For "what makes image generation slow", BM25 returns diffusion correctly, bge-small returns grounding-dino, and the other three return diffusion. For "find objects and draw boxes", BM25 returns dino while the other four return yolo correctly. For "search photos by typing a sentence", BM25 returns bert while the other four return clip correctly. For "what attention does to each token", every column returns transformer correctly. For "get a probability, not text", BM25 returns enc-dec-arch and bge-small returns t5-enc-dec, both wrong, the fusion returns jev correctly, the reranker returns enc-dec-arch, and ColBERT returns jev correctly. For "what the model does with audio", every column is wrong, returning llms, bert, transformer, transformer and transformer. The totals are 4 of 8 for BM25, 4 of 8 for bge-small, 7 of 8 for the fusion, 6 of 8 for the reranker and 7 of 8 for ColBERT. Per query the costs are 0.11 ms, a few ms, a few ms, 1.2 seconds and 30 ms, and the indexes for 287 passages are an inverted index, 0.44 MB, both of those, 0.44 MB and 12.19 MB. A panel below reads: BM25 and the embedding model both score 4 of 8 and they are not the same 4, each one answers a question the other misses; fusing those two lists costs nothing and scores 7 of 8, including one question both halves got wrong on their own; ColBERT reaches the same 7 of 8 on its own and its index is 27.7 times the size; and the last row is wrong everywhere with the question at fault, because only two passages in the whole corpus mention audio. The caption reads: a retrieval score is a score for your corpus and your questions, and not for the model.
The same corpus and the same eight questions run through the text-embeddings, rerankers and bi-encoders articles, so the numbers in all four line up.

ColBERT answered seven of the eight from the right article, which no single-vector method managed. It also produced 23,811 token vectors for those 287 passages, averaging 83 per passage, which is 12.19 MB against the 0.44 MB the embedding model needed for the same text.

The reranker column is the interesting comparison. Both ColBERT and a reranker get their accuracy from reading the query against the document at a finer grain than one vector allows, and they pay for it at different times. The reranker pays at query time, taking 1.2 seconds to rescore ten candidates. ColBERT pays at indexing time and in storage, and then answers in 30 milliseconds.

All of those costs come from one simple setup, and they compare with each other only because the setup was the same. Everything ran on a Mac CPU. Both indexes are raw 32-bit floats with no compression, so the 27.7 times is one vector per token against one per passage at the same precision. ColBERT's 30 milliseconds is a brute-force MaxSim against all 287 passages, with no candidate index, and the reranker's 1.2 seconds is ten pairs scored two at a time. The official ColBERT indexing code, a GPU, or a collection of millions of passages would move every one of those numbers, and the sections below say in which direction.

Where it shows up

Where ColBERT shows up

First-stage retrieval where recall has to be high

The usual place. ColBERT replaces the embedding model in stage one, so what reaches the reranker is a better shortlist. When the answer has to be in the top 10 or the whole pipeline fails, this is the cheapest way to raise that number.

Domains where the vocabulary is specific

Product codes, drug names, legal citations and internal jargon all suffer under a single pooled vector, because the pooling averages away the one word that mattered. Keeping a vector per token keeps that word addressable.

Retrieval you want to explain

Because the score is a sum over query tokens, you can print which passage word each query word matched. The bi-encoders article's late-interaction figure does exactly that, and it is the reason a ColBERT result is easier to debug than a cosine.

Document retrieval over page images

ColPali applies the same MaxSim machinery to the patches of a rendered page rather than the tokens of a passage. It is the largest current use of late interaction, and it grew faster than ColBERT itself.

The middle option when a reranker is too slow

A cross-encoder reranker is more accurate and cannot be precomputed. ColBERT precomputes everything it can and gives up the per-pair attention, which puts it between the two designs the bi-encoders article covers.

How it works

MaxSim, query augmentation, and the index that follows

The score is a sum of best matches

Both texts go through the same BERT encoder, with a [Q] marker prepended to queries and a [D] marker to documents so the model knows which it is reading. Every output vector is projected down to 128 numbers and scaled to length 1.

Scoring one pair works like this. Each query vector is compared against every vector in the passage, the single highest similarity is kept, and those kept values are added together. The paper calls the operation MaxSim, and a passage's score is the sum of it over all query positions.

Nothing in that step involves the model. The encoding happened earlier for the passage and just now for the query, and the comparison is arithmetic over vectors that already exist, which is why ColBERT can run over a whole collection where a reranker cannot.

Every query is padded with mask tokens, and they matter more than you'd think

A ColBERT query is padded out to a fixed length with [mask] tokens, and the encoder produces a vector for each of those positions too. The paper calls this query augmentation and describes it as "a soft, differentiable mechanism for learning to expand queries with new terms or to re-weigh existing terms based on their importance for matching the query".

On the one question measured in detail for the bi-encoders article, the twelve words actually typed scored the wrong passage higher, 7.67 against 7.58. The seventeen mask positions scored the right one higher, 8.73 against 8.00, and they decided the result. That is worth knowing before you conclude that a ColBERT result came from word matching.

The index is the price

A bi-encoder index holds one vector per passage. A ColBERT index holds one per token, so the size follows the length of your documents rather than their number. On the corpus here that was 27.7 times larger, and the ColBERTv2 paper puts the general case at "an order of magnitude" over single-vector models.

Santhanam et al. (2021) spent most of ColBERTv2 on that problem. Their residual compression stores each vector as the identifier of its nearest centroid plus a residual quantized to one or two bits per dimension, and recovers an approximation at search time by adding the two back together. The paper reports a 6 to 10 times reduction in footprint from it.

So the 12.19 MB above is the uncompressed number, which is what you get from running the encoder yourself. Use the official indexing code and it drops by most of an order of magnitude.

Searching without comparing everything

The run behind the figure compared the query against all 287 passages, which took 30 milliseconds and does not scale. Real deployments put the token vectors in an approximate nearest neighbour index, retrieve candidate passages by finding query tokens' nearest document tokens, and run full MaxSim only on the candidates. That machinery is what the ColBERT repository provides, and it is the reason most people run the library rather than the raw checkpoint.

Versions

The checkpoints you'll see, as of September 2026

Open ColBERT-family models, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
colbert-ir/colbertv2.0June 2023MIT2.42M
answerdotai/answerai-colbert-small-v1August 2024Apache-2.0838K
jinaai/jina-colbert-v2August 2024CC-BY-NC-4.068K
lightonai/GTE-ModernColBERT-v1April 2025Apache-2.027K

colbertv2.0 is still the default, three years after release, which is the same pattern the BERT article found in encoders. answerai-colbert-small-v1 is the one to look at if size matters, at 33M parameters against colbertv2.0's 110M.

Check the licence before you commit. jina-colbert-v2 is CC-BY-NC-4.0, which rules out commercial use, and the other three on that list do not.

GTE-ModernColBERT-v1 is built on gte-modernbert-base, so it carries the longer context and faster inference that came with ModernBERT. It is the newest of the four and the least used, and at 149M parameters it is also the largest.

Choosing

When late interaction is worth its index

Which retriever fits the job
The situationReach forWhy
A first retrieval system, with no measurements yetAn embedding modelSmallest index, simplest code, and it is the baseline everything else is measured against
Recall at the shortlist is the thing failingColBERTIt raised 4 of 8 to 7 of 8 here, at stage one, where a reranker cannot help
Exact names, codes and identifiersBM25, or hybrid searchCheaper than ColBERT and better at the specific thing ColBERT is being hired for
The top few results are in the wrong orderA rerankerIt reads the pair together, which is more accurate than any precomputed comparison
Storage or memory is the binding constraintAn embedding model, with Matryoshka truncationA ColBERT index is an order of magnitude larger before compression
Searching PDFs and slides as they lookColPaliThe same late interaction, over page images rather than text
You need one number you can thresholdA reranker, or JevA MaxSim sum grows with query length and is a ranking, not a probability

Measure hybrid search before you reach for ColBERT. On the corpus behind the figure, fusing BM25 with a small embedding model matched ColBERT's 7 of 8 and cost nothing extra to store. ColBERT earns its index when that fusion is already in place and still misses.

Try it

How to try it

The official library handles the indexing, the compression and the candidate generation, and it is what you should use in anything real.

index.py — ColBERT through its own library
from ragatouille import RAGPretrainedModel

rag = RAGPretrainedModel.from_pretrained("colbert-ir/colbertv2.0")
rag.index(index_name="site", collection=passages)   # writes a compressed index

hits = rag.search(query="how do I train a classifier on my own labels", k=3)
for hit in hits:
    print(hit)          # score, rank and the passage text

Scoring one pair by hand is worth doing once, because it shows that MaxSim is arithmetic and nothing more.

maxsim.py — one query against one passage
import torch

# qv: (32, 128) query token vectors, dv: (n, 128) passage token vectors,
# both already L2-normalised, so a dot product is a cosine.
sim = qv @ dv.T                     # every query token against every doc token
best = sim.max(dim=1).values        # each query token keeps its best match
print("MaxSim score:", float(best.sum()))
Ask your AI coding tool

Measure whether ColBERT is worth the index on my own documents. Take a folder of Markdown or text files, split it into passages of at least 25 words, and build three retrieval systems over it: BM25, BAAI/bge-small-en-v1.5, and colbert-ir/colbertv2.0 through RAGatouille. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy for each system, plus indexing time, query time and on-disk index size in megabytes. Then add a fourth system that fuses the BM25 and embedding rankings with reciprocal rank fusion at k of 60, and report the same numbers for it. Print a table of the four, and list every question where the systems disagree, with what each one returned.

Limits

What it can't do

  • The index is an order of magnitude larger. 12.19 MB against 0.44 MB for the same 287 passages here, before compression. ColBERTv2's residual compression takes 6 to 10 times of that back, and it is still the biggest index of the options on this page.
  • It cannot be thresholded. A MaxSim score is a sum over query positions, so a longer query scores higher on everything. The scores on this run ranged from 12.53 to 22.46 for correct answers, and comparing them across queries means nothing.
  • Indexing is a real cost. 38.6 milliseconds per passage on a CPU here, against 14.3 for the embedding model. For a million passages that gap is hours.
  • It is still a bi-encoder at heart. The two texts never see each other inside the model, so a reranker can still catch things it misses. The two are usually run together rather than one instead of the other.
  • Most of the tooling assumes the library. Vector databases are built around one vector per document, and multi-vector support is uneven, so check that your database handles it before planning around it.
  • It was measured here on 287 passages and eight questions. That is enough to show the shape and nowhere near enough to choose a model. The last row of the figure, where every method is wrong and the question is at fault, is the reason to run this on your own data.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.