What it is
It is the most accurate retriever here, and the most expensive to store
The argument for late interaction is easiest to see next to the alternatives. The figure runs five ways of searching over the same 287 passages, taken from this site's own articles, with the same eight questions asked of each.

ColBERT answered seven of the eight from the right article, which no single-vector method managed. It also produced 23,811 token vectors for those 287 passages, averaging 83 per passage, which is 12.19 MB against the 0.44 MB the embedding model needed for the same text.
The reranker column is the interesting comparison. Both ColBERT and a reranker get their accuracy from reading the query against the document at a finer grain than one vector allows, and they pay for it at different times. The reranker pays at query time, taking 1.2 seconds to rescore ten candidates. ColBERT pays at indexing time and in storage, and then answers in 30 milliseconds.
All of those costs come from one simple setup, and they compare with each other only because the setup was the same. Everything ran on a Mac CPU. Both indexes are raw 32-bit floats with no compression, so the 27.7 times is one vector per token against one per passage at the same precision. ColBERT's 30 milliseconds is a brute-force MaxSim against all 287 passages, with no candidate index, and the reranker's 1.2 seconds is ten pairs scored two at a time. The official ColBERT indexing code, a GPU, or a collection of millions of passages would move every one of those numbers, and the sections below say in which direction.
Where it shows up
Where ColBERT shows up
First-stage retrieval where recall has to be high
The usual place. ColBERT replaces the embedding model in stage one, so what reaches the reranker is a better shortlist. When the answer has to be in the top 10 or the whole pipeline fails, this is the cheapest way to raise that number.
Domains where the vocabulary is specific
Product codes, drug names, legal citations and internal jargon all suffer under a single pooled vector, because the pooling averages away the one word that mattered. Keeping a vector per token keeps that word addressable.
Retrieval you want to explain
Because the score is a sum over query tokens, you can print which passage word each query word matched. The bi-encoders article's late-interaction figure does exactly that, and it is the reason a ColBERT result is easier to debug than a cosine.
Document retrieval over page images
ColPali applies the same MaxSim machinery to the patches of a rendered page rather than the tokens of a passage. It is the largest current use of late interaction, and it grew faster than ColBERT itself.
The middle option when a reranker is too slow
A cross-encoder reranker is more accurate and cannot be precomputed. ColBERT precomputes everything it can and gives up the per-pair attention, which puts it between the two designs the bi-encoders article covers.
How it works
MaxSim, query augmentation, and the index that follows
The score is a sum of best matches
Both texts go through the same BERT encoder, with a [Q] marker prepended to queries and a [D] marker to documents so the model knows which it is reading. Every output vector is projected down to 128 numbers and scaled to length 1.
Scoring one pair works like this. Each query vector is compared against every vector in the passage, the single highest similarity is kept, and those kept values are added together. The paper calls the operation MaxSim, and a passage's score is the sum of it over all query positions.
Nothing in that step involves the model. The encoding happened earlier for the passage and just now for the query, and the comparison is arithmetic over vectors that already exist, which is why ColBERT can run over a whole collection where a reranker cannot.
Every query is padded with mask tokens, and they matter more than you'd think
A ColBERT query is padded out to a fixed length with [mask] tokens, and the encoder produces a vector for each of those positions too. The paper calls this query augmentation and describes it as "a soft, differentiable mechanism for learning to expand queries with new terms or to re-weigh existing terms based on their importance for matching the query".
On the one question measured in detail for the bi-encoders article, the twelve words actually typed scored the wrong passage higher, 7.67 against 7.58. The seventeen mask positions scored the right one higher, 8.73 against 8.00, and they decided the result. That is worth knowing before you conclude that a ColBERT result came from word matching.
The index is the price
A bi-encoder index holds one vector per passage. A ColBERT index holds one per token, so the size follows the length of your documents rather than their number. On the corpus here that was 27.7 times larger, and the ColBERTv2 paper puts the general case at "an order of magnitude" over single-vector models.
Santhanam et al. (2021) spent most of ColBERTv2 on that problem. Their residual compression stores each vector as the identifier of its nearest centroid plus a residual quantized to one or two bits per dimension, and recovers an approximation at search time by adding the two back together. The paper reports a 6 to 10 times reduction in footprint from it.
So the 12.19 MB above is the uncompressed number, which is what you get from running the encoder yourself. Use the official indexing code and it drops by most of an order of magnitude.
Searching without comparing everything
The run behind the figure compared the query against all 287 passages, which took 30 milliseconds and does not scale. Real deployments put the token vectors in an approximate nearest neighbour index, retrieve candidate passages by finding query tokens' nearest document tokens, and run full MaxSim only on the candidates. That machinery is what the ColBERT repository provides, and it is the reason most people run the library rather than the raw checkpoint.
Versions
The checkpoints you'll see, as of September 2026
| Model | Released | License | Downloads a month |
|---|---|---|---|
colbert-ir/colbertv2.0 | June 2023 | MIT | 2.42M |
answerdotai/answerai-colbert-small-v1 | August 2024 | Apache-2.0 | 838K |
jinaai/jina-colbert-v2 | August 2024 | CC-BY-NC-4.0 | 68K |
lightonai/GTE-ModernColBERT-v1 | April 2025 | Apache-2.0 | 27K |
colbertv2.0 is still the default, three years after release, which is the same pattern the BERT article found in encoders. answerai-colbert-small-v1 is the one to look at if size matters, at 33M parameters against colbertv2.0's 110M.
Check the licence before you commit. jina-colbert-v2 is CC-BY-NC-4.0, which rules out commercial use, and the other three on that list do not.
GTE-ModernColBERT-v1 is built on gte-modernbert-base, so it carries the longer context and faster inference that came with ModernBERT. It is the newest of the four and the least used, and at 149M parameters it is also the largest.
Choosing
When late interaction is worth its index
| The situation | Reach for | Why |
|---|---|---|
| A first retrieval system, with no measurements yet | An embedding model | Smallest index, simplest code, and it is the baseline everything else is measured against |
| Recall at the shortlist is the thing failing | ColBERT | It raised 4 of 8 to 7 of 8 here, at stage one, where a reranker cannot help |
| Exact names, codes and identifiers | BM25, or hybrid search | Cheaper than ColBERT and better at the specific thing ColBERT is being hired for |
| The top few results are in the wrong order | A reranker | It reads the pair together, which is more accurate than any precomputed comparison |
| Storage or memory is the binding constraint | An embedding model, with Matryoshka truncation | A ColBERT index is an order of magnitude larger before compression |
| Searching PDFs and slides as they look | ColPali | The same late interaction, over page images rather than text |
| You need one number you can threshold | A reranker, or Jev | A MaxSim sum grows with query length and is a ranking, not a probability |
Measure hybrid search before you reach for ColBERT. On the corpus behind the figure, fusing BM25 with a small embedding model matched ColBERT's 7 of 8 and cost nothing extra to store. ColBERT earns its index when that fusion is already in place and still misses.
Try it
How to try it
The official library handles the indexing, the compression and the candidate generation, and it is what you should use in anything real.
from ragatouille import RAGPretrainedModel
rag = RAGPretrainedModel.from_pretrained("colbert-ir/colbertv2.0")
rag.index(index_name="site", collection=passages) # writes a compressed index
hits = rag.search(query="how do I train a classifier on my own labels", k=3)
for hit in hits:
print(hit) # score, rank and the passage textScoring one pair by hand is worth doing once, because it shows that MaxSim is arithmetic and nothing more.
import torch
# qv: (32, 128) query token vectors, dv: (n, 128) passage token vectors,
# both already L2-normalised, so a dot product is a cosine.
sim = qv @ dv.T # every query token against every doc token
best = sim.max(dim=1).values # each query token keeps its best match
print("MaxSim score:", float(best.sum()))Measure whether ColBERT is worth the index on my own documents. Take a folder of Markdown or text files, split it into passages of at least 25 words, and build three retrieval systems over it: BM25, BAAI/bge-small-en-v1.5, and colbert-ir/colbertv2.0 through RAGatouille. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy for each system, plus indexing time, query time and on-disk index size in megabytes. Then add a fourth system that fuses the BM25 and embedding rankings with reciprocal rank fusion at k of 60, and report the same numbers for it. Print a table of the four, and list every question where the systems disagree, with what each one returned.
Limits
What it can't do
- The index is an order of magnitude larger. 12.19 MB against 0.44 MB for the same 287 passages here, before compression. ColBERTv2's residual compression takes 6 to 10 times of that back, and it is still the biggest index of the options on this page.
- It cannot be thresholded. A MaxSim score is a sum over query positions, so a longer query scores higher on everything. The scores on this run ranged from 12.53 to 22.46 for correct answers, and comparing them across queries means nothing.
- Indexing is a real cost. 38.6 milliseconds per passage on a CPU here, against 14.3 for the embedding model. For a million passages that gap is hours.
- It is still a bi-encoder at heart. The two texts never see each other inside the model, so a reranker can still catch things it misses. The two are usually run together rather than one instead of the other.
- Most of the tooling assumes the library. Vector databases are built around one vector per document, and multi-vector support is uneven, so check that your database handles it before planning around it.
- It was measured here on 287 passages and eight questions. That is enough to show the shape and nowhere near enough to choose a model. The last row of the figure, where every method is wrong and the question is at fault, is the reason to run this on your own data.
Go deeper
Khattab and Zaharia (2020): ColBERT, efficient and effective passage search via contextualized late interaction · Santhanam et al. (2021): ColBERTv2, effective and efficient retrieval via lightweight late interaction · Faysse et al. (2024): ColPali, which applies the same idea to page images · The ColBERT repository, which is where the indexing code lives
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
