MLGuerrillaStart with M1 →
Architectures·12 min read·Updated 24 September 2026

Bi-encoders and cross-encoders

These are the two ways to score how well two pieces of text go together. A bi-encoder reads each one separately, which makes it fast, and a cross-encoder reads them together, which makes it accurate.

Almost everything a retrieval system does comes down to one question, which is how well a query and a document go together. There are two ways to ask an encoder that question, and they differ in one respect. Either the document is read on its own, before anyone has asked anything, or it is read together with the query.

That one difference decides the speed, the accuracy and the shape of the system you build around it. A bi-encoder reads separately and can precompute, so it can search millions of documents. A cross-encoder reads together, so there is nothing about a document it can compute before the query arrives, and it can only score a shortlist.

What it is

The query is present, or it isn't

A bi-encoder runs the encoder twice, once over each text, and compares the two vectors it gets back. The document half runs once when you build your index, and the query half runs when somebody searches. Everything about its speed follows from the document vector existing before the query does.

A cross-encoder runs the encoder once over both texts joined into a single input, with a separator between them, and reads a score off the end. Every layer of the model can look from a word in the query to a word in the document, which is the accuracy, and the score exists only for that one pair, which is the cost.

A figure titled "Encode the two texts separately, or encode them together", noting that the same encoder architecture is used both times and that what changes is whether the query is present when the document is read. The left panel, BI-ENCODER, is labelled "the document is read before the query exists" and shows two columns. The document goes into an encoder marked "once, ever" and becomes a stored vector. The query goes into an encoder marked "once per query" and becomes a vector. The two vectors are joined by a line marked "dot". A note reads: 61.8M comparisons a second once the vectors exist. The right panel, CROSS-ENCODER, outlined in terracotta and labelled "both texts go in together, as one input", shows a single chip reading "the query, then a separator token, then the document" feeding an encoder marked "once per pair, every time", which produces one score. A note reads: nothing can be stored, because the score needs both. A strip below, titled "what that costs on 10,000 documents" using the measured 14.3 ms to embed a passage and 120.7 ms to score a pair on the same laptop CPU, gives four figures. Building the index takes 143 seconds, once ever. Answering one query takes 15 milliseconds, being one embed then the comparisons. Cross-encoding all 10,000 takes 20 minutes for a single query. Cross-encoding 100 queries takes 1.4 days. The caption reads: a search engine uses the first, a reranker uses the second, and neither replaces the other.
The 20 minutes is the same measured 120.7 ms per pair, multiplied by 10,000. Nothing about it gets cheaper with a better index, because there is no index to build.

The numbers in that strip are the argument. Building a searchable index of 10,000 passages took 143 seconds once. Answering a query off that index took 15 milliseconds. Scoring the same 10,000 passages with a cross-encoder takes 20 minutes, for one question, every time it is asked.

All of those are sequential timings, one text or one pair at a time on a laptop CPU, which is how the 120.7 ms per pair was measured. A real reranker sends pairs through the model in batches, and on a GPU a batch of dozens of pairs usually takes only a little longer than a single pair, so the 20 minutes could shrink a lot. This page didn't measure a batched run, so treat the size of that drop as something to time on your own hardware. What batching can't change is that the work still grows with the number of documents scored for every query, because each pair is a full pass through the model. The bi-encoder's per-query work stays at one embedding and a cheap comparison whatever the collection size, so batching narrows the gap between the two designs without removing it.

Which models use this

Which models are which

The two designs across the index
DesignModelsWhat it is used for
Bi-encoderText embedding models such as bge, E5 and Qwen3-EmbeddingSearch, clustering, deduplication
Bi-encoderCLIP and SigLIP, with one tower for images and one for textImage search and zero-shot classification
Bi-encoderTwo-tower retrieval models in recommendation systemsPicking a few hundred candidates out of millions of items
Cross-encoderRerankers such as bge-reranker and Cohere RerankRescoring a shortlist
Cross-encoderNatural language inference models, which judge whether one sentence follows from anotherFact checking and contradiction detection
Cross-encoderJev and the classifiers that read a document and a question togetherTyped decisions over one piece of text

CLIP makes that trade most visibly. Its two towers never see each other's input, which is exactly why you can embed a million photos once and then search them with any sentence you type. The price is the one the CLIP article measures, where the raw scores sit in a narrow band and mean little on their own.

How it works

What reading together buys, and what it costs

The bi-encoder has to guess what you will ask

A document vector has to be useful for every question anyone might ask of that document, because it is computed before any of them arrive. It is a summary with no idea what it will be asked.

That is why two passages about training on labelled data end up near each other even when one is about images and the other about text. The distinction was never in the question when the vector was made.

The cross-encoder is told what you asked

With both texts in one input, attention can run between them at every layer. The model can check the specific words of your question against the specific words of the passage, which is how the rerankers run separates "labelled examples" in a question about text from "labeled images" in a passage about object detection.

The measured gain on that run was 4 of 8 questions answered from the right article becoming 6 of 8, on the same corpus, with the same first stage.

Precomputation is the whole difference

The bi-encoder's document pass depends on nothing but the document, so its result can be stored and reused forever, for every query anyone asks. The cross-encoder has no query-independent step to store. Every layer mixes query tokens with document tokens, so even the document's first-layer vectors change when the query changes.

A cross-encoder's output can still be cached, but only as a score for one exact pair. If the same query comes back against the same passage, a cache keyed on both texts returns the old score without running the model. That helps with repeated popular queries, and it does nothing for a query that differs by one word, since that is a new pair. A cache like that can't be indexed or searched either, because each score belongs to a single query.

This is also why a bi-encoder comparison is nearly free. Once both vectors exist, comparing them is a dot product, and a million of those took 16 milliseconds on the machine behind the figure.

ColBERT sits between the two

Khattab and Zaharia (2020) built ColBERT on a third arrangement. It runs the encoder over the document ahead of time the way a bi-encoder does, and it keeps the vector for every token rather than pooling them into one. At query time it does the same for the query, then compares the two sets of vectors token by token, which the paper calls late interaction.

The comparison itself is called MaxSim. Every query vector finds the closest vector in the passage, that one similarity is kept, and the kept scores are added up to give the passage its score. Running it on the question the embedding model got wrong shows both the mechanism and a piece of ColBERT that is easy to miss.

A five-step figure titled "Late interaction keeps every token, and scores them one at a time", from a real run of colbert-ir/colbertv2.0 on the one question the bi-encoder got wrong. Step one shows the query cut into 12 token chips, how, do, i, train, a, class, ifier, on, my, own, labelled, examples, each becoming a vector of 128 numbers, with the note that an embedding model would pool these into one vector and drop the 12, and ColBERT keeps all 12. Step two shows the two candidate passages side by side. The left one is from /models/yolo, stored as 146 token vectors, reading "To detect your own classes, like hard hats or a particular product, you fine-tune on labeled images", and is labelled the stage-1 winner. The right one, outlined in terracotta, is from /models/bert, stored as 74 token vectors, reading "You load the pretrained encoder, attach a fresh output layer with one slot per label, and train on your labeled", and is labelled the right answer. Step three lists each of the 12 typed tokens with its best match in each passage. Against the YOLO passage, train matches train at 0.85, own matches own at 0.84, labelled matches labeled at 0.82 and examples matches only like at 0.35. Against the BERT passage, train matches train at 0.91, labelled matches labeled at 0.85 and examples matches examples at 0.92. The 12 add up to 7.67 for YOLO and 7.58 for BERT, so the YOLO passage is ahead on the words that were typed. Step four, outlined in terracotta, is headed: ColBERT also scores 17 empty positions. It explains that every query is padded out to 32 positions with mask tokens, that the encoder gives each of those a vector too, and that the paper calls this query augmentation and describes it as a learned way to expand the query with terms you did not type. Four mask rows are shown, where mask 1 finds only like at 0.40 in the YOLO passage and finds examples at 0.83 in the BERT one. The 17 mask positions add up to 8.00 for YOLO and 8.73 for BERT. Step five totals all 32 positions. YOLO scores 7.67 plus 8.00 plus 1.92 for the 3 marker tokens, giving 17.58. BERT scores 7.58 plus 8.73 plus 2.14, giving 18.45, and wins. A closing line says the words you typed favoured the wrong passage, and the positions ColBERT added on top of them are what moved the right one to the front. The caption reads: both texts are turned into vectors ahead of time, and the matching is the part that waits for the query.
The mask rows in step four are not a quirk of this run. Every ColBERT query has them, and on this pair they carry more of the score than the words do.

Two things in that run are worth taking away. The first is that word-level matching does what you would expect, and "examples" in the question finds the word examples in the BERT passage at 0.92 while finding nothing better than "like" in the YOLO one at 0.35.

The second is that the typed words alone still put the wrong passage first, at 7.67 against 7.58. What moved the right passage to the front was the 17 mask positions ColBERT pads every query with. The paper describes that padding as "a soft, differentiable mechanism for learning to expand queries with new terms or to re-weigh existing terms", and on this pair it contributed more of the score than the question did.

The trade is index size. A passage that cost one vector now costs one per token, so the 287 passages behind these figures would go from 287 vectors to tens of thousands. ColBERT's own later work spends most of its effort on compressing exactly that.

Choosing

Choosing, which usually means using both

Which design fits the step
The stepDesignWhy
Finding candidates in a large collectionBi-encoderThe document vectors already exist, so the search is a comparison
Reordering the top 10 to 50Cross-encoderIt reads the pair, and the shortlist is small enough to afford it
Comparing two fixed texts, one pair at a timeCross-encoderNo index is involved, so nothing is gained by precomputing
Clustering or deduplicating a collectionBi-encoderVectors can be compared with each other, and pair scores cannot
Matching text against imagesA bi-encoder such as CLIPTwo towers is what makes a searchable image index possible
Deciding one yes-or-no about one documentCross-encoder, or JevBoth read the pair and return a score you can threshold

The usual answer is both, in that order. Retrieve with the bi-encoder, rerank the top of it with the cross-encoder, and choose the shortlist depth by measuring how often the right answer appears in it. That structure is why the text embeddings and rerankers articles use the same corpus and the same eight questions.

Try it

How to try it

Both designs are a few lines in transformers, and running them side by side on your own pairs is the fastest way to feel the difference.

both.py — the same pair scored two ways
import torch
from transformers import (AutoModel, AutoModelForSequenceClassification,
                          AutoTokenizer)

query, doc = "how do I train a classifier", "You load the pretrained encoder..."

# bi-encoder: two passes, and the document half is reusable
bt = AutoTokenizer.from_pretrained("BAAI/bge-small-en-v1.5")
bm = AutoModel.from_pretrained("BAAI/bge-small-en-v1.5").eval()
def vec(text):
    with torch.no_grad():
        h = bm(**bt(text, return_tensors="pt")).last_hidden_state[:, 0]
    return torch.nn.functional.normalize(h, dim=-1)
print("cosine:", float(vec(query) @ vec(doc).T))

# cross-encoder: one pass over both, and nothing is reusable
ct = AutoTokenizer.from_pretrained("BAAI/bge-reranker-v2-m3")
cm = AutoModelForSequenceClassification.from_pretrained(
    "BAAI/bge-reranker-v2-m3").eval()
with torch.no_grad():
    print("score:", float(cm(**ct([[query, doc]],
          return_tensors="pt")).logits))

The two numbers are on different scales and are not comparable. The cosine sits in a narrow band across all pairs, and the cross-encoder score spreads out, which is the practical reason one sorts well and the other thresholds well.

Ask your AI coding tool

Show me the speed and accuracy trade between the two designs on my own data. Take a CSV of query and document pairs with a relevance label. Score every pair two ways, once with BAAI/bge-small-en-v1.5 as a bi-encoder using cosine similarity, and once with BAAI/bge-reranker-v2-m3 as a cross-encoder. Report the area under the ROC curve for each. Time the cross-encoder twice, once scoring pairs one at a time and once in batches of 32, and time the bi-encoder's document embedding separately from its query embedding, since the documents only need embedding once. Then estimate how long each design would take to answer one new query against 10,000 documents, counting only the work that has to happen after the query arrives. Then plot both score distributions for relevant and irrelevant pairs on one chart, so I can see which of the two has a usable threshold.

Limits

What neither design fixes

  • A bi-encoder cannot recover a distinction it was never told about. The document vector is made before the query exists, so a question that turns on a detail the summary dropped cannot be answered by better comparison.
  • A cross-encoder cannot search. Scoring every document for every query is the 20 minutes in the figure when run one pair at a time, and batching on a GPU cuts that time without changing how the work grows with the collection, so it always depends on something else handing it a shortlist.
  • Neither score is calibrated out of the box. A cosine sits in a narrow band, and a cross-encoder logit is on no fixed scale, so any threshold has to be fitted on your own labelled pairs.
  • Both truncate. A bi-encoder truncates the document to its context limit, and a cross-encoder splits that limit between the query and the document, so long passages lose their tails either way.
  • The accuracy gap depends on your corpus. The measured gain behind these pages was two questions out of eight on one small corpus, and on a corpus where retrieval already ranks well, a reranker adds latency and little else.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.