MLGuerrillaStart with M1 →
Embeddings and retrieval·14 min read·Updated 24 September 2026

Rerankers

A reranker reads your query and one document together and scores that pair. It runs on the shortlist that search already returned, and it is usually the cheapest accuracy you can add to a retrieval system.

A reranker takes one query and one document, reads them together in a single pass, and returns a score for how well that document answers that query. It does the same job a search index does, and it does it far more accurately and far more slowly.

That cost is why a reranker never searches. Search runs first over everything you have and returns a shortlist, and the reranker rescores only that shortlist. The text embeddings article covers the first stage, and this page covers what the second one recovers.

What it is

It rescores a shortlist, and it moves the top of it

The mechanism is what makes it accurate. An embedding model has to turn the document into a vector before it has seen your query, so the vector has to be good for every question anyone might ask. A reranker sees both at once, so it can judge this document against this question and nothing else.

Cohere's own documentation describes the input plainly, saying Rerank "combines the tokens from the query with the tokens from the document and the combined total counts toward the context limit for a single document". You can send many documents in one request, and each one is still read against the query on its own and given its own score.

Running both stages over the same corpus and the same eight questions as the embeddings article shows what the second one is worth.

A figure titled "The reranker reads the query and the passage together, and it moves the top", using the same 287 passages and the same eight questions as the embeddings article with a second stage added. The top strip shows a five-box pipeline: 287 passages, everything you have; then bge-small-en-v1.5 at 14 ms each, built once, labeled STAGE 1 SEARCH; then top 10, one dot product each; then bge-reranker-v2-m3 at 121 ms per pair, labeled STAGE 2 RESCORE and outlined in terracotta; then the answer, one passage. Below, eight questions are listed with the article each one landed on after search and after reranking. "How do I train a classifier on my own labelled examples" went from yolo to bert and is marked fixed. "Why can a model not predict the next word" stayed on bert. "What makes image generation slow" went from grounding-dino to diffusion-models and is marked fixed. "How does a model find objects in a photo" stayed on yolo, "can I search my photos by typing a sentence" stayed on clip, and "what is attention doing to each token" stayed on transformer. "How do I get a probability instead of generated text" moved from t5-and-encoder-decoder-models to encoder-decoder-architectures and is marked still wrong. "What does the model do with audio" moved from bert to transformer and is marked never in the top 10. A summary line reads: 4 of 8 right after search, 6 of 8 after reranking, and the top result changed on 6 of 8. The caption reads: the reranker only reorders what search handed it, and on the one question whose answer never reached the top 10 it had nothing to fix.
Both stages ran on a laptop CPU. The reranker is the same size class as the retriever and takes roughly nine times as long per document, because it runs once for every pair, where the retriever ran once per document and kept the result.

Two of the eight questions changed from a wrong article to the right one. The first is the near miss from the embeddings article, where "how do I train a classifier on my own labelled examples" landed on a YOLO passage about fine-tuning on labelled images. Reading the question and the passage together is enough to tell that the question is about text.

Where it shows up

Where rerankers show up

The second stage of RAG

The passages an LLM receives decide what it can answer with. Retrieving 50 and reranking down to 5 puts better passages in front of the model without making the prompt longer, and a shorter prompt with better passages usually beats a longer one.

Search results a person reads

Order counts for more when a human is scanning, because almost nobody looks past the first few results. This is the job the design was built for, and Nogueira and Cho (2019) is where it comes from.

Merging results from several searches

When you run keyword search and embedding search together, you get two lists scored on scales that have nothing to do with each other. A reranker scores every candidate on one scale, which is a cleaner way to merge than tuning weights by hand.

Filtering, when the score is calibrated on your data

A reranker's score separates relevant from irrelevant much better than a cosine does, so a threshold on it can work. It still has to be fitted on your own labelled pairs, the way the Jev article's calibration figure shows.

How it works

How it works, and what each pair costs

It is a cross-encoder with a scoring head

A reranker is an encoder that takes the query and the document as one input, with a separator between them, and puts a small output layer on the final representation that produces one number. bge-reranker-v2-m3 is built on an XLM-RoBERTa encoder, which is the same family as the models in the BERT article, which produces a single score where a classifier would produce a list of labels.

Reading both together is what buys the accuracy. The model can attend from a word in the question to a word in the document at every layer, so "labelled examples" in the query can be matched against "labeled images" in the passage and marked as the wrong kind of label.

Pointwise, pairwise and listwise rerankers

Everything on this page so far describes a pointwise reranker, which scores one query and one document at a time and sorts by that score. It is the design in Nogueira and Cho (2019), in bge-reranker-v2-m3 and in Cohere Rerank, and it is what most production systems run.

A pairwise reranker reads the query with two documents and says which of the two is more relevant. Pradeep, Nogueira and Lin (2021) put one after a pointwise stage, as monoT5 then duoT5, because comparing every pair costs a number of calls that grows with the square of the shortlist, so it only runs on the top handful.

A listwise reranker reads the query and the whole shortlist in one context and returns an order. Sun et al. (2023) did this by asking an LLM to output a permutation of the passage numbers. jina-reranker-v3 is listwise too, and its card says it takes up to 64 documents in one 131K-token context, so the documents can be compared with each other as well as with the query.

A reranker built on an LLM is not automatically listwise. Qwen3-Reranker reads one query and one document, and its card turns the probability of the model answering "yes" against "no" into the score, which makes it pointwise. The cost section below is about the pointwise case, where the work grows with the number of pairs.

Nothing can be precomputed, and that sets the cost

An embedding model runs once per document, ever. A reranker runs once per query-document pair, because the score depends on both halves and neither exists without the other. Every new query brings new pairs, so the scoring waits until the query arrives.

On the run behind the figure, scoring 10 pairs took 1,206 milliseconds on a laptop CPU, which averages to 121 milliseconds per pair. Those pairs were fed to the model two at a time, cut to 320 tokens, on four CPU threads. Embedding a passage with the retriever took 14 milliseconds, and that happened once when the index was built.

Batching changes the per-pair number, so a second run kept every other setting and varied the batch size and the shortlist depth. The machine was busier during that run and everything came out slower, including the 10 pairs in batches of 2, so compare its rows with each other only.

Reranking one query on a laptop CPU, second run, mean over the eight queries
ShortlistBatch sizeTime per queryPer pair
Top 1012.13 s213 ms
Top 1021.98 s198 ms
Top 10101.91 s191 ms
Top 50111.39 s228 ms
Top 50108.39 s168 ms
Top 505010.06 s201 ms

On a CPU, batching saved between 10 and 26 percent, and the largest batch was not the fastest. A likely reason is padding, since every pair in a batch is padded to the length of the longest one, and a bigger batch pads more short pairs. The cost still grew with the shortlist, with 50 pairs taking 4.4 times as long as 10 at the best batch size on each. At that rate all 287 passages would take around 50 seconds a query on this machine, which is an estimate from the table and was not run.

A GPU changes the picture more than any batch size does on a CPU. It runs a batch of pairs in close to the time of one until the batch fills the card, so per-pair cost can fall a long way, and it was not measured here. On any hardware the pairs are real work that no index removes, and reranking everything is always possible and never sensible.

The shortlist depth decides both the cost and the ceiling

On this run the right article was inside the top 10 for seven of the eight questions, so 10 was deep enough for all but one. The one it missed can never be fixed by reranking, which the figure marks.

So measure recall at your shortlist depth before you tune anything else. If the answer is not in the top 50, no reranker will find it, and the fix belongs in the first stage.

Versions

The models you'll see, as of September 2026

Open rerankers, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
BAAI/bge-reranker-v2-m3March 2024Apache-2.017.4M
Qwen/Qwen3-Reranker-4BJune 2025Apache-2.02.5M
jinaai/jina-reranker-v3September 2025CC-BY-NC-4.00.8M
mixedbread-ai/mxbai-rerank-large-v2March 2025Apache-2.00.1M

Check the licence before you pick. jina-reranker-v3 is CC-BY-NC-4.0, which rules out commercial use, where the other three on that list are Apache-2.0.

Hosted rerankers, from vendor docs read 2026-09-21
ModelWhat Cohere says it is for
rerank-v4.0-proMultilingual, built for quality on harder material, taking text and semi-structured JSON
rerank-v4.0-fastThe same model family tuned for low latency and high throughput
rerank-v3.5The previous generation, with a 4,096-token context

The split between a quality tier and a latency tier is the shape of this category now. Reranking sits directly in the path of a user waiting for an answer, so the vendors hand you the choice and let you make it.

Choosing

Choosing, and when to skip the second stage

Which model to reach for
The jobReach forWhy
A first reranker in any languagebge-reranker-v2-m3Apache-2.0, multilingual, and by far the most used, so the integrations exist
The best open accuracy you can self-hostA current Qwen3-Reranker sizeLarger, same licence, and it costs more per pair
A managed reranker with a latency choiceCohere rerank-v4.0-pro or -fastThe pro and fast tiers are the same decision made explicit
Finding the candidates in the first placeAn embedding modelA reranker cannot search, and it only reorders what it is handed
Accuracy between the two, with precomputationColBERT-style late interactionIt keeps a vector per token, so some of the work can happen ahead of time
Deciding a yes or no about one pairA reranker with a fitted threshold, or a fine-tuned encoderBoth read the pair, and both need calibrating on your data

Skip the second stage when the first one is already returning the right passage at the top, which you only know by measuring. On the run behind the figure, search alone got 4 of 8 and the reranker brought it to 6 of 8, and on a corpus where search already gets 7 of 8, the same reranker would add a second of latency for almost nothing.

The bi-encoders and cross-encoders article covers the underlying difference between the two stages, and it is the page to read if this one left you wondering why the fast model cannot be more accurate on its own.

Try it

How to try it

A reranker is a sequence-classification model with one output, so it is three lines to load and one to score.

rerank.py — rescore a shortlist
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "BAAI/bge-reranker-v2-m3"           # Apache-2.0, multilingual
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()

pairs = [[query, passage] for passage in shortlist]
batch = tok(pairs, padding=True, truncation=True,
            max_length=512, return_tensors="pt")
with torch.no_grad():
    scores = model(**batch).logits.view(-1).float()

for i in scores.argsort(descending=True)[:5]:
    print(round(float(scores[i]), 2), shortlist[i][:90])

The code sends the whole shortlist as one batch, which is fine for a few dozen pairs. For longer lists, split it into batches of 8 to 32 and measure, since the table above found a middle size fastest on a CPU. The scores are raw numbers on no fixed scale, so they sort and they do not compare across queries. Put a sigmoid on them if you want something between 0 and 1, and fit the threshold on your own labelled pairs before you trust it.

Ask your AI coding tool

Measure whether a reranker is worth adding to my retrieval system. Read my documents from a folder, split them into passages, and index them with BAAI/bge-small-en-v1.5. Take a CSV of questions with the document that should answer each one. For shortlist depths of 5, 10, 25 and 50, report recall at that depth from the retriever alone, then rerank with BAAI/bge-reranker-v2-m3 and report top-1 accuracy after reranking, plus the added latency per query at each depth. Print a table with one row per depth so I can see where the accuracy stops improving and the latency keeps growing. Then list the questions the reranker fixed and the ones where the right answer never made the shortlist, since those need a better first stage.

Limits

What it can't do

  • It can't search. It only reorders what the first stage handed it, so an answer missing from the shortlist stays missing. One of the eight questions in the figure failed this way.
  • It costs a forward pass per document. Batching shares the passes out more efficiently, and latency still grows with the shortlist. On the CPU here, 50 pairs took 4.4 times as long as 10, and there is no index that avoids it.
  • Only an exact repeated pair can be cached. A cache keyed on the query text and the document text returns the stored score when that same pair comes back, which helps with popular repeated queries. A query that differs by one word is a new pair, so every document on its shortlist goes through the cross-encoder again.
  • Its scores have no fixed scale. They sort candidates for one query and mean nothing across queries until you fit a threshold on labelled pairs of your own.
  • Long documents get truncated. The query and the document share one context window, so a long passage loses its tail, and chunking decisions from the first stage carry straight through.
  • A gain on a public benchmark is not a gain on your corpus. The run on this page moved 4 of 8 to 6 of 8 on one small corpus, and the only number that settles it for you is the same measurement on your own documents and questions.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.