What it is
It rescores a shortlist, and it moves the top of it
The mechanism is what makes it accurate. An embedding model has to turn the document into a vector before it has seen your query, so the vector has to be good for every question anyone might ask. A reranker sees both at once, so it can judge this document against this question and nothing else.
Cohere's own documentation describes the input plainly, saying Rerank "combines the tokens from the query with the tokens from the document and the combined total counts toward the context limit for a single document". You can send many documents in one request, and each one is still read against the query on its own and given its own score.
Running both stages over the same corpus and the same eight questions as the embeddings article shows what the second one is worth.

Two of the eight questions changed from a wrong article to the right one. The first is the near miss from the embeddings article, where "how do I train a classifier on my own labelled examples" landed on a YOLO passage about fine-tuning on labelled images. Reading the question and the passage together is enough to tell that the question is about text.
Where it shows up
Where rerankers show up
The second stage of RAG
The passages an LLM receives decide what it can answer with. Retrieving 50 and reranking down to 5 puts better passages in front of the model without making the prompt longer, and a shorter prompt with better passages usually beats a longer one.
Search results a person reads
Order counts for more when a human is scanning, because almost nobody looks past the first few results. This is the job the design was built for, and Nogueira and Cho (2019) is where it comes from.
Merging results from several searches
When you run keyword search and embedding search together, you get two lists scored on scales that have nothing to do with each other. A reranker scores every candidate on one scale, which is a cleaner way to merge than tuning weights by hand.
Filtering, when the score is calibrated on your data
A reranker's score separates relevant from irrelevant much better than a cosine does, so a threshold on it can work. It still has to be fitted on your own labelled pairs, the way the Jev article's calibration figure shows.
How it works
How it works, and what each pair costs
It is a cross-encoder with a scoring head
A reranker is an encoder that takes the query and the document as one input, with a separator between them, and puts a small output layer on the final representation that produces one number. bge-reranker-v2-m3 is built on an XLM-RoBERTa encoder, which is the same family as the models in the BERT article, which produces a single score where a classifier would produce a list of labels.
Reading both together is what buys the accuracy. The model can attend from a word in the question to a word in the document at every layer, so "labelled examples" in the query can be matched against "labeled images" in the passage and marked as the wrong kind of label.
Pointwise, pairwise and listwise rerankers
Everything on this page so far describes a pointwise reranker, which scores one query and one document at a time and sorts by that score. It is the design in Nogueira and Cho (2019), in bge-reranker-v2-m3 and in Cohere Rerank, and it is what most production systems run.
A pairwise reranker reads the query with two documents and says which of the two is more relevant. Pradeep, Nogueira and Lin (2021) put one after a pointwise stage, as monoT5 then duoT5, because comparing every pair costs a number of calls that grows with the square of the shortlist, so it only runs on the top handful.
A listwise reranker reads the query and the whole shortlist in one context and returns an order. Sun et al. (2023) did this by asking an LLM to output a permutation of the passage numbers. jina-reranker-v3 is listwise too, and its card says it takes up to 64 documents in one 131K-token context, so the documents can be compared with each other as well as with the query.
A reranker built on an LLM is not automatically listwise. Qwen3-Reranker reads one query and one document, and its card turns the probability of the model answering "yes" against "no" into the score, which makes it pointwise. The cost section below is about the pointwise case, where the work grows with the number of pairs.
Nothing can be precomputed, and that sets the cost
An embedding model runs once per document, ever. A reranker runs once per query-document pair, because the score depends on both halves and neither exists without the other. Every new query brings new pairs, so the scoring waits until the query arrives.
On the run behind the figure, scoring 10 pairs took 1,206 milliseconds on a laptop CPU, which averages to 121 milliseconds per pair. Those pairs were fed to the model two at a time, cut to 320 tokens, on four CPU threads. Embedding a passage with the retriever took 14 milliseconds, and that happened once when the index was built.
Batching changes the per-pair number, so a second run kept every other setting and varied the batch size and the shortlist depth. The machine was busier during that run and everything came out slower, including the 10 pairs in batches of 2, so compare its rows with each other only.
| Shortlist | Batch size | Time per query | Per pair |
|---|---|---|---|
| Top 10 | 1 | 2.13 s | 213 ms |
| Top 10 | 2 | 1.98 s | 198 ms |
| Top 10 | 10 | 1.91 s | 191 ms |
| Top 50 | 1 | 11.39 s | 228 ms |
| Top 50 | 10 | 8.39 s | 168 ms |
| Top 50 | 50 | 10.06 s | 201 ms |
On a CPU, batching saved between 10 and 26 percent, and the largest batch was not the fastest. A likely reason is padding, since every pair in a batch is padded to the length of the longest one, and a bigger batch pads more short pairs. The cost still grew with the shortlist, with 50 pairs taking 4.4 times as long as 10 at the best batch size on each. At that rate all 287 passages would take around 50 seconds a query on this machine, which is an estimate from the table and was not run.
A GPU changes the picture more than any batch size does on a CPU. It runs a batch of pairs in close to the time of one until the batch fills the card, so per-pair cost can fall a long way, and it was not measured here. On any hardware the pairs are real work that no index removes, and reranking everything is always possible and never sensible.
The shortlist depth decides both the cost and the ceiling
On this run the right article was inside the top 10 for seven of the eight questions, so 10 was deep enough for all but one. The one it missed can never be fixed by reranking, which the figure marks.
So measure recall at your shortlist depth before you tune anything else. If the answer is not in the top 50, no reranker will find it, and the fix belongs in the first stage.
Versions
The models you'll see, as of September 2026
| Model | Released | License | Downloads a month |
|---|---|---|---|
BAAI/bge-reranker-v2-m3 | March 2024 | Apache-2.0 | 17.4M |
Qwen/Qwen3-Reranker-4B | June 2025 | Apache-2.0 | 2.5M |
jinaai/jina-reranker-v3 | September 2025 | CC-BY-NC-4.0 | 0.8M |
mixedbread-ai/mxbai-rerank-large-v2 | March 2025 | Apache-2.0 | 0.1M |
Check the licence before you pick. jina-reranker-v3 is CC-BY-NC-4.0, which rules out commercial use, where the other three on that list are Apache-2.0.
| Model | What Cohere says it is for |
|---|---|
rerank-v4.0-pro | Multilingual, built for quality on harder material, taking text and semi-structured JSON |
rerank-v4.0-fast | The same model family tuned for low latency and high throughput |
rerank-v3.5 | The previous generation, with a 4,096-token context |
The split between a quality tier and a latency tier is the shape of this category now. Reranking sits directly in the path of a user waiting for an answer, so the vendors hand you the choice and let you make it.
Choosing
Choosing, and when to skip the second stage
| The job | Reach for | Why |
|---|---|---|
| A first reranker in any language | bge-reranker-v2-m3 | Apache-2.0, multilingual, and by far the most used, so the integrations exist |
| The best open accuracy you can self-host | A current Qwen3-Reranker size | Larger, same licence, and it costs more per pair |
| A managed reranker with a latency choice | Cohere rerank-v4.0-pro or -fast | The pro and fast tiers are the same decision made explicit |
| Finding the candidates in the first place | An embedding model | A reranker cannot search, and it only reorders what it is handed |
| Accuracy between the two, with precomputation | ColBERT-style late interaction | It keeps a vector per token, so some of the work can happen ahead of time |
| Deciding a yes or no about one pair | A reranker with a fitted threshold, or a fine-tuned encoder | Both read the pair, and both need calibrating on your data |
Skip the second stage when the first one is already returning the right passage at the top, which you only know by measuring. On the run behind the figure, search alone got 4 of 8 and the reranker brought it to 6 of 8, and on a corpus where search already gets 7 of 8, the same reranker would add a second of latency for almost nothing.
The bi-encoders and cross-encoders article covers the underlying difference between the two stages, and it is the page to read if this one left you wondering why the fast model cannot be more accurate on its own.
Try it
How to try it
A reranker is a sequence-classification model with one output, so it is three lines to load and one to score.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "BAAI/bge-reranker-v2-m3" # Apache-2.0, multilingual
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()
pairs = [[query, passage] for passage in shortlist]
batch = tok(pairs, padding=True, truncation=True,
max_length=512, return_tensors="pt")
with torch.no_grad():
scores = model(**batch).logits.view(-1).float()
for i in scores.argsort(descending=True)[:5]:
print(round(float(scores[i]), 2), shortlist[i][:90])The code sends the whole shortlist as one batch, which is fine for a few dozen pairs. For longer lists, split it into batches of 8 to 32 and measure, since the table above found a middle size fastest on a CPU. The scores are raw numbers on no fixed scale, so they sort and they do not compare across queries. Put a sigmoid on them if you want something between 0 and 1, and fit the threshold on your own labelled pairs before you trust it.
Measure whether a reranker is worth adding to my retrieval system. Read my documents from a folder, split them into passages, and index them with BAAI/bge-small-en-v1.5. Take a CSV of questions with the document that should answer each one. For shortlist depths of 5, 10, 25 and 50, report recall at that depth from the retriever alone, then rerank with BAAI/bge-reranker-v2-m3 and report top-1 accuracy after reranking, plus the added latency per query at each depth. Print a table with one row per depth so I can see where the accuracy stops improving and the latency keeps growing. Then list the questions the reranker fixed and the ones where the right answer never made the shortlist, since those need a better first stage.
Limits
What it can't do
- It can't search. It only reorders what the first stage handed it, so an answer missing from the shortlist stays missing. One of the eight questions in the figure failed this way.
- It costs a forward pass per document. Batching shares the passes out more efficiently, and latency still grows with the shortlist. On the CPU here, 50 pairs took 4.4 times as long as 10, and there is no index that avoids it.
- Only an exact repeated pair can be cached. A cache keyed on the query text and the document text returns the stored score when that same pair comes back, which helps with popular repeated queries. A query that differs by one word is a new pair, so every document on its shortlist goes through the cross-encoder again.
- Its scores have no fixed scale. They sort candidates for one query and mean nothing across queries until you fit a threshold on labelled pairs of your own.
- Long documents get truncated. The query and the document share one context window, so a long passage loses its tail, and chunking decisions from the first stage carry straight through.
- A gain on a public benchmark is not a gain on your corpus. The run on this page moved 4 of 8 to 6 of 8 on one small corpus, and the only number that settles it for you is the same measurement on your own documents and questions.
Go deeper
Nogueira and Cho (2019): Passage Re-ranking with BERT · Khattab and Zaharia (2020): ColBERT, late interaction over BERT · Reimers and Gurevych (2019): Sentence-BERT, which exists because cross-encoders are too slow to search with · Cohere: the Rerank model docs · BAAI: the bge-reranker-v2-m3 model card · Pradeep, Nogueira and Lin (2021): the Expando-Mono-Duo design pattern, pointwise then pairwise · Sun et al. (2023): Is ChatGPT Good at Search?, listwise reranking with an LLM
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
