MLGuerrillaStart with M1 →
Embeddings and retrieval·16 min read·Updated 24 September 2026

BM25 and sparse retrieval

BM25 counts how many of your query's rare words a document contains and adjusts for length. It has no idea what any of them mean, and on the corpus behind this article it answered a question the embedding model got wrong.

BM25 is a scoring function. It takes a query and a document as input and returns a number for how well they match, computed from nothing but word counts. It came out of the Okapi retrieval work in the 1990s, and it is the ranking Elasticsearch still uses by default today.

The scoring rests on three observations about words. A word that appears in few documents tells you more than a word that appears in most of them. A document that uses a query word several times is a better match than one that uses it once, though the second use counts for less than the first. And a long document containing a word is less impressive than a short one containing it, so length gets divided out. The formula is those three ideas multiplied together and summed over the query's words.

What it is

It scores the words, and it never learned what they mean

Working through one real query makes the whole thing concrete. This is the same question and the same 287 passages the text embeddings article searches, with BM25 in place of the embedding model.

A figure titled "BM25 scores the words, and it has never heard of meaning", from one question run against 287 passages from this site. The question is "how do I train a classifier on my own labelled examples", and a note says stopwords are dropped and the eight words left are scored one at a time. The passage that won is from /models/bert, 43 words long, scoring 10.72, and reads "The second step is fine-tuning. You take the pretrained model, put one small output layer on top of it, and train it again on your own labeled examples. The BERT paper makes the point that this is ...". The runner-up, from /models/jev, scored 10.22. A table then shows where the score came from, word by word, with how many of the 287 passages contain it, its rarity, how many times it appears in this passage, and what it adds. The word how is in 40 passages, rarity 1.96, appears 0 times, adds 0.00. The word do is in 10 passages, rarity 3.31, appears once, adds 3.28. The word train is in 17 passages, rarity 2.80, appears once, adds 2.77. The word classifier is in 9 passages, rarity 3.41, appears 0 times, adds 0.00. The word my is in 8 passages, rarity 3.52, appears 0 times, adds 0.00. The word own is in 42 passages, rarity 1.91, appears once, adds 1.89. The word labelled, shown in terracotta, is in 0 passages, rarity 0.00, appears 0 times, adds 0.00. The word examples is in 17 passages, rarity 2.80, appears once, adds 2.77. Added up, the total is 10.72, and a note says three of the eight words scored nothing and the passage still won. A final panel headed "the word that mattered most scored zero" explains that labelled appears in 0 of the 287 passages because this site spells it labeled, that BM25 compares strings so the two are unrelated words, that classifier is in 9 passages and in none of this one so it added nothing either, and that the score came from do, train, own and examples. It closes by noting that an embedding model got this same question wrong, and it knows perfectly well that labelled and labeled are the same word. The caption reads: exact matching is the strength and the failure, and they are the same property.
The rarity column is the inverse document frequency, and it is why do at 3.31 outscores own at 1.91 even though both appear once here.

Two things in that run are worth sitting with.

The first is that BM25 got this question right. The embedding model put a YOLO passage about fine-tuning on labelled images first, because both passages are about training on your own labels and one vector has no room to keep them apart. BM25 had no such problem, because "fine-tuning" and "layer" are just different strings from "images" and "boxes".

The second is how it got there. Three of the eight scored words contributed nothing at all, and one of those was "labelled", the word carrying most of the question's meaning. This site spells it "labeled", and with the plain tokenizer this run used, those are two unrelated tokens. The right answer came back anyway, on the strength of four ordinary words. The analyzer section further down shows how a stemmer would have joined the two spellings.

Where it shows up

Where BM25 shows up

The default in the search engine you already have

Elasticsearch ships BM25 as its default similarity, Lucene provides it with k1 at 1.2 and b at 0.75, and SQLite's FTS5 maps its rank column to a built-in bm25(). If you already have a search box, this is likely what is behind it, so the useful question is what to add rather than what to replace.

Names, codes and identifiers

Part numbers, error codes, SKUs, ticket ids and people's names are the case where exact matching wins outright. An embedding blurs exactly the detail those depend on, and a customer searching for ERR_CONN_4021 wants that string and not something nearby.

The other half of hybrid search

Most production retrieval runs BM25 and an embedding model side by side and merges the two rankings. The measurement further down is the reason, and it is larger than most people expect.

A baseline that is embarrassingly hard to beat

BM25 needs no model, no GPU and no training, and it scored 4 of 8 here against the embedding model's 4 of 8. Any retrieval system worth deploying should be measured against it. The ColPali paper's benchmark table has BM25 at 65.1 nDCG@5 against 67.0 for the BGE-M3 embedding model on visually rich documents, with both reading text pulled out by OCR plus generated captions for the figures.

Filtering before the expensive stage

Running BM25 first and embedding only what survives is a cheap way to cut the cost of indexing or reranking a large collection.

How it works

The formula, learned sparse retrieval, and how to merge two rankings

Three terms multiplied together

A document's score is the sum, over the query's words, of three factors.

  • Inverse document frequency. A word in few documents is worth more. In the figure, "my" appears in 8 of the 287 passages and scores 3.52, while "own" appears in 42 and scores 1.91.
  • Saturating term frequency. More occurrences help, with each one helping less than the last. The k1 parameter sets how fast that flattens out, and Elasticsearch's default is 1.2.
  • Length normalisation. A word in a 40-word passage counts for more than the same word in a 400-word one. The b parameter sets how much, and both Elasticsearch and Lucene default it to 0.75.

Everything lives in an inverted index, which is a map from each word to the list of documents containing it. Scoring a query means walking a few short lists rather than touching the collection, which is why BM25 answered in 0.11 milliseconds on this corpus and why it scales to billions of documents on ordinary hardware.

The analyzer decides what counts as a word

BM25 counts tokens, and something has to turn the raw text into tokens first. In Elasticsearch and Lucene that step is called the analyzer, and it runs on every document at indexing time and on every query at search time. An analyzer is a chain of small steps.

  • A tokenizer splits the text into pieces, usually at spaces and punctuation.
  • A lowercase filter makes "BERT" and "bert" the same token.
  • A stopword filter drops words like "the" and "of", if you turn one on.
  • A stemmer cuts words back to a common stem, so "labels", "labelled" and "labeled" all become "label".

Elasticsearch's default standard analyzer does the first two and nothing else, since its stopword list is empty unless you set one. Its english analyzer adds English stopwords and the Porter stemmer. The run behind the figure used a plain tokenizer, lowercasing and a short stopword list, which is close to standard, and that is why "labelled" matched nothing. Under the Porter stemmer both spellings become "label", so the same query would have scored that word against this site's "labeled".

Stemming has its own cost. It joins words that happen to share a stem and mean different things, like "university" and "universe", which Porter reduces to the same token, and it can mangle part numbers and names. The usual answer is two fields over the same text, one stemmed for ordinary words and one left exact for identifiers, searched together.

The analyzer has to be the same on both sides. Stem the documents and not the query, and every inflected query word misses.

Learned sparse retrieval puts a model in front of the same index

The weakness that no analyzer fixes is vocabulary. A stemmer can join "labelled" to "labeled", and nothing in the analyzer can join "car" to "automobile", or a question to a passage that answers it in different words. A synonym list can cover a handful of known pairs, and someone has to write and maintain it.

SPLADE attacks that while keeping the inverted index. It runs a BERT encoder over the text and produces a weight for every word in the model's vocabulary, so a passage about cars ends up with weight on "automobile" and "vehicle" even when those words never appear in it. The naver/splade-v3 card describes the output as "a 30522-dimensional sparse vector space", which is one dimension per token in BERT's vocabulary, with almost all of them zero.

Because the output is still a set of weighted words, it goes into the same inverted index BM25 uses. Formal et al. (2021) put the appeal plainly, saying learned sparse representations "inherit from the desirable properties of bag-of-words models such as the exact matching of terms and the efficiency of inverted indexes". The v3 card reports 40.2 MRR@10 on the MS MARCO dev set and 51.7 average nDCG@10 on BEIR-13.

Fusing two rankings, which is where the measured gain is

When you run keyword search and embedding search together you get two lists whose scores are on scales that have nothing to do with each other. A BM25 score of 10.72 and a cosine of 0.762 cannot be added, averaged or compared.

Reciprocal rank fusion sidesteps that by ignoring the scores and using only the positions. Each list votes for a document with one divided by a constant plus its rank, and the votes are added. The constant is usually 60, and it flattens the difference between first and second place so a document both lists rank highly beats a document one list loves and the other ignores.

A figure titled "Neither list had it first, and fusing them did", from one of the eight questions run through keyword search and embedding search over the same 287 passages. The question is "how do I get a probability instead of generated text". BM25 returned a passage from the encoder-decoder architectures article and bge-small returned one from the T5 article, and a terracotta line notes the right answer is in /models/jev and is neither of those. The next panel, headed reciprocal rank fusion with k at 60, explains that each list votes with one over sixty plus its rank and the two votes are added, so a passage high in both lists beats a passage first in one of them. Four rows follow. The jev passage, outlined in terracotta, was ranked 2 by BM25 and 3 by bge, giving one over 62 plus one over 63, which is 0.03200 and the highest. The encoder-decoder architectures passage was ranked 1 by BM25 and 7 by bge, giving 0.03132. Two separate passages from the large language models article scored 0.03054 and 0.02920. A line says second and third beat first and seventh, so the passage both lists nearly liked is the one that comes back. A final panel lists where the right article ranked on all eight questions. For train a classifier on my own labels, BM25 ranked it 1, bge 2, fused 1. For why can't it predict the next word, 1, 1 and 1. For what makes image generation slow, 1, 2 and 1. For find objects and draw boxes, 3, 1 and 1. For search photos by typing a sentence, 7, 1 and 1. For what attention does to each token, 1, 1 and 1. For get a probability not text, 2, 3 and 1. For what the model does with audio, 32, 28 and 21. The totals are 4 of 8 for BM25, 4 of 8 for bge, and 7 of 8 fused. The caption reads: two mediocre lists that fail on different questions fuse into one that does not.
The two rows both labelled llms are two different passages from the large language models article, which is why their ranks differ.

Four of eight became seven of eight. The question that makes the case is the one in the top panel, where the right passage was second in one list and third in the other, and neither list would have returned it. Fusing them put it first.

That gain is bigger than the reranker bought on the same corpus, which took 4 of 8 to 6 of 8 at 121 milliseconds per pair. Fusion costs a second query against an index you already have.

Why the two methods fail on different questions

The pattern in the rank table is the whole argument for running both. BM25 ranked the right article first on the two questions where the embedding model's single vector blurred a distinction, and it put the right article at rank 7 on "can I search my photos by typing a sentence", where the answering passage never uses the word "photos".

So the failures are not correlated, and uncorrelated failures are exactly what fusion fixes. When two methods are wrong about the same things, merging them gains nothing.

Versions

What to run, as of September 2026

BM25 itself has no model to download. It ships inside the search engine, and the numbers that matter are the parameters.

BM25 parameters, with Elasticsearch defaults read 2026-09-21
ParameterDefaultWhat raising it does
k11.2Lets repeated words keep counting for longer before the score flattens
b0.75Penalises long documents more heavily

The run behind these figures used k1 of 1.5, which is the other common default and comes from the original Okapi work. Both are defensible and neither is tuned, which is the point: BM25 is competitive before anyone touches it.

Learned sparse models are ordinary Hugging Face checkpoints.

Open sparse retrieval models, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
naver/splade-cocondenser-ensembledistilMay 2022CC-BY-NC-SA-4.0330K
naver/splade-v3March 2024CC-BY-NC-SA-4.0110K
opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distillJuly 2024Apache-2.027K
prithivida/Splade_PP_en_v1February 2024Apache-2.025K

Read that licence column before you plan around SPLADE. Every Naver checkpoint, including the newest and best one, is CC-BY-NC-SA-4.0, which rules out commercial use. The OpenSearch and Splade_PP models are Apache-2.0 and exist largely because of that.

Choosing

Choosing, which usually means running two things

Which retrieval to reach for
The situationReach forWhy
You have a search box and no retrieval yetBM25 aloneIt is already in your database, it needs no model, and it is the baseline for everything else
Users search for names, codes or exact stringsBM25, weighted heavilyAn embedding blurs the one detail those queries depend on
Users ask questions in their own wordsAn embedding modelNothing has to match word for word
You can afford two queriesBM25 fused with embeddings4 of 8 and 4 of 8 became 7 of 8 here, for one extra index lookup
You want vocabulary matching without a second indexSPLADE, if the licence allowsExpanded terms go into the same inverted index
The top few results are in the wrong orderA rerankerFusion fixes recall, and ordering the top 10 is a different job
Recall at stage one is the binding constraintColBERTIt matched the fusion here on its own, at 27.7 times the index
Searching PDFs and slides as imagesColPaliNo amount of BM25 tuning helps when the text was never extracted

Start with BM25 and add embeddings, rather than the other way round. The keyword index is already there, the fusion is twenty lines, and the measurement above says that combination is where most of the accuracy is.

Try it

How to try it

BM25 is short enough to write out, and writing it out once is the fastest way to understand what it does.

bm25.py — the whole scoring function
import math, re
from collections import Counter

def toks(s): return re.findall(r"[a-z0-9]+", s.lower())

docs = [toks(t) for t in passages]
N, avgdl = len(docs), sum(len(d) for d in docs) / len(docs)
df = Counter(w for d in docs for w in set(d))
idf = {w: math.log(1 + (N - n + 0.5) / (n + 0.5)) for w, n in df.items()}
tf = [Counter(d) for d in docs]
k1, b = 1.5, 0.75

def score(query, i):
    dl, total = len(docs[i]), 0.0
    for w in toks(query):
        f = tf[i].get(w, 0)
        if f:
            total += idf[w] * f * (k1 + 1) / (f + k1 * (1 - b + b * dl / avgdl))
    return total

Fusing two rankings is shorter still, and it needs neither list's scores.

fuse.py — reciprocal rank fusion
def rrf(*rankings, k=60):
    """Each ranking is a list of document ids, best first."""
    votes = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):
            votes[doc_id] = votes.get(doc_id, 0) + 1 / (k + rank)
    return sorted(votes, key=votes.get, reverse=True)

merged = rrf(bm25_ranking, embedding_ranking)
Ask your AI coding tool

Measure whether hybrid search is worth it on my own documents. Read every Markdown file under a folder I pass in, split it into passages of at least 25 words, and build two retrieval systems over them: BM25 written out by hand with k1 of 1.5 and b of 0.75, and BAAI/bge-small-en-v1.5 with its query prefix. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy for each system separately. Then fuse the two rankings with reciprocal rank fusion at k of 60 and report the same numbers for the fusion. Print, for every question, the rank the correct passage reached in each of the three lists, so I can see which questions each method fails on. Finally, list the questions where BM25 and the embedding model disagree entirely, with what each one returned.

Limits

What it can't do

  • It matches tokens and nothing else. "labelled" scored zero against a corpus that writes "labeled", as the first figure shows, because the analyzer in that run had no stemmer. A stemming analyzer handles plurals and spelling variants like that one, and synonyms, translations and paraphrases stay invisible without a synonym list or learned expansion.
  • It cannot answer a question asked in different words. It put the right article at rank 7 for "can I search my photos by typing a sentence", because the passage that answers it never says "photos".
  • The score is not comparable across queries. A long query scores higher on everything, so the 10.72 in the figure means nothing next to another question's number, and there is no threshold to set.
  • Tuning k1 and b buys little. They matter at the margin, and the gap between BM25 and BM25 plus embeddings is far larger than the gap between any two settings of them.
  • SPLADE is mostly non-commercial. The Naver checkpoints, which are the strongest, are all CC-BY-NC-SA-4.0. Check before you build on one.
  • The measurement here is 287 passages and eight questions. It is enough to show that the two methods fail differently and nowhere near enough to pick parameters. Run it on your own corpus before you trust any of these numbers.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.