What it is
It scores the words, and it never learned what they mean
Working through one real query makes the whole thing concrete. This is the same question and the same 287 passages the text embeddings article searches, with BM25 in place of the embedding model.

Two things in that run are worth sitting with.
The first is that BM25 got this question right. The embedding model put a YOLO passage about fine-tuning on labelled images first, because both passages are about training on your own labels and one vector has no room to keep them apart. BM25 had no such problem, because "fine-tuning" and "layer" are just different strings from "images" and "boxes".
The second is how it got there. Three of the eight scored words contributed nothing at all, and one of those was "labelled", the word carrying most of the question's meaning. This site spells it "labeled", and with the plain tokenizer this run used, those are two unrelated tokens. The right answer came back anyway, on the strength of four ordinary words. The analyzer section further down shows how a stemmer would have joined the two spellings.
Where it shows up
Where BM25 shows up
The default in the search engine you already have
Elasticsearch ships BM25 as its default similarity, Lucene provides it with k1 at 1.2 and b at 0.75, and SQLite's FTS5 maps its rank column to a built-in bm25(). If you already have a search box, this is likely what is behind it, so the useful question is what to add rather than what to replace.
Names, codes and identifiers
Part numbers, error codes, SKUs, ticket ids and people's names are the case where exact matching wins outright. An embedding blurs exactly the detail those depend on, and a customer searching for ERR_CONN_4021 wants that string and not something nearby.
The other half of hybrid search
Most production retrieval runs BM25 and an embedding model side by side and merges the two rankings. The measurement further down is the reason, and it is larger than most people expect.
A baseline that is embarrassingly hard to beat
BM25 needs no model, no GPU and no training, and it scored 4 of 8 here against the embedding model's 4 of 8. Any retrieval system worth deploying should be measured against it. The ColPali paper's benchmark table has BM25 at 65.1 nDCG@5 against 67.0 for the BGE-M3 embedding model on visually rich documents, with both reading text pulled out by OCR plus generated captions for the figures.
Filtering before the expensive stage
Running BM25 first and embedding only what survives is a cheap way to cut the cost of indexing or reranking a large collection.
How it works
The formula, learned sparse retrieval, and how to merge two rankings
Three terms multiplied together
A document's score is the sum, over the query's words, of three factors.
- Inverse document frequency. A word in few documents is worth more. In the figure, "my" appears in 8 of the 287 passages and scores 3.52, while "own" appears in 42 and scores 1.91.
- Saturating term frequency. More occurrences help, with each one helping less than the last. The
k1parameter sets how fast that flattens out, and Elasticsearch's default is 1.2. - Length normalisation. A word in a 40-word passage counts for more than the same word in a 400-word one. The
bparameter sets how much, and both Elasticsearch and Lucene default it to 0.75.
Everything lives in an inverted index, which is a map from each word to the list of documents containing it. Scoring a query means walking a few short lists rather than touching the collection, which is why BM25 answered in 0.11 milliseconds on this corpus and why it scales to billions of documents on ordinary hardware.
The analyzer decides what counts as a word
BM25 counts tokens, and something has to turn the raw text into tokens first. In Elasticsearch and Lucene that step is called the analyzer, and it runs on every document at indexing time and on every query at search time. An analyzer is a chain of small steps.
- A tokenizer splits the text into pieces, usually at spaces and punctuation.
- A lowercase filter makes "BERT" and "bert" the same token.
- A stopword filter drops words like "the" and "of", if you turn one on.
- A stemmer cuts words back to a common stem, so "labels", "labelled" and "labeled" all become "label".
Elasticsearch's default standard analyzer does the first two and nothing else, since its stopword list is empty unless you set one. Its english analyzer adds English stopwords and the Porter stemmer. The run behind the figure used a plain tokenizer, lowercasing and a short stopword list, which is close to standard, and that is why "labelled" matched nothing. Under the Porter stemmer both spellings become "label", so the same query would have scored that word against this site's "labeled".
Stemming has its own cost. It joins words that happen to share a stem and mean different things, like "university" and "universe", which Porter reduces to the same token, and it can mangle part numbers and names. The usual answer is two fields over the same text, one stemmed for ordinary words and one left exact for identifiers, searched together.
The analyzer has to be the same on both sides. Stem the documents and not the query, and every inflected query word misses.
Learned sparse retrieval puts a model in front of the same index
The weakness that no analyzer fixes is vocabulary. A stemmer can join "labelled" to "labeled", and nothing in the analyzer can join "car" to "automobile", or a question to a passage that answers it in different words. A synonym list can cover a handful of known pairs, and someone has to write and maintain it.
SPLADE attacks that while keeping the inverted index. It runs a BERT encoder over the text and produces a weight for every word in the model's vocabulary, so a passage about cars ends up with weight on "automobile" and "vehicle" even when those words never appear in it. The naver/splade-v3 card describes the output as "a 30522-dimensional sparse vector space", which is one dimension per token in BERT's vocabulary, with almost all of them zero.
Because the output is still a set of weighted words, it goes into the same inverted index BM25 uses. Formal et al. (2021) put the appeal plainly, saying learned sparse representations "inherit from the desirable properties of bag-of-words models such as the exact matching of terms and the efficiency of inverted indexes". The v3 card reports 40.2 MRR@10 on the MS MARCO dev set and 51.7 average nDCG@10 on BEIR-13.
Fusing two rankings, which is where the measured gain is
When you run keyword search and embedding search together you get two lists whose scores are on scales that have nothing to do with each other. A BM25 score of 10.72 and a cosine of 0.762 cannot be added, averaged or compared.
Reciprocal rank fusion sidesteps that by ignoring the scores and using only the positions. Each list votes for a document with one divided by a constant plus its rank, and the votes are added. The constant is usually 60, and it flattens the difference between first and second place so a document both lists rank highly beats a document one list loves and the other ignores.

Four of eight became seven of eight. The question that makes the case is the one in the top panel, where the right passage was second in one list and third in the other, and neither list would have returned it. Fusing them put it first.
That gain is bigger than the reranker bought on the same corpus, which took 4 of 8 to 6 of 8 at 121 milliseconds per pair. Fusion costs a second query against an index you already have.
Why the two methods fail on different questions
The pattern in the rank table is the whole argument for running both. BM25 ranked the right article first on the two questions where the embedding model's single vector blurred a distinction, and it put the right article at rank 7 on "can I search my photos by typing a sentence", where the answering passage never uses the word "photos".
So the failures are not correlated, and uncorrelated failures are exactly what fusion fixes. When two methods are wrong about the same things, merging them gains nothing.
Versions
What to run, as of September 2026
BM25 itself has no model to download. It ships inside the search engine, and the numbers that matter are the parameters.
| Parameter | Default | What raising it does |
|---|---|---|
k1 | 1.2 | Lets repeated words keep counting for longer before the score flattens |
b | 0.75 | Penalises long documents more heavily |
The run behind these figures used k1 of 1.5, which is the other common default and comes from the original Okapi work. Both are defensible and neither is tuned, which is the point: BM25 is competitive before anyone touches it.
Learned sparse models are ordinary Hugging Face checkpoints.
| Model | Released | License | Downloads a month |
|---|---|---|---|
naver/splade-cocondenser-ensembledistil | May 2022 | CC-BY-NC-SA-4.0 | 330K |
naver/splade-v3 | March 2024 | CC-BY-NC-SA-4.0 | 110K |
opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distill | July 2024 | Apache-2.0 | 27K |
prithivida/Splade_PP_en_v1 | February 2024 | Apache-2.0 | 25K |
Read that licence column before you plan around SPLADE. Every Naver checkpoint, including the newest and best one, is CC-BY-NC-SA-4.0, which rules out commercial use. The OpenSearch and Splade_PP models are Apache-2.0 and exist largely because of that.
Choosing
Choosing, which usually means running two things
| The situation | Reach for | Why |
|---|---|---|
| You have a search box and no retrieval yet | BM25 alone | It is already in your database, it needs no model, and it is the baseline for everything else |
| Users search for names, codes or exact strings | BM25, weighted heavily | An embedding blurs the one detail those queries depend on |
| Users ask questions in their own words | An embedding model | Nothing has to match word for word |
| You can afford two queries | BM25 fused with embeddings | 4 of 8 and 4 of 8 became 7 of 8 here, for one extra index lookup |
| You want vocabulary matching without a second index | SPLADE, if the licence allows | Expanded terms go into the same inverted index |
| The top few results are in the wrong order | A reranker | Fusion fixes recall, and ordering the top 10 is a different job |
| Recall at stage one is the binding constraint | ColBERT | It matched the fusion here on its own, at 27.7 times the index |
| Searching PDFs and slides as images | ColPali | No amount of BM25 tuning helps when the text was never extracted |
Start with BM25 and add embeddings, rather than the other way round. The keyword index is already there, the fusion is twenty lines, and the measurement above says that combination is where most of the accuracy is.
Try it
How to try it
BM25 is short enough to write out, and writing it out once is the fastest way to understand what it does.
import math, re
from collections import Counter
def toks(s): return re.findall(r"[a-z0-9]+", s.lower())
docs = [toks(t) for t in passages]
N, avgdl = len(docs), sum(len(d) for d in docs) / len(docs)
df = Counter(w for d in docs for w in set(d))
idf = {w: math.log(1 + (N - n + 0.5) / (n + 0.5)) for w, n in df.items()}
tf = [Counter(d) for d in docs]
k1, b = 1.5, 0.75
def score(query, i):
dl, total = len(docs[i]), 0.0
for w in toks(query):
f = tf[i].get(w, 0)
if f:
total += idf[w] * f * (k1 + 1) / (f + k1 * (1 - b + b * dl / avgdl))
return totalFusing two rankings is shorter still, and it needs neither list's scores.
def rrf(*rankings, k=60):
"""Each ranking is a list of document ids, best first."""
votes = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
votes[doc_id] = votes.get(doc_id, 0) + 1 / (k + rank)
return sorted(votes, key=votes.get, reverse=True)
merged = rrf(bm25_ranking, embedding_ranking)Measure whether hybrid search is worth it on my own documents. Read every Markdown file under a folder I pass in, split it into passages of at least 25 words, and build two retrieval systems over them: BM25 written out by hand with k1 of 1.5 and b of 0.75, and BAAI/bge-small-en-v1.5 with its query prefix. Take a CSV of questions with the file that should answer each one, and report top-1 and top-5 accuracy for each system separately. Then fuse the two rankings with reciprocal rank fusion at k of 60 and report the same numbers for the fusion. Print, for every question, the rank the correct passage reached in each of the three lists, so I can see which questions each method fails on. Finally, list the questions where BM25 and the embedding model disagree entirely, with what each one returned.
Limits
What it can't do
- It matches tokens and nothing else. "labelled" scored zero against a corpus that writes "labeled", as the first figure shows, because the analyzer in that run had no stemmer. A stemming analyzer handles plurals and spelling variants like that one, and synonyms, translations and paraphrases stay invisible without a synonym list or learned expansion.
- It cannot answer a question asked in different words. It put the right article at rank 7 for "can I search my photos by typing a sentence", because the passage that answers it never says "photos".
- The score is not comparable across queries. A long query scores higher on everything, so the 10.72 in the figure means nothing next to another question's number, and there is no threshold to set.
- Tuning k1 and b buys little. They matter at the margin, and the gap between BM25 and BM25 plus embeddings is far larger than the gap between any two settings of them.
- SPLADE is mostly non-commercial. The Naver checkpoints, which are the strongest, are all CC-BY-NC-SA-4.0. Check before you build on one.
- The measurement here is 287 passages and eight questions. It is enough to show that the two methods fail differently and nowhere near enough to pick parameters. Run it on your own corpus before you trust any of these numbers.
Go deeper
Robertson and Zaragoza (2009): The Probabilistic Relevance Framework, BM25 and Beyond · Formal et al. (2021): SPLADE v2, sparse lexical and expansion model for information retrieval · Cormack et al. (2009): Reciprocal Rank Fusion, which is the merge step measured here · Elasticsearch's similarity settings, where BM25 is the default · Elasticsearch's built-in analyzers, including standard and english
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
