MLGuerrillaStart with M1 →
Embeddings and retrieval·15 min read·Updated 24 September 2026

ColPali

ColPali indexes a document by embedding a picture of each page, then matches your typed question against the picture. Nothing is parsed, chunked or read, and on the benchmark its authors built it scores 81.3 against 65.1 for keyword search.

ColPali is a retrieval model for documents. It takes a picture of each page as input, along with your typed question, and returns the pages most likely to answer it. The page is never converted to text at any point.

That is a large departure from how document search normally works. Faysse et al. (2024) describe the usual pipeline as a parser or OCR system to pull the words out, then a layout model to find the titles and tables and figures, then a chunking strategy, and sometimes a captioning step to describe the pictures in words so an embedding model can index them. Their paper reports that tuning that pipeline helps more than tuning the embedding model does, which is a fair summary of where most document RAG projects spend their time.

What it is

It looks at the page, and the layout survives

The simplest way to see what this means is to index something you can go and look at. The figure renders eight pages of this site to images and searches those images.

A three-step figure titled "It searches a picture of the page, and never reads the text", from eight pages of this site rendered to images at 1000 by 1400 and indexed by vidore/colSmol-256M. Step one shows a thumbnail of the rendered Whisper page of this site, labelled "one page of this site, as a picture", feeding a box marked "a vision-language model, 238M parameters, no OCR, no parser", which produces 871 vectors for this page with 128 numbers in each of them, taking 435.5 KB per page at fp32 and scored with MaxSim as in ColBERT. A line reads: the page was never turned into text, and its headings, its tables, its figures and its code blocks are all just parts of one image. Step two lists eight questions typed as text, with the page each one returned, its score, and where the right page ranked. "How do I train a classifier on my own labels" returned text-embeddings at 13.79 with the right page at rank 3. "What makes image generation slow" returned diffusion-models at 11.96, correct. "How does a model draw boxes around objects in a photo" returned clip at 14.06 with the right page at rank 2. "Can I search my photos by typing a sentence" returned clip at 13.58, correct. "What does the model do with audio" returned whisper at 15.00, correct. "How do I get a probability instead of generated text" returned jev at 13.00, correct. "What is attention doing to each token" returned transformer at 14.77, correct. "How do I turn a paragraph into one vector" returned text-embeddings at 14.17, correct. The right page came first 6 times out of 8. Step three gives what that cost: 52.8 seconds per page to index on a Mac CPU, 435.5 KB per page uncompressed, and 84 milliseconds per query against eight pages, with a note that the indexing number is the one that matters and a GPU is what makes this design practical. The caption reads: nothing here was parsed, chunked or read, and the pictures were compared to the words.
The two misses both returned a page that genuinely discusses the subject, with the expected page second and third. Eight pages is a small enough index that one wrong answer moves the score by 12 points.

Six of the eight questions came back with the right page. The run used colSmol-256M, which is the smallest model in the family and the only one that fits comfortably on a laptop CPU, so treat the accuracy as a sign that the idea works rather than as a measurement of the approach.

The cost panel is the part worth taking seriously. Indexing took 52.8 seconds per page on a CPU, which is the number that decides whether this design is usable on your collection. A GPU changes it by a large factor, and it is still the expensive half.

Where it shows up

Where ColPali shows up

RAG over PDFs that have tables and charts in them

Financial reports, scientific papers, regulatory filings and datasheets all carry meaning in their layout. A parser that turns a table into a run of words loses the columns, and this design never takes that step.

Slide decks

Slides are close to the worst case for a text pipeline, because the text is short, the layout is the argument and half the content is a diagram. They are close to the best case for indexing the image.

Scanned documents and anything OCR handles badly

Handwriting, stamps, old typefaces and photographed pages all degrade OCR before the retrieval model ever sees them. Skipping the step skips the degradation.

Getting a document search system running quickly

The whole ingestion pipeline collapses into rendering pages to images. That is one library call, and it is the reason the paper describes its own approach as "drastically simpler".

Any collection where the figures carry the answer

A question whose answer is in a chart has nothing to match against in an extracted-text index, unless someone captioned the chart. Here the chart is part of the page vector.

How it works

A VLM for the pages and MaxSim for the scoring

The page becomes patches, and every patch becomes a vector

A vision-language model cuts the page image into patches and produces one vector per patch, which is the same step the VLM article's image-to-tokens figure walks through. ColPali keeps all of them and projects each down to 128 numbers.

The query is encoded by the same model's language side, one vector per token. Scoring a page is then ColBERT's MaxSim, where every query vector takes its best match among the page's vectors and those best matches are added up.

So the retrieval machinery is not new. What changed is that the thing being matched is a picture of a page rather than a passage of text.

Late interaction is where most of the accuracy comes from

The paper's ablation makes that unusually clear. On ViDoRe, its own benchmark of visually rich document retrieval, plain SigLIP scores 51.4 average nDCG@5. Fine-tuning it lifts that to 58.6. Putting a language model on top moves it to 58.8. Adding late interaction takes it to 81.3.

For comparison on the same benchmark, BM25 scores 65.1 and the BGE-M3 text embedding model scores 67.0. Those two baselines had a full text pipeline in front of them, with OCR to extract the words and a vision-language model writing captions for the figures, which is the expensive pipeline ColPali skips. Two contrastive image-text models, Jina-CLIP and Nomic-vision, score 17.7 and 12.9, which is the measurement behind the warning in the multimodal embeddings article.

A page costs about a quarter of a megabyte

The paper reports 257.5 KB per page, which is 1,024 image patches plus six extra text tokens for the phrase "Describe the image", 1,030 vectors of 128 numbers, stored at 16 bits each. The run behind the figure came out at 435.5 KB per page because it stored 32-bit floats. Its 871 vectors at the paper's 16 bits would be 217.8 KB, a little smaller than the paper's page because colSmol cuts the page into fewer patches.

Put that next to a text index at the same 16 bits, where a page might be three passages of 384 numbers each, or about 2.3 KB. A factor of roughly a hundred is the trade this design asks you to accept, and ColBERTv2's compression work is what people apply to bring it down.

Querying is cheap, and indexing is not

Section 5.2 of the paper puts query encoding at about 30 milliseconds for ColPali's language model, against about 22 for a text embedding model on a 15-token query, and the late interaction at roughly 1 millisecond per 1,000 pages. Those are small numbers.

Indexing is where the cost sits, because every page goes through a multi-billion-parameter vision-language model once. On the CPU behind the figure that was 52.8 seconds a page with a 256M model, so plan the ingestion around a GPU and around the fact that re-indexing a large collection is a job rather than a script.

Versions

The checkpoints you'll see, as of September 2026

Open ColPali-family models, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
vidore/colqwen2-v1.0November 2024Apache-2.0 backbone, MIT adapters371K
vidore/colqwen2.5-v0.2January 2025MIT317K
vidore/colpali-v1.2August 2024MIT57K
vidore/colpali-v1.3November 2024MIT36K
vidore/colSmol-256MJanuary 2025MIT4.3K

The name on the paper is not the checkpoint people run. The two ColQwen2 models together are pulled more than ten times as often as both ColPali checkpoints, because it is built on Qwen2-VL-2B-Instruct and its card says it "takes dynamic image resolutions in input and does not resize them, changing their aspect ratio as in ColPali".

colSmol-256M is the one to reach for when the model has to run somewhere small. It is two orders of magnitude less used than ColQwen2, and it was the only member of the family that would index a page on the laptop behind this article.

Read the ViDoRe leaderboard before choosing on accuracy. It moves, and the numbers in the 2024 paper are a floor rather than a current ranking.

Choosing

Choosing, and when OCR is still right

Which document retrieval to reach for
The documentsReach forWhy
Plain text, already extracted and cleanAn embedding model, or hybrid searchTwo orders of magnitude cheaper to store and to index
PDFs with tables, charts and layout that mattersColQwen2The layout is part of the vector rather than something a parser threw away
Slide decksColQwen2Short text, heavy layout, and diagrams that carry the point
Scans, handwriting and photographed pagesColPali-familyIt skips the OCR step that was degrading everything downstream
You need the text itself, not just the right pageAn OCR or document parsing modelColPali returns a page image and no words
Millions of pages, on a budgetExtracted text first, this secondThe index is the constraint, and compression only takes back some of it
It has to run on a laptop or a phonecolSmol-256M, measured carefullyIndexing is the slow half, and it is slow enough to notice
Exact figures, codes or identifiersBM25 over extracted textExact string matching is a thing this design does not do

Run both and compare on your own documents. The ViDoRe gap is large on visually rich material and it narrows to nothing on clean text, and which of those you have is the only question that decides this.

Try it

How to try it

Rendering pages to images is the whole ingestion pipeline, and then the model does the rest.

index_pages.py — PDF pages in, retrieved page out
from pdf2image import convert_from_path
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("vidore/colqwen2-v1.0")     # Apache-2.0
pages = convert_from_path("report.pdf", dpi=150)       # one image per page

doc_vecs = model.encode_document(pages, batch_size=4)  # the whole index
q_vecs = model.encode_query(["what was revenue in Q3"])

scores = model.similarity(q_vecs, doc_vecs)            # MaxSim
best = int(scores[0].argmax())
pages[best].save("answer_page.png")                    # a page, not a paragraph

What comes back is a page image, so something still has to read it. Handing that page to a vision-language model with the original question is the normal second half of the system.

From a PDF to an answer that cites its page

The full system has four steps. The PDF becomes page images, ColQwen2 picks the pages that match the question, a vision-language model reads those pages and writes the answer, and the answer carries the page numbers it came from so a person can check it. The code below extends the one above to do all four, keeping each page's number from the moment it is rendered.

answer_with_pages.py — PDF in, answer with page citations out
import re
from pdf2image import convert_from_path
from sentence_transformers import MultiVectorEncoder
from transformers import pipeline

retriever = MultiVectorEncoder("vidore/colqwen2-v1.0")
reader = pipeline("image-text-to-text", model="Qwen/Qwen2.5-VL-3B-Instruct")

pages = convert_from_path("report.pdf", dpi=150)
page_no = list(range(1, len(pages) + 1))                   # PDF page numbers, from 1
doc_vecs = retriever.encode_document(pages, batch_size=4)  # index once, keep it

def answer(question, k=2):
    q_vecs = retriever.encode_query([question])
    scores = retriever.similarity(q_vecs, doc_vecs)[0]
    top = scores.argsort(descending=True)[:k].tolist()
    content = []
    for i in top:
        content.append({"type": "text", "text": f"Page {page_no[i]}:"})
        content.append({"type": "image", "image": pages[i]})
    content.append({"type": "text", "text":
        f"Question: {question}\n"
        "Answer only from the pages above. After each fact, cite its page "
        "as (p. N). If the pages do not contain the answer, say so."})
    out = reader(text=[{"role": "user", "content": content}],
                 max_new_tokens=200, return_full_text=False)
    text = out[0]["generated_text"]
    cited = {int(n) for n in re.findall(r"\(p\. (\d+)\)", text)}
    shown = {page_no[i] for i in top}
    return text, sorted(cited), sorted(cited - shown)   # last one should be empty

text, cited, bad = answer("what was revenue in Q3")

The prompt labels each image with its page number, because the model has no other way to know which page it is looking at. A good reply to the revenue question names the figure and ends with a citation such as "(p. 14)", cited holds the pages the answer points to, and bad holds any page it cites that it was never shown, which should be empty. A non-empty bad, or an answer with no citation at all, is a reply to send back or flag.

The check covers only part of the problem. It confirms the model cited a page it was given, and it cannot confirm the model read the number off that page correctly, which is a separate failure for tables and charts. So the interface should put the cited page image next to the answer, where a person can see the number for themselves in a second. Two pages is also a choice. Raising k makes it likelier the answer is in front of the model, and every extra page adds image tokens to the reader's input and time to the reply.

Ask your AI coding tool

Measure whether visual document retrieval beats my text pipeline on my own PDFs. Take a folder of PDFs and build two retrieval systems over them. The first extracts text with a PDF parser, splits it into passages of at least 25 words, and indexes it with BAAI/bge-small-en-v1.5 plus BM25 fused with reciprocal rank fusion at k of 60. The second renders every page to a 150 dpi image and indexes it with vidore/colqwen2-v1.0 through sentence-transformers. Take a CSV of questions with the page number that answers each one, and report top-1 and top-5 page accuracy for both systems, plus indexing time per page, query time and index size in megabytes. Then split the results by whether the answering page contains a table or a figure, so I can see where the visual system earns its cost.

Limits

What it can't do

  • It returns a page, and not an answer. There is no text in the output, so a reading step has to follow, usually a vision-language model call on the retrieved pages. A page citation from that step shows where the answer came from, and a person still has to look at the page to confirm it was read correctly.
  • Indexing is expensive. 52.8 seconds a page on a laptop CPU here, and a real vision-language model forward pass per page on any hardware. Re-indexing a large collection is a scheduled job.
  • The index is large. The paper reports 257.5 KB a page at 16 bits, and the run here came out at 435.5 KB at 32 bits, against a few kilobytes for extracted text. Compression helps and does not close it.
  • It does not do exact matching. A part number or a case reference is the thing BM25 is best at and this is worst at, so a hybrid setup is still the sensible shape.
  • Page granularity is coarse. A dense page of text has one vector set for the whole thing, so a long page that mentions your subject once competes on the same footing as one that is about it.
  • The measurement here is eight pages and eight questions on the smallest model. It shows the mechanism working. Use the ViDoRe leaderboard and your own documents for anything you plan around.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.