MLGuerrillaStart with M1 →
Vision-language·15 min read·Updated 24 September 2026

Vision-language models (VLMs)

A vision-language model takes images and text as input and writes text back. It is what answers questions about a screenshot, a chart or a photo, and it is the same LLM loop with an image encoder bolted onto the front.

A vision-language model takes one or more images along with text, and writes text back. You hand it a screenshot and ask what the error says, or a chart and ask which quarter fell, and it answers in words.

Inside, it is an LLM with an extra step in front. An image encoder turns the picture into vectors, a small layer reshapes those vectors so the language model can read them, and from there the model is doing exactly what it does with text, which is predicting one token at a time.

What it is

It turns the picture into tokens, then writes an answer about them

The language model never sees pixels. It sees a sequence of vectors that occupy the same slots its text tokens occupy, so the image becomes, in effect, a few hundred extra tokens at the start of the prompt.

That framing explains most of what follows. Image tokens cost the same as text tokens, a higher-resolution image means more of them, and the model's attention treats a region of the photo the way it treats a word.

A five-step figure titled "The picture becomes vectors, and the vectors sit in token slots", from a real pass of the street photo through the SigLIP vision tower that many open VLMs use as their eyes. Step one, cut into patches, shows the photo overlaid with a 14 by 14 grid of 16-pixel squares, one square outlined in terracotta, giving 196 patches. Step two, vision encoder, is a SigLIP box at 93M parameters that reads all 196 patches at once. Step three, one vector per patch, shows two real rows. Patch 1 begins +0.87, -1.90, +2.49, +0.52, -1.89 and patch 98 begins +2.73, +0.62, -0.48, -0.36, +1.26, with a note that there are 194 more and each is 768 numbers long, and that these are what the language model receives and never pixels. Step four, projector, is one small trained layer mapping 768 to the language model's own width, so the vectors now fit the same slots as words. A strip below, step five, what the language model actually receives, shows six chips labelled img followed by a note saying 190 more, then the word chips What, does, the, sign, say and a question mark, all in one row. It notes that the image tokens and the word tokens go through the same layers and the model writes its answer from the end of the sequence, and that a patch of pavement occupies a slot exactly the way the word sign does. A final panel says that Qwen2.5-VL, given the same photo, cuts it into 104 by 138 patches, which is 14,352, merges each 2 by 2 block into one token, and so receives that one photo as 3,588 tokens before any question is typed. The caption reads: resolution and cost are the same dial, because a bigger picture is more tokens in the prompt.
The vector values in step three are the real output of the encoder on this photo. Nothing in the numbers says image. The language model tells them apart from word vectors by what it learned in training and by where they sit in the sequence, and it processes both through the same layers.
A figure titled "It describes the scene correctly and misreads the small signs", from Qwen2.5-VL-3B on the street photo the vision articles use, asked four questions at temperature 0. On the left is the photo, a crosswalk in Amsterdam with people walking across it, shops behind them, 1920 by 1445 pixels. On the right are four question and answer cards. Asked what is happening in this photo, it answered "People are walking across a crosswalk in a city", marked correct. Asked exactly how many people are in the photo, it answered 7, marked correct, with a note that YOLO found 7 too. Asked whether there is any text visible and what it says, it answered with a list, RAVANAA BHAVAN, FOU FOW RAMER, Pasta e Pizza, SARAYANAA BHAVAN and FOU FOW RAMER again, and this card is outlined in terracotta and marked "two signs misread". Asked whether the bicycle is left or right of the stroller, it answered "Right", with a note that the boxes overlap so neither is true. A panel below, titled "what the signs actually say", shows a crop of the shopfronts and lists the truth. The small hanging sign SARAVANAA BHAVAN was read as SARAYANAA BHAVAN. The round sign FOU FOW RAMEN was read as FOU FOW RAMER, twice. RAVANAA BHAVAN matches the window sign, where a pole hides the SA, and Pasta e Pizza was read exactly. The caption reads: the wrong sign readings come in the same tone as the right ones.
The same model, the same photo, four questions in one session. Nothing in the wording of the wrong answers signals that they are less reliable than the right ones.

The scene description is right and the count matches what YOLO found on the same photo. It misread two of the signs. It wrote SARAYANAA for the small hanging sign that says SARAVANAA, and turned RAMEN into RAMER twice, in the same even tone as the answers that were correct. Its RAVANAA BHAVAN looks like a third mistake until you check the photo, where a lamp post covers the first two letters of the window sign, so that answer matches what the pixels show.

A rerun of the same question with the same model file on 2026-09-24 read RAMEN correctly and then repeated "FOU FOW RAMEN" dozens of times in a loop. Temperature 0 kept each run on one path, and the two runs still disagreed, so a single run is one sample of what the model does.

Where it shows up

Where VLMs show up

Answering questions about documents and screenshots

Reading an invoice, pulling a total off a receipt, or telling you what an error dialog says. This is the largest commercial use, and it is also where the failure above bites, so the accuracy has to be measured on your own documents.

Describing images at scale

Generating alt text, captioning a product catalogue, or writing the descriptions that a text embedding model then indexes for search.

Charts, diagrams and tables

A VLM will read a bar chart and answer questions about it, which is genuinely useful and genuinely unreliable at the level of exact values.

Agents that see a screen

A model driving a browser or an app needs to read the interface it is acting on. The screenshot goes in, and the next action comes out as text.

Moderation and classification with a written rule

You describe what you are looking for in words, with no labelled examples to collect, which is the same argument Grounding DINO makes for detection.

How it works

Three parts, and one of them is doing the talking

A vision encoder turns the image into vectors

The encoder is usually a Vision Transformer, which cuts the image into square patches and treats each patch the way a language model treats a word. Many open VLMs use a SigLIP encoder for this, which the SigLIP article covers, because it was trained to align images with text in the first place.

A projector reshapes those vectors for the language model

The encoder's vectors are the wrong size and the wrong shape for the language model's input, so a small trained layer maps them across. Liu et al. (2023) showed that this layer can be a single linear projection and still work, which is what made the LLaVA recipe spread so fast.

The language model does everything after that

Once the image vectors sit in the input sequence, the model generates text one token at a time exactly as it would from a text-only prompt. Everything the LLM article says about sampling, context and cost applies unchanged.

Flamingo (Alayrac et al., 2022) is the earlier design worth knowing, where a frozen language model receives image information through added cross-attention layers. The projector approach won on simplicity, and both are still around.

Resolution is a token count, which makes it a cost

A larger image produces more patches, and more patches means more tokens in the context. Qwen2-VL handles images at their native resolution, where earlier models forced every image into a fixed square, which helps on documents where the small text is the point.

This is the setting to check first when a VLM misreads a document. If the image was downscaled before it reached the encoder, the letters were gone before the model saw them.

Diagnosing a misread, one step at a time

When a VLM gets the text in an image wrong, the cause can sit in the photo or in the resizing before the model ever runs, or in the model itself, and each one has a different fix. These steps separate them, and the example is the sign misreads above.

  1. 01Check the pixels first. Zoom into the source image at the spot the model misread. Its RAVANAA BHAVAN turned out to be correct, because a pole hides the SA. The hanging sign that says SARAVANAA is 75 pixels wide in a 1920-pixel photo, with letters a few pixels tall, so a person zooming in can barely read it either.
  2. 02Find out what size the model received. Runtimes often resize images before the encoder. Log the prompt token count the server reports, because it tracks the image size the encoder saw. Ollama reported 3,626 prompt tokens for the full photo, close to the 3,588 image tokens the Qwen2.5-VL processor produces for it.
  3. 03Crop the region and ask again. A crop spends the model's image tokens on the region that holds the answer. Here it didn't help. The hanging-sign crop came back as SARAVAN BHAVAS, and the ramen sign as TOUICOW DAIKAIWA.
  4. 04Upscale the crop to see whether size is the limit. The same crops at 4 times the size gave SARAVAN BHAVAS again and IOWA, and Ollama reported the same prompt token count as for the small crops, about 1,060 to 1,090, which suggests it resized both to the same input. Upscaling adds no detail that wasn't in the source, so when a crop fails at every size, the fix is a sharper source image.
  5. 05Run a dedicated OCR model on the same crop. If OCR reads it and the VLM doesn't, the model is the limit, and the pipeline should read text with OCR and hand the text to the VLM or an LLM.
  6. 06Repeat the question. One run at temperature 0 is one sample, as the rerun above showed. Ask the same question a few times, or across two models, and treat disagreement as a sign the answer isn't readable from that image.

This is one photo and one small model, so it doesn't say how often cropping helps in general. On documents scanned at a readable size, cropping to the field in question is a cheap first thing to try, and the steps above tell you whether it worked.

Versions

The models you'll see, as of September 2026

Open VLMs, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
Qwen/Qwen3-VL-8B-InstructOctober 2025Apache-2.020.0M
google/gemma-4-26B-A4B-itMarch 2026Check the card10.1M
Qwen/Qwen3.5-9BFebruary 2026Check the card9.3M
google/gemma-4-31B-itMarch 2026Check the card9.1M
Qwen/Qwen2.5-VL-7B-InstructJanuary 2025Apache-2.06.9M

Something changed in this category during 2026, and that table is where you can see it. The Gemma 4 family and the current Qwen text models all carry Hugging Face's image-text tag, which means the general-purpose open model now takes images by default. A separate "VL" model is becoming the exception.

The closed frontier models from OpenAI, Anthropic and Google all take image input as standard too, which the LLM article covers in its versions table.

Benchmark numbers for these models move constantly and the ones circulating in blog posts are usually stale or wrong. MMMU, DocVQA, ChartQA and OCRBench are the four names you will see. Take any figure from the model's own card or technical report, and treat a number with no source attached as marketing.

Choosing

Choosing, and when a specialist beats a generalist

Which model to reach for
The jobReach forWhy
Questions about an image in free-form textA VLMWriting the answer is what it does
Dense documents, forms and receiptsA document parsing or OCR model first, then a VLM or LLM on the textThe measured failure above is exactly this job
Searching a photo library by descriptionCLIP or SigLIPTwo towers give you an index, and a VLM gives you one answer at a time
Boxes around what a phrase namesGrounding DINOA detector is built to return boxes, and a VLM's boxes need checking on your own images before you rely on them
Counting objects reliablyYOLO or another detectorA detector returns one box per object, so counting becomes arithmetic
One label on a whole image, at high volumeA fine-tuned image classifierFar smaller, far cheaper, and it only answers the one question
Fixed visual features to build onDINOv3It returns vectors, and a VLM returns sentences

The pattern repeats from the rest of this index. A VLM is the most general option and usually the most expensive per answer, so the question is whether you need the generality. When the output is a number or a box, a specialist is smaller, cheaper and easier to check.

Try it

How to try it

Ollama runs a small VLM locally, which is what produced the figure.

ask_image.py — one question about one image
import base64, json, urllib.request

image = base64.b64encode(open("photo.jpg", "rb").read()).decode()
body = {"model": "qwen2.5vl:3b", "stream": False,
        "prompt": "What does the text in this image say?",
        "images": [image], "options": {"temperature": 0}}

req = urllib.request.Request("http://localhost:11434/api/generate",
        data=json.dumps(body).encode(),
        headers={"Content-Type": "application/json"})
print(json.loads(urllib.request.urlopen(req).read())["response"])

Hosted models take the same shape through their own SDKs, with the image sent as base64 or a URL alongside the text.

Ask your AI coding tool

Test whether a VLM is accurate enough for my documents before I build on it. Take a folder of images and a CSV giving the correct answer for one specific question about each one, such as the invoice total or the date. Run each image through a local qwen2.5vl:3b via Ollama and through one hosted model I name, both at temperature 0, and record the answer and the latency. Report exact-match accuracy per model, and list every image where the two models disagree so I can look at those myself. Log the prompt token count each call reports, so I can see what image size the model received. Then repeat the whole run with each image downscaled to half its resolution, and once more on a crop around the answer's region from a box column in the CSV, and report the accuracy for each, so I can see how much of the result depends on resolution.

Limits

What it can't do

  • It reads text in images unreliably, and it sounds the same when it is wrong. In the run above it misread two of the four signs, and a crop of each sign read no better. Use an OCR or document parsing model when the letters are the point, and check the source pixels before blaming the model.
  • Its coordinates are only as good as your evaluation shows. Some VLMs will output boxes, and how close they land depends on the model and the images. Score them against boxes you drew by hand on your own images before relying on them, the same check you would run on a detector.
  • Counting is a guess that sometimes lands. It matched a detector's count of 7 on this photo, which is not a method you can rely on.
  • Spatial questions get confident answers whether or not they are answerable. Asked whether the bicycle was left or right of the stroller, it said "Right", where the two objects overlap horizontally in the frame.
  • Resolution silently decides what it can see. An image downscaled before the encoder loses the small text, and nothing in the answer tells you that happened.
  • Images cost tokens. A large image can take more context than the question and the answer combined, so cost and resolution are the same dial.
  • Benchmarks are not your documents. MMMU and DocVQA say almost nothing about your invoices, and the only number that settles it is the one you measure on your own files.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.