MLGuerrillaStart with M1 →
Free · in beta·intermediate·M20·60 min read·Prereq: The Model as a Component (M4). Tool Use (M18) covers the tool-call loop the agent in this module runs, and RAG & Retrieval, Part 1 (M19-1) covers the image embeddings and similarity scores used to check and rank candidates.

Visual Grounding & VLMs

The capability

What visual grounding is

As you already know from M4, a vision-language model, a VLM, takes an image and some text and writes text back. Ask one "is there a horse in this picture?" and it'll say yes. For years that was about as far as it went, because if you then asked it to draw a box around the horse, the box would land somewhere near the horse, or on the wrong side of the picture, or the model would say it couldn't give positions at all. It could recognise the horse and it couldn't say where it was.

Visual grounding is the second half. The input is an image plus a text description of something in it, like "the horse" or "the Render menu" or "every icon", and the output is where that thing is, as a bounding box, which is four numbers giving the left, top, right and bottom edges of a rectangle around it, or as a single point. The description is what makes it grounding. A classic object detector finds the categories it was trained on, like "person" or "car", while a grounding model finds whatever the text names, including things nobody trained it on by name.

Whether you want a point or a box depends on what happens next. An agent that clicks only needs a point, and many computer-use tools ask for exactly that. A box carries more, though. It tells you how big the target is, which is what you crop around when you want a closer look, and it's what you compare against a correct box when you measure accuracy. Most of this module asks for boxes and clicks their centre, because the extra two numbers pay for themselves the first time something goes wrong.

Grounding isn't something only VLMs do, either. There are models built for nothing else, grounding models like Grounding DINO, which take an image and a phrase and return boxes, and are small enough to run locally. There are also grounding models trained on screens, which find UI elements from a description. For years those dedicated models were far better at placing boxes than any VLM, and VLMs have only caught up recently. Section 6 compares them, and this module follows the VLM route because that's where most agents now start.

In September 2026 I took a screenshot of Blender, the open-source 3D program, and gave GPT-5.6 Terra one instruction, "draw a red bounding box around every icon and a blue bounding box around every label". No positions went in with it. The model found nearly every icon and piece of text on a crowded professional screen and drew the boxes by editing the screenshot, and to draw them it had to know where each element was first.

A Blender 2.80 window with its splash screen open, annotated by GPT-5.6 Terra. Red boxes sit tightly around icons, including the toolbar tools on the left, the viewport buttons at the top right, the object icons in the outliner and the property tabs on the right edge. Blue boxes sit around text labels, including the menu names File, Edit, Render, Window and Help, the workspace tabs from Layout to Scripting, the splash screen links from New File to Development Fund, the outliner entries Camera, Cube and Light, the Transform fields Location, Rotation and Scale, and the frame numbers along the timeline from 10 to 250.
One instruction and no positions. The Blender interface is open source, and the splash artwork is Blender Studio's Spring.

Grounding still can't read your intent. Look at the outliner in the top right of that screenshot. Collection, Camera, Cube and Light each have an eye icon beside them, and the four eyes are pixel-for-pixel the same. Ask for "the eye icon" and the model picks one, and nothing in the box tells you it had a choice. That problem is solvable, and a lot of my own work went into it. What worked was building a good harness around the model, using image-to-image encoders and multimodal encoders to match candidates, with VLMs as part of the loop too, and section 7 walks through it. The model also can't see what isn't in the image, like a menu item that only appears after a hover, and its boxes are estimates that can be a few pixels off, which is a lot when the target is a 12-pixel icon. And the model only knows the image it was sent. If your code resized the screenshot on the way, the numbers describe the resized image, which section 4 is about.

Some places it shows up

  • Computer-use agents that operate an application from screenshots, finding the button to click next. This is the running example in this module, and the loop M18 described for tools, with a click as the tool.
  • UI test automation that checks a screen looks right and clicks through it, without depending on element ids that change between releases.
  • Document work, like finding where a field sits on a scanned form so the value can be read or redacted, which M23 builds on.
  • Robots and cameras, where "the red bin on the left" has to become a position an arm can reach.

These are examples, and grounding shows up anywhere a program needs a position from a description. The inputs and outputs are the same in all of them, an image and a description in, a box or a point out, and the same things go wrong, mostly in how the numbers are read back.

The running example

Northlight's agent clicked beside every button

Northlight is a small animation studio. It's invented, and so are its numbers. Its artists spend part of every week on the same Blender chores, opening a scene, setting the render resolution and frame range, exporting, and saving under a naming scheme. So an engineer built an agent to do them. The agent has no access to Blender's code. It works the way a person does, by looking at a screenshot of the current screen, deciding what to click, clicking and looking again.

The first version, built in the spring on the models available then, did everything right except find things. Asked for the position of "the Render menu", the model answered with a box that was near the top of the screen and nowhere near the word Render. On small icons the boxes were mostly guesses. The agent logged its reasoning, and the reasoning was fine, since it knew it wanted the Render menu and why. It just couldn't say where the Render menu was.

The second version switched to a newer model and got good boxes back, and the clicks still missed, every one of them by the same pattern. A click meant for a button near the middle of the screen landed far down and to the right of it, and a click near the top left landed close to its target. That pattern is a coordinate bug, a factor of two between the screenshot and the screen, and section 4 finds it. The team fixed it, and then switched providers for a price test and every box came back transposed, with x and y swapped. The model was fine both times, and the code reading its numbers was wrong.

The third version, on GPT-5.6 and then GPT-6, finds nearly everything Northlight asks for, including the small icons. Its remaining failures took the longest to fix. Asked to hide the light before a test render, it clicked an eye icon in the outliner, and on some scenes it was the camera's eye. Asked for "Release Notes" on the splash screen, it hit the link in the wrong column about half the time, since the splash shows two of them. And a two-minute export task that worked on most days failed often enough that the artists stopped trusting it, even though each single click was right almost every time. The rest of this module is the work of fixing those, which means converting coordinates correctly, telling apart targets that look identical, deciding when to trust a prediction, measuring accuracy on Northlight's own screens, and knowing when a cheaper component would do the job.

How the model sees an image

How a VLM turns a screenshot into something a language model can read

A language model works on tokens, as M3 showed, and its core is a decoder, a transformer that reads a sequence of token vectors and predicts the next token. Nothing in that design knows about pixels. A VLM gets around this by turning the image into vectors of the same size as the language model's token vectors, so the decoder can read image and text in one sequence. Most VLMs you'll meet are built from three parts, and it's worth being able to name them, because each one explains a different grounding failure.

Four cards left to right with a token strip below. The screenshot card shows a grid of patches and notes that the image is resized and cut into patches of 14 to 32 pixels, and that a 1440 by 900 screenshot at 32 pixels is 1,305 patches. The vision encoder card, a Vision Transformer pretrained on image and caption pairs, shows each patch becoming a feature vector, such as patch 1 becoming 0.12, minus 0.40 and so on, one vector per patch, each aware of its neighbours. The highlighted projector card, also called a connector or adapter, maps an encoder vector to a language-model token vector, and notes that it was a projection matrix in the original LLaVA, and that some designs merge 2 by 2 patches into one token or compress to 32 tokens. The language model card says visual tokens sit next to the prompt's text tokens and attention links the words to the patches they describe. The strip below shows the decoder's input, several image tokens followed by the text tokens the, eye, icon, in, the, Light, row, and then its output, a JSON box with x_min 1613, y_min 157, x_max 1633, y_max 175.
The box at the end is generated one token at a time, like any other text, which is why a model never trained to write boxes writes plausible numbers that mean nothing.

The vision encoder is usually a Vision Transformer, which cuts the image into a grid of patches and treats each patch the way a text transformer treats a word piece (Dosovitskiy et al., 2020). Every layer lets each patch look at every other patch, so by the end, the vector for a patch holding half an icon also carries information about the whole icon and the label beside it. Those vectors are the visual features. The encoder is typically pretrained on huge numbers of image and caption pairs, the way CLIP and SigLIP are, which is why its features line up with words well enough for a language model to use them.

The projector, also called a connector or adapter, is the small piece that makes a text-only decoder multimodal. The encoder's vectors have a different size and meaning from the decoder's token vectors, so the projector maps one into the other. It can be surprisingly simple. The original LLaVA connected a CLIP ViT-L/14 encoder to the Vicuna language model "using a simple projection matrix", and in its first training stage "only the projection matrix is updated", with the encoder and the language model left as they were (LLaVA project page). Other designs compress harder. BLIP-2's Q-Former uses 32 learned query vectors that pull what they need from the encoder's features, so every image becomes the same 32 tokens whatever its size (Li et al., 2023).

Once projected, the visual tokens sit in the decoder's input next to the text tokens, and the decoder's attention treats them like any other tokens. When you write "the eye icon beside Light", the words can attend to the patches that hold the word Light and the icon next to it, and that attention is where grounding happens. The answer comes out as text. Some models write coordinates as ordinary digits. PaliGemma added 1,024 special location tokens, <loc0000> to <loc1023>, and writes a box as four of them, y first (Beyer et al., 2024). Either way, a box is something the model generates one token at a time, and a model that was never trained to write boxes will write plausible-looking numbers that mean nothing, which is what the older models did.

Fixed resolution, dynamic resolution and tiles

How the image gets cut up decides how much detail survives, and designs differ most here. Early VLMs resized every image to the one square size their encoder was trained on, a few hundred pixels on a side. A 1440-by-900 screenshot squeezed into that loses most of its small text and icons before the model sees anything, and nothing afterwards can bring them back. Newer designs keep more. Qwen2-VL's dynamic resolution processes "images of varying resolutions into different numbers of visual tokens", so a bigger screenshot becomes more tokens, and then "a simple MLP layer is employed after the ViT to compress adjacent 2×2 tokens into a single token", which is how a 224-by-224 image ends up as 66 tokens (Wang et al., 2024). Another common approach cuts a large image into tiles, encodes each tile at the encoder's native size, and adds a small whole-image thumbnail for context.

This is why architecture decides how well a model grounds small things. Every step that shrinks the image or merges tokens trades detail for cost. A 12-pixel icon might start as one 14-pixel patch, get merged with three neighbours into one token, and share that token with part of a label, so the decoder has to locate something smaller than the unit it reads in. A model with fixed low resolution can't ground that icon at all, a model with dynamic resolution can if you send it enough pixels, and any model does better when you crop, which is section 8's coarse-to-fine design.

What resolution costs, and what the provider does to your image

With hosted models you don't choose the architecture, but you do choose the pixels, and they cost money. OpenAI's docs, as of September 2026, count how many 32-by-32-pixel patches cover the image and bill image tokens from that, with a multiplier of 1.2 for most current models (OpenAI images and vision guide). A 1440-by-900 screenshot needs 45 by 29 patches, 1,305 of them, so about 1,570 image tokens. The same screen from a Mac with a Retina display is 2880 by 1800, four times the patches.

The image the model sees is also often a smaller copy of the one you sent. Each detail setting has a limit, and larger images get scaled down to fit. For GPT-5.5 and similar models, the high setting allows "up to 2,500 patches and a 2048-pixel maximum dimension", so the 2880-by-1800 Retina screenshot, which needs 5,130 patches, gets shrunk. OpenAI's advice for anything coordinate-sensitive is to "resize images to fit those limits before sending them and map returned coordinates back to the original image". Anthropic's computer-use docs go further and reject a screenshot that's too big, because "the API doesn't downscale for you" (Anthropic computer use docs). Either way, the rule is the same. Decide the size of the image yourself, remember the scale factor, and never let a size change happen somewhere you can't see it.

Small targets are where all of this shows up. ScreenSpot-Pro, a benchmark of 1,581 instructions on screenshots from 23 professional applications, has targets that "occupy 0.07% of the screenshot area on average" (Liu et al., 2025). When it came out in April 2025, the best grounding model scored 18.9% and GPT-4o scored under 1%, and the authors' best result, 48.1%, came from a method that searches a smaller region at a time. That's one benchmark from 2025, and the models in section 1 are much better, but the lesson about small targets still holds, which is that cropping to a smaller region and asking again gives the model more pixels per icon.

Coordinates

Ask for numbers in one format, and convert them in your code

A box is four numbers, and there are several ways to write them, and nothing warns you when your code reads one format as another. The run in section 1 drew its boxes onto the image, and when an agent needs to click, it asks for the numbers themselves.

A chat reply explaining coordinate format. It says the model uses pixel coordinates, each box written as x_min, y_min, x_max, y_max measured from the image's top-left corner, and gives an example on a 1440 by 960 image, icon_box equals 12, 68, 55, 108, meaning the box starts 12 pixels from the left and 68 from the top and ends at x 55, y 108. It adds that normalized fractions such as 0.008, 0.071, 0.038, 0.113 are also possible, and that pixels are the most direct for drawing boxes back onto the original screenshot.
The model's own description of the format it used in the Blender run. Other providers use different conventions, which is where agents usually go wrong.
Common box formats, as of September 2026
FormatOrder of the numbersScaleWhere you meet it
Pixel cornersx_min, y_min, x_max, y_maxPixels of the image the model sawThe run above, most detection libraries
Normalized cornersx_min, y_min, x_max, y_max0 to 1, fractions of width and heightMany annotation tools and datasets
Gemini's box_2dymin, xmin, ymax, xmax0 to 1000Gemini's docs (Google)
PaliGemma's location tokensymin, xmin, ymax, xmax0 to 1023, as <loc> tokensPaliGemma and models fine-tuned from it
Centre and sizex_center, y_center, width, heightOften 0 to 1YOLO training labels
A single pointx, yPixels of the screenshotComputer-use tools, like Anthropic's left_click

The y-first rows are the ones that catch people. Read [ymin, xmin, ymax, xmax] as if it were x first and every box comes back reflected across the diagonal, which is what Northlight's price test did. Getting the format wrong costs more than a few boxes. Roboflow found that on GPT-5.6, "Using the wrong coordinate format reduced GPT-5.6 detection performance by around 15 mAP points", out of about 46 (Roboflow, July 2026). So tell the model the format you want in the prompt, ask for JSON with named fields like x_min so the order is written down, and check the reply against a schema, as M4 did for any structured output.

Then convert everything to one format in your own code, in one place, with tests. Three conversions cover most agents.

ground.py — one place for every coordinate conversion
def gemini_box_to_pixels(box_2d, width, height):
    """Gemini's box_2d is [ymin, xmin, ymax, xmax] on a 0-1000 grid."""
    ymin, xmin, ymax, xmax = box_2d
    return (xmin / 1000 * width, ymin / 1000 * height,
            xmax / 1000 * width, ymax / 1000 * height)


def crop_box_to_full(box, crop_origin, crop_scale=1.0):
    """A box found inside a crop, mapped back to the full screenshot's pixels."""
    x0, y0, x1, y1 = box
    ox, oy = crop_origin
    return (ox + x0 / crop_scale, oy + y0 / crop_scale,
            ox + x1 / crop_scale, oy + y1 / crop_scale)


def click_point(box, sent_size, screen_size):
    """Centre of a pixel box in the image the model saw, in the screen's click units."""
    x0, y0, x1, y1 = box
    sx, sy = screen_size[0] / sent_size[0], screen_size[1] / sent_size[1]
    return ((x0 + x1) / 2 * sx, (y0 + y1) / 2 * sy)

The last function is what fixed Northlight's missed clicks. On a Mac with a Retina display, a screenshot has twice as many pixels in each direction as the coordinates the operating system uses for clicks. Anthropic's docs put it plainly, "macOS Retina displays capture screenshots at a device pixel ratio of 2, so the image is twice the resolution of the logical screen coordinates", so "halve the coordinates Claude returns before issuing the click". Northlight's agent sent the full 2880-by-1800 screenshot and clicked at the pixel numbers that came back, on a screen that's 1440 by 900 in click units. Every click landed at twice its intended distance from the top-left corner, which is why the misses grew toward the bottom right. Passing sent_size=(2880, 1800) and screen_size=(1440, 900) halves them, and the same function covers a screenshot your code shrank before sending it. The middle function handles the other place coordinates change frame, which is a crop. If you cut out a region starting at (1360, 60) and enlarged it twice before asking again, a box in the crop maps back with crop_origin=(1360, 60) and crop_scale=2.

Test the conversions with numbers you can check by hand. A box at the exact centre of the image should click the exact centre of the screen, and a Gemini box of [0, 0, 500, 1000] on a 1440-by-900 image is the top half, (0, 0, 1440, 450). Those two tests would have caught both of Northlight's bugs before any click happened.

The agent loop

Screenshot, box, click, then look again

A computer-use agent repeats one loop. It takes a screenshot of the application as it is right now, asks the model where the next target is, converts the box to a click point, clicks, and takes another screenshot. The screenshot is the agent's only view of the application's state, so each loop starts from a fresh one, and a box from an old screenshot is only good until something on the screen moves.

This is M18's tool loop with the screen as the tool. The model decides what to do next and says where, and your code does the clicking. Anthropic's computer-use tool, for example, offers actions like screenshot, left_click with a coordinate, type and zoom, and its coordinates are "in the pixel space of the full-display screenshots you return, with the origin at the top left".

The step people skip is the last one, looking again. A click can land exactly where the model said and still do the wrong thing, because the box was on the wrong element, or a dialog opened between the screenshot and the click, or the click worked and the next screen took two seconds to draw. So after every click, the agent checks that the screen changed the way it expected. After clicking the Render menu, the menu should be open, and a small grounding call on the new screenshot, "is the Render menu's dropdown open?", or a crop compared with the menu's expected look, answers that. If it didn't change, the agent retries once from a fresh screenshot and then stops and reports, with the screenshots attached, which is M9's stop rule applied to clicks.

Ask for the target by what it says or looks like, "the Render menu", never by position, "the fourth menu", because positions change when a window is resized. And when a target is small or the screen is crowded, crop the region around the model's first answer, send the crop at full resolution, ask again, and map the answer back through the crop's offset with crop_box_to_full. Anthropic's tool has a zoom action for exactly this, and after a zoom "Claude still expresses coordinates in the full screenshot's space".

Ask your AI coding tool

Write a Python click loop for a desktop app agent. Each step takes a screenshot, records its size in pixels and the screen's size in click units, and asks the vision model for the target's box as JSON with fields x_min, y_min, x_max, y_max in pixels of the image sent. Validate the JSON against a schema, convert the box centre to screen coordinates with a single function that takes the sent image size and the screen size, click, then take a new screenshot and ask the model whether the expected change happened. Retry once from a fresh screenshot on failure, then stop and save both screenshots. Write pytest tests for the conversion function, including a Retina test where the screenshot is 2880x1800 and the screen is 1440x900, and a crop test where a box found in a 2x enlarged crop maps back to the full screenshot.

Choosing the component

Pick a detector for what it was trained on, and a VLM for the rest

A VLM isn't the only thing that can find objects in an image, and for years it was the worst option. Which component fits depends mostly on whether your images look like what it was trained on.

A classic object detector, like the YOLO family, is trained on labeled boxes for a fixed list of categories and finds only those (Redmon et al., 2016). It's small and fast and runs locally for almost nothing per image, and on the categories it was trained on it's accurate. The list of categories is the catch. A detector trained on everyday photos knows people and cars and knows nothing about a Blender viewport button.

An open-vocabulary grounding model, like Grounding DINO, takes a text description and finds matching objects, including categories it was never trained on by name (Liu et al., 2023). It's the closest thing to a VLM's grounding in a local model. It still learned from photos, though, and UI icons are a different kind of image, flat and small and made to be read at a glance, so in my experience Grounding DINO isn't good at UI icons. The same goes for general image encoders like SigLIP and DINO, which weren't trained on much UI.

Then there are models trained on screens. Microsoft's OmniParser fine-tuned a YOLOv8 detector on "interactable regions extracted from popular webpage DOM trees" and pairs it with a captioning model that describes what each detected element does (Lu et al., 2024). It finds UI elements well, and what it returns is a list of boxes, one per element. Deciding which of those boxes is "the Render menu" is still a separate step. One common way to take that step is to draw a numbered mark on each detected box and ask a VLM which number matches the description, a technique called Set-of-Mark prompting (Yang et al., 2023).

Text on screen has its own cheap component, which is OCR. An OCR engine returns every piece of text with its box, exactly, in milliseconds, and for a target that's a word, like the Render menu or a "Release Notes" link, it's often the most reliable way to find it. It can't find an icon, and it can't tell you which of two identical words you meant, but it turns every label on the screen into an anchor you can reason from, which section 7 depends on.

And there are VLMs themselves, which since late 2024 have gone from bad at this to good. Google shipped Gemini 2.0 Flash in December 2024 with bounding-box output and a spatial-understanding demo, and in my experience Gemini was the first VLM whose boxes were good enough to use. Roboflow measured OpenAI's jump in July 2026, from 13.8 mAP@50 for GPT-5.5 to 46.2 for GPT-5.6 Sol on its detection benchmark, where mAP@50 is the standard detection score that counts a box as right when it overlaps the true one by at least half. In my experience GPT-5.6 and the GPT-6 models released in September 2026 ground UI elements well, and they're too new for anyone to know yet where they fail.

Which component to reach for
Your situationStart withWhy
Your images look like the detector's training data, and volume is highA trained detector like YOLOFast, local, nearly free per image, and accurate on what it knows
Everyday objects named by text, run locallyGrounding DINOFinds text-described objects without training, on photos
Every element on a screen, as a listOmniParserTrained on UI elements, and returns boxes you can pick from
The target is a word or labelOCRExact text and boxes in milliseconds
Things detectors weren't trained on, and you don't want to fine-tuneA VLMGrounds almost anything a description names, for a few cents a call

The line between these is moving fast. Fine-tuning a detector on your own UI used to be the obvious choice for accuracy and speed and cost. VLMs are now cheap enough, and good enough, that for many teams the question is only whether the volume and latency justify training and maintaining a detector. Northlight runs a few hundred agent steps a day, which is nowhere near that point, so it uses a VLM for every step. And as section 8 shows, the components also combine, so the choice is often which ones to chain together.

Use the VLM to label, and train the detector on its labels

You don't always have to choose between the two. A VLM that grounds well can do the slow, expensive part of building a detector, which is drawing the boxes on thousands of training images, and then a small detector trained on those boxes does the fast, cheap part in production. Roboflow, one of the biggest tools for labeling detection datasets, made this the default in September 2026. For object detection projects, its Auto Label feature's "default model for object detection projects is now GPT-6 Astra", which takes the class names you type and draws the boxes zero-shot, with no example boxes and no training (Roboflow, 8 September 2026).

That changes the trade above. Fine-tuning a detector on your own UI used to mean paying people to draw boxes for weeks. Now the labeling costs a VLM call per image, and the detector you train from it keeps the speed and low cost per image that make detectors worth having at high volume. The catch is that the labels are only as good as the model that drew them, and Roboflow points out that "every box comes back at full confidence, nothing is filtered out for you, so this pass is where the quality of the dataset is set." So a person still reviews the labels, ideally with the checks from section 10, before any detector learns from them. For Northlight, if the agent ever runs thousands of steps a minute, this is how it would get a detector for Blender's icons without anyone drawing a box by hand.

Repeated targets

Identical targets need a reference, and pixels can't give you one

Go back to the outliner in the Blender screenshot. It's a list of the scene's objects, one row each, with an eye icon at the right end of every row that shows or hides that object.

Blender's outliner, as the agent sees it
RowLabelIcon at the right
1Collectioneye
2Cameraeye
3Cubeeye
4Lighteye

Northlight's agent had to hide the light before a test render, so the request was "click the eye icon for Light". Recognising an eye icon is easy, and every model since the spring found all four. The hard part is picking the one that belongs to the Light row. That's a different problem called reference resolution, working out which instance of a thing the description refers to, and it needs relational grounding, meaning the answer depends on the target's relationship to something else on the screen, here the word Light in the same row.

This is where an image embedding model, the SigLIP-style encoder from section 10, is honest and useless at the same time. Crop the four eyes, embed them, compare each with a reference eye icon, and all four score almost the same, because they are the same. The encoder is answering "which of these looks like an eye icon?" correctly, and that isn't the question. No amount of visual similarity can pick Light's eye, because the thing that makes it Light's is outside the crop.

What does pick it is context. The row's text is the anchor, and OCR finds "Light" exactly. The relationship is the row, since an icon and a label that share a vertical centre and sit in the same panel belong together, so the eye whose centre is closest to Light's centre on the y axis, and to the right of it, is the one. Hierarchy helps where you can get it. The outliner is a tree, and a nested object's eye sits in its own row under its parent. If the application exposes an accessibility tree, which lists every element with its name and its parent, the eye might even be labeled "Light visibility", and then there's nothing to ground. When none of that is available, a VLM can do the relational reasoning itself, given the crop of the whole panel and a request that names the relationship, "the eye icon in the same row as the label Light".

The Release Notes links on the splash screen are the same problem with text. There are two, one in the "Getting Started" column and one lower down beside the recovery options, and OCR finds both with perfect confidence. The description has to carry the difference, "the Release Notes link under Getting Started", and the harness has to check which one it got. This is the problem I spent the most time on, and what worked was a harness around the model, with encoders to find and filter candidates by look and the VLM to decide among them from the context, so that each part does the job it's good at. Section 8 turns that into an architecture.

Northlight changed two things after the camera incident. Its descriptions now always name the anchor, "the eye icon in the Light row of the outliner", and its test dataset has a slice of repeated targets on purpose, since a model can score well overall while getting every one of them wrong.

Grounding architectures

Three ways to build the grounding step

Sending the whole screenshot and a description to a VLM and clicking what comes back is one design. It's the simplest, and for many screens it's enough, but it isn't the only design, and on crowded professional screens the other two often do better.

Three rows, one per architecture, with local steps drawn on beige and VLM calls on white. Single-shot goes from the screenshot to a VLM on the full image, asked for the eye icon in the Light row, to a box, with one call, weak on small targets and the least to build. Coarse-to-fine goes from a VLM finding the region, the outliner panel, to a local crop and enlarge at full resolution, to a VLM on the crop giving a precise box, to a local map back with crop_box_to_full, with two calls, better on small targets and a little more to build. Candidate-based goes from a local generator such as OmniParser, OCR or the DOM, to a local embedding filter that keeps what looks like the icon, to local rules such as same row as Light, to a highlighted VLM step that chooses a mark among a few real boxes, with one call plus local steps, best on repeated targets and the most to build.
Beige steps run locally in milliseconds, and each white step is a model call you wait seconds for.

Coarse-to-fine, for small targets

Coarse-to-fine grounding runs two passes. The first asks for the region that contains the target, like "the outliner panel", on a screenshot small enough to be cheap. The second cuts that region out of the full-resolution screenshot, enlarges it, and asks for the target inside it, so a 12-pixel eye icon that was a fraction of a token in the first pass now covers several. Then crop_box_to_full maps the answer back. It's the ScreenSpot-Pro lesson from section 3 turned into a design, and it costs a second call and a second wait on every step that uses it, so most agents only take the second pass when the target is small or the first answer looks unsure.

Candidate generation and ranking, for crowded screens

The candidate-based design splits grounding into two questions, where are the things that could be targets, and which one is it. A candidate generator answers the first, and it can be any component that proposes boxes. OmniParser proposes every interactable element. OCR proposes every word. An accessibility tree or a web page's DOM lists elements with names and positions, when the application exposes one. A detector or a segmentation model proposes objects. Then ranking narrows them down, cheaply first. Image-to-image embeddings compare each candidate crop with a reference crop, and image-to-text embeddings compare it with the description. Nearby text from OCR and spatial rules like "same row as Light" score the relationship. Interaction history counts too, since the button that worked last time on this screen is a strong candidate. Only the last few candidates go to a VLM, often as numbered marks in the Set-of-Mark style, with the question the embeddings can't answer.

On the left, four stages with bars showing how many candidates remain. The UI parser proposes every interactable element, 80. The embedding filter keeps crops that look like the reference eye icon, 5. The spatial and text rule keeps those in the outliner beside an OCR'd label, 4. The VLM with numbered marks, asked which mark is in the same row as Light, keeps 1, shown as the only highlighted bar. A note says the counts are illustrative and that the filter keeps a fifth eye from the properties panel, which the rule drops. On the right, a dark outliner panel shows the four candidates, rows Collection, Camera, Cube and Light, each with an identical eye icon and the same embedding similarity of 0.91 to the reference eye, and the Light row's eye outlined as the chosen one.
Four identical scores are the encoder's honest answer, and the reason the last step has to read the row.

Narrowing the search helps in four ways. Accuracy goes up, because the VLM chooses among five marked boxes it can see clearly, where it would otherwise have to write coordinates for one icon among thousands of pixels. Latency and cost go down, since the filters are local and fast and the VLM sees one small request. Reliability goes up, because a candidate's box comes from the parser, so the model can't return a box a few pixels off or in empty space. And debugging gets easier, since you can see which stage dropped the right candidate, the same walk-back M19-1 taught for retrieval.

It's also more to build and more to break. The generator has to propose the target, or no ranking can find it, which is exactly the retrieval problem from M19-1, where a reranker can't recover what the search missed. You now maintain a parser, an encoder, thresholds and a VLM prompt. If your screens are simple, your targets are big or text, and single-shot passes your test dataset, the candidate pipeline is complexity you don't need. Northlight uses single-shot for menus, OCR for text targets, and the candidate pipeline only for icons in panels with repeats, which is where its failures were.

Embedding models and VLMs answer different questions

The hybrid pipeline works because it uses each model for the question it can answer. It's worth saying the boundary plainly, because mixing them up is a common design mistake.

An image encoder and a VLM, side by side
Image encoder, like SigLIP or CLIPVLM
InputAn image, or a text for the paired text encoderAn image and an instruction
OutputA vector, compared with other vectors by cosine similarityGenerated text, including boxes, choices or explanations
Answers"Which candidate crop looks most like this reference icon?""Which eye icon belongs to the row labeled Light?"
Can't answerAnything that depends on context outside the cropAnything, fast and cheaply, a thousand times a second
Cost per comparisonMicroseconds once the vectors exist, and runs locallySeconds and a fraction of a cent per call

An encoder compares looks. A VLM reads a scene and reasons about it. When the question is visual similarity, the encoder is faster, cheaper and deterministic. When the question is which instance, or what an element does, or whether the screen is in the expected state, only the VLM can answer it, and the art is in giving it only the few candidates the encoder couldn't separate.

Picking an architecture
Single-shotCoarse-to-fineCandidate-based
Accuracy on big or text targetsGoodGoodGood
Accuracy on small targetsDrops with screen sizeBetter, the crop adds pixelsGood if the generator finds them
Repeated targetsDepends on the descriptionDepends on the descriptionBest, relations are explicit
Model calls per step121, plus local steps
LatencyOne callTwo callsOne call plus milliseconds
Boxes land on real elementsNot guaranteedNot guaranteedYes, they come from the generator
Build and maintenanceLeastA little moreMost

Uncertainty and abstention

Don't click when the answer is uncertain

A generative VLM hands back a box with no reliable sense of how sure it is. You can ask it for a confidence number and it will write one, but that number is generated text, like the box, and it isn't calibrated, which means a stated 0.9 doesn't come true nine times in ten. So a production grounding step needs its own evidence of certainty, and the good sources of it are the parts of the system you can measure.

The clearest one is the margin between the top two candidates. Here are two illustrative outcomes of the embedding filter for the same request.

The top score alone can't tell you how sure to be, illustrative scores
Candidate ACandidate BMarginWhat it means
Screen 10.910.900.01Two near-identical candidates, ambiguous
Screen 20.910.470.44One clear match

Both top scores are 0.91, and only one of these should be clicked without more work. Screen 1 is the outliner eyes again, where a high similarity for everything is the sign of a reference problem, and the next step is the relational check from section 7. That's the same lesson M19-1 taught about similarity scores, which is that a score ranks candidates and says nothing about whether the top one is the answer.

Other proxies are worth combining. Agreement between two methods is strong evidence, like a VLM's box that contains the OCR box for the same word, or a candidate that wins on both the embedding ranker and the VLM. Agreement between two models costs a second call and catches the rare wild answer. Asking the VLM to verify, "is the element in this crop the eye icon for Light?", is cheaper than the original grounding because the crop is small. And the expected-state check after the click, from section 5, is the final proxy, since a click that produced the right screen was right whatever the scores said.

When the evidence is weak, the agent has better options than clicking and hoping. It can retry with a crop at higher resolution, which fixes perception problems. It can escalate to a stronger or more expensive model for that one step. It can ask for a better observation, like scrolling the panel so the row is fully visible, hovering to reveal a tooltip, or widening the window. Or it can abstain and hand the step to a person with the screenshot and the candidates it couldn't separate. Northlight's rule is to click only when the top candidate beats the second by a margin picked from its test dataset, or when two methods agree, and otherwise to crop and retry once and then ask. A wrong click in Blender can delete an artist's afternoon, and an extra second of checking can't.

Evaluation

Measure grounding against boxes or crops you trust

As M1 said about every model, you need your own dataset before you can say a grounding model works on your screens. For grounding that means screenshots of your application, each with a description of a target and a way to check the answer, and which way works depends on what reference you have.

When you have the correct box, check the overlap

If someone has drawn the correct box on the same screenshot, you can compare the model's box with it directly. Intersection over union, IoU, is the area the two boxes share divided by the area they cover together, so it's 1 for a perfect match and 0 when they don't touch, and 0.5 is the usual pass mark in detection benchmarks.

ground.py — intersection over union for two pixel boxes
def iou(a, b):
    ix = max(0, min(a[2], b[2]) - max(a[0], b[0]))
    iy = max(0, min(a[3], b[3]) - max(a[1], b[1]))
    inter = ix * iy
    area = lambda r: (r[2] - r[0]) * (r[3] - r[1])
    return inter / (area(a) + area(b) - inter)

For an agent that clicks, a simpler check often counts for more, which is whether the centre of the model's box falls inside the correct box, because that's where the click lands. A box can have a low IoU and still click the right button, and a box with a decent IoU on a tiny icon can still put the click on the neighbour. Record both.

When you have a reference crop, compare the images

Often you don't have boxes on the same screenshot. You have a reference image of the target, like a crop of the Render icon from an earlier screenshot, and the screen you're testing on is a different size, or a different theme, or the icon moved. IoU needs both boxes on the same image, so it can't compare these. What you can do is crop whatever the model boxed and ask whether that crop shows the same thing as the reference.

My approach is to embed both crops with an image encoder and compare them with cosine similarity, the same comparison M19-1 used for text, and to test which encoder works for the images at hand. For UI elements SigLIP has worked best for me (Zhai et al., 2023), and for other images CLIP or DINO can be better. As M19-1 showed for text, a similarity score isn't a probability, so the threshold for "same element" has to come from labeled pairs of your own crops. The other approach is to send both crops to a VLM and ask whether they show the same element, which handles flipped or restyled icons well and costs a call per check. And keep section 7 in mind when you read a crop check, since a crop of the wrong eye icon matches the reference perfectly, so repeated targets need their boxes checked by position.

Checking a grounding answer
CheckNeedsGood for
IoU of the two boxesThe correct box on the same screenshotComparing models on a labeled test dataset
Centre inside the correct boxThe correct box on the same screenshotAgents that click the centre
Encoder similarity of the cropsA reference crop, and a threshold from labeled pairsDifferent screen sizes, themes or positions, for targets with no identical twin
A VLM asked whether two crops matchA reference crop, one extra callRestyled or flipped icons, spot checks

Slice the results, because the average hides the failures

Every example in the test dataset should carry a few labels besides its box, the kind of target, its size, whether it's text or an icon, how many look-alikes share the screen, the screenshot's resolution, the theme and the application. They cost seconds to add when you label, and they're what turns one accuracy number into a map of where the system fails. This is an illustrative result from Northlight's test dataset.

One model on Northlight's test dataset, illustrative numbers
SliceTargetsCentre inside the correct box
All targets40091%
Text buttons and menus15098%
Icons at 24 pixels or larger12096%
Icons under 16 pixels7071%
Repeated targets, like the outliner eyes6063%

A 91% model sounds ready to ship. The slices say it's ready for menus, needs coarse-to-fine for small icons and needs the candidate pipeline for repeated ones, and that nine failures in ten come from two slices that are a third of the dataset. So the question to ask of a grounding eval is where the model fails, and "what's its accuracy" comes second. Northlight's test dataset is 120 screenshots of its own Blender scenes at three window sizes, with 400 targets described the way the agent describes them, and every model or prompt change runs against all of it, sliced, with the failures read one by one, the same way M19-1's retrieval failures were.

Grounding accuracy isn't the same as agent success

There are four different things you can measure, and each sits on top of the one before.

On the left, four stacked levels. Grounding accuracy asks whether we located the requested thing, measured by IoU or centre-inside-box on the test dataset. Click accuracy asks whether the click landed on it on the real screen, after the coordinate conversion. Step success asks whether the app reached the expected state, measured by the look-again check. Task success asks whether the whole task finished, such as the exported file or the saved setting. On the right, a line chart of task success against the number of steps from 1 to 30, with one line per per-step success rate. At 30 steps, 0.995 per step gives 86.0%, 0.99 gives 74.0%, the highlighted 0.98 gives 54.5% and 0.95 gives 21.5%.
The curves assume independent steps, which real agents don't quite have, so they show the shape of the problem more than a forecast.

The gaps between levels are where bugs hide. Northlight's second version had good grounding accuracy and terrible click accuracy, which was the Retina factor. And the levels multiply. If each step succeeds 98% of the time, independently, a 30-step export task succeeds 0.98 ** 30, about 54.5% of the time, so a system that's right on almost every click fails nearly half its tasks. That's the arithmetic behind Northlight's artists losing trust in the two-minute export. It's also why agents need per-step reliability close to 100%, and why verification after each step and recovery on failure matter as much as a better model. A failed step the agent notices and retries costs seconds, while a failed step it doesn't notice ruins the task.

Failure taxonomy

Name the failure before you fix it

When a click goes wrong, "the grounding failed" is too vague to act on, because the fixes for different failures are different, and the wrong fix wastes a week. Every failed step falls into one of seven categories, and each points at a different remedy.

Seven ways a grounding step fails
FailureWhat happenedHow you tellWhat fixes it
PerceptionThe target was too small, blurred, covered or lost in resizingThe target is barely visible in the image the model sawCrop and enlarge, send more pixels, coarse-to-fine
RecognitionThe model saw the region and misread what the element isThe box is on a real element of the wrong kind, like a lock icon for an eyeA clearer description, a reference crop, a stronger model
Reference resolutionRight kind of element, wrong instanceThe box is on an identical twin of the targetName the anchor in the description, OCR, spatial rules, candidates with a VLM chooser
Spatial or relational reasoningIt misread "beside", "under" or "in the toolbar"The box is near the anchor on the wrong side, or in the wrong panelCrop to the panel, state the relation plainly, candidate ranking by rule
LocalizationRight element, bad boxThe box is on the target but offset or oversized, and the centre missesRefinement in a crop, a box from a parser, click the parser's centre
Coordinate systemThe model's box was right, and your conversion was wrongThe box looks right drawn on the image the model saw, and the click misses in a patternFix and test the conversion function
Execution or stateThe click landed on the right target and the app didn't do the expected thingThe click is correct, and the next screenshot isn'tWait for the redraw, verify, retry, handle dialogs and focus

The order you check them in saves time, and the first check is always the same. Draw the model's box onto the exact image the model received, which may be smaller than your original screenshot. If the box is right there and the click was wrong, it's a coordinate failure, and no model change will fix it. If the box is wrong on that image, look at the target inside it. Barely visible means perception, visible but a different kind of thing means recognition, the right kind of thing but a twin means reference resolution or relational reasoning, and the right element with a sloppy box means localization. Only if the box and the click were both right do you look at the application's state, which is an execution failure.

Once grounding models got strong, many of the failures that matter in production moved into the system around them, into resizing, coordinate conventions, stale screenshots, reference ambiguity, verification and execution. That's where most of Northlight's time went. But the models still fail too, on tiny targets, dense interfaces, spatial relationships, repeated controls, poor images and screens unlike anything they trained on, so both kinds belong in the taxonomy. The newest models are too new for anyone to have mapped their failures, and I don't know yet where they fail. What's known comes from slightly older models and the benchmarks. Roboflow found that "Sol becomes less stable on images around 2,000 by 2,000 pixels or larger, especially at lower reasoning effort", and that at those sizes "GPT-5.6 Sol returned boxes in seemingly random parts of the image", often in "unnatural layouts, such as straight rows or evenly spaced groups" (Roboflow, July 2026). That's a perception and localization failure triggered by size, so a 4K screenshot should either go in smaller or go through coarse-to-fine, and boxes laid out in a neat grid that doesn't match the screen are a sign to check for.

The whole diagnosis fits on one page.

A decision tree that starts from a highlighted instruction, draw the model's box on the exact image the model received, which may be smaller than your screenshot. If the box is on the target, the left branch asks whether the click landed on the target on the real screen, and if not, it's a coordinate system failure, fixed by fixing and testing the conversion function. If it did, it asks whether the app reached the expected next state, and if not, it's an execution or state failure, fixed by waiting for the redraw, verifying, retrying once and handling dialogs. If the box is wrong, the right branch asks whether the target is clearly visible in that image, and if not it's perception, fixed by cropping and enlarging or coarse-to-fine. If it is, it asks whether the box is on the right kind of element, and if not it's recognition, fixed by a clearer description, a reference crop or a stronger model. If it is, it asks whether it's the right instance in the right place, and if not it's a reference or relational failure, fixed by naming the anchor, OCR, spatial rules and candidates with a VLM. If it is, the right element has a sloppy box, a localization failure, fixed by refining in a crop or clicking a parser's box.
Reference resolution and spatial reasoning share a leaf here, since the check that finds them is the same.

Keep the evidence for every failure, the screenshot the model saw, its raw answer, the converted click and the next screenshot. With those four, classifying a failure takes a minute, and without them it's a guess, because a grounding failure is impossible to debug from a log line and obvious from the pictures.

When to fine-tune

Fine-tune only when systematic failures remain

Fine-tuning is the last step on a long ladder, and most grounding problems are solved on the lower rungs. Before training anything, make sure the prompt asks for a stated format as structured output, that the coordinate conversion is tested, that screenshots reach the model at a resolution where the targets survive, with crops where they don't, that candidates come from a generator where the screen is crowded, and that a sliced test dataset says exactly where the remaining failures are. If one slice still fails systematically after all that, like every icon in one custom panel, fine-tuning is worth considering, and M21 covers how.

What you fine-tune depends on which part is failing. Fine-tuning a detector like YOLO on your own UI gives you a fast candidate generator, and it's the cheapest training to run, especially with VLM-drawn labels from section 6. Fine-tuning an image encoder on your own crops makes the embedding ranker separate your icons better, which helps when two different icons look alike to a general encoder. Fine-tuning the VLM itself on your screens and descriptions is the most expensive and the only one that fixes reasoning, like reference resolution in an unusual layout. And some open VLMs let you train only the projector, the way LLaVA's first stage did, which is cheap and helps the language model make better use of features the encoder already has, though it can't fix an encoder that never saw your kind of image.

Whichever you train, the examples that teach the most are hard negatives, wrong answers that look almost right. For the target "the eye icon in the Light row", the positive example is Light's eye, and the hard negatives are the Camera, Cube and Collection eyes, which are identical in pixels and wrong in position. Easy negatives, like the Render menu, teach the model nothing it doesn't already know. When visually similar candidates dominate your screens, as they do in any list, table or toolbar, a training dataset without hard negatives teaches the model to find eye icons, which it can already do, and leaves it unable to pick the right one, which was the problem.

Operations

Cost and speed per click

A grounding call costs what any VLM call costs, the image tokens plus the text, and at current prices that's a few cents or less. Roboflow's measurements from July 2026 put GPT-5.6 Sol at about 10 seconds and 2.5 cents per image, Terra at about 6 seconds and 1 cent, and Luna at about 5 seconds and under half a cent. OpenAI released GPT-6.1 Sol on 29 September 2026 at $2 per million input tokens and $10 per million output tokens, so the roughly 1,570 image tokens of a 1440-by-900 screenshot cost about a third of a cent on the way in (TechCrunch, 29 September 2026). Reasoning tokens and the answer come on top, and with a reasoning model those can be most of the cost.

For an agent, speed is a bigger problem than cost, because every step waits for a call. At several seconds per grounding call, a 20-click task takes minutes, where a person takes seconds. A local detector, OCR engine or embedding ranker answers in milliseconds, which is why detectors still win at high volume, why the candidate pipeline in section 8 is often faster than single-shot, and why the component question in section 6 comes down to volume and latency as much as accuracy. Coarse-to-fine doubles the calls on the steps that use it, so it's worth gating on target size or uncertainty. For Northlight, a chore that took an artist ten minutes of attention now takes the agent three minutes of nobody's attention, which is the trade that made it worth building.

Resolution is the main lever on both. A smaller screenshot is fewer image tokens and a faster call, and smaller icons, so the right size is the smallest one at which your sliced test dataset still passes on the slices you care about, found by running it at two or three sizes.

Hands-on

Build a grounding system for one application, and take it apart

Pick a desktop or web application you use, one with small icons and some repeated controls, like a layers panel or a list with an action per row. Blender, a code editor or a design tool all work. By the end you'll have numbers and failures from your own system, which is what makes this something you can talk about in an interview.

Build the test dataset first

Take 30 or more screenshots across at least two window sizes, and on each write down two or three targets the way an agent would describe them. Make sure the dataset includes large targets, tiny icons under 16 pixels, text buttons, icons, and repeated targets like one eye icon among four identical ones. Draw the correct box for each, save a crop of each target as a reference image, and label every example with its slice, meaning target type, size, text or icon, look-alike count and screenshot resolution.

Compare two approaches, at two resolutions

Implement single-shot grounding with a VLM and one other approach, either coarse-to-fine or a candidate pipeline with OCR or OmniParser. Run both at two screenshot resolutions. Ask for boxes in a stated format with JSON checked against a schema, convert every answer in one tested function, and score IoU and centre-inside-box against your boxes, sliced by target type and size. Record the time and cost of every call.

Probe embeddings and uncertainty

Rank candidate crops with an image encoder, SigLIP to start. Find one example where it clearly helps, like picking an icon among many different ones, and one where it can't solve the problem alone, like the repeated targets. For the repeated ones, record the top-two margin, and decide the rule you'd use to click, crop and retry, escalate or abstain.

Then build the loop, and break it

Build the screenshot, ground, click, verify loop from section 5 on top of your better approach, with the coordinate conversion tested for your display's scaling and for crops, and saved screenshots on every failure. Run it on one real task of five or more steps, a few times. Then collect at least three failures and classify each with the taxonomy from section 11.

What to hand in

  • The test dataset, with its slice labels and a line on why each screen is in it.
  • A table of both approaches at both resolutions, overall and per slice, with IoU, centre-inside-box, time and cost.
  • The tests for your coordinate conversion, including your display's scaling, a crop mapped back, and one provider's format that isn't pixel corners.
  • One coarse-to-fine example, with the screenshot, the crop and both answers.
  • The embedding example that helped, the one it couldn't solve, and your confidence rule with the margins that justify it.
  • At least three failures with their screenshots, each classified and paired with the fix it points to.
  • The task's success rate over your runs, next to your per-step success rate, and whether the two agree with the multiplication in section 10.
  • A short note on whether a detector or fine-tuning would be justified for this application, and what evidence would change your mind.

Putting it together

Putting it together

Visual grounding turns a description into a position. A VLM does it by cutting the screenshot into patches, encoding them into visual features, projecting those into tokens its language model can read next to your text, and writing a box back one token at a time, and every step that shrinks the image or merges tokens costs small targets some of their detail. For years VLMs could recognise things without placing them. That changed with Gemini 2.0 and then sharply with GPT-5.6 and the GPT-6 models, which ground UI elements on a crowded professional screen from one instruction.

Once grounding got that good, many of the failures that matter in production moved into the system around the model. The image the model saw has to be the image your code thinks it saw, the box has to be read in the format the model used, and the click has to be scaled to the screen's own units, all in one tested function. Identical targets need a reference, found from nearby text, rows and hierarchy, because pixels alone can't separate twins. The models still fail on tiny targets, repeated controls, spatial relationships and unfamiliar screens, so the design counts too. Single-shot is simplest, coarse-to-fine buys pixels for small targets, and candidate generation with ranking uses embeddings for looks and a VLM for reasoning, with fewer and better choices for the VLM to make.

An agent built on that clicks only when it has evidence to, checks every click by looking again, and names each failure before fixing it. It's measured on its own application's screens, sliced by the kinds of target that fail differently, and judged by task success, since per-step accuracy multiplies. And it reaches for fine-tuning last, with hard negatives, when a slice keeps failing after everything else.

The finished grounding step in Northlight's agent
PartWhat it does
ScreenshotCaptured fresh each step, its pixel size and the screen's click size recorded
RequestThe target described by what it says or looks like, with its anchor named, the box format stated and JSON required
ArchitectureSingle-shot for menus, OCR for text targets, candidates with a VLM chooser for repeated icons, coarse-to-fine for small ones
ConversionOne tested function from the model's format, and from crops, to screen click coordinates
ConfidenceClick when the top-two margin is wide or two methods agree, otherwise crop and retry once, then ask a person
ChecksBox inside the image, and a look-again after every click
Measurement400 targets on 120 screenshots, sliced by target type and size, rerun on every change, with task success tracked apart

M18 covers the tool loop the agent runs, M19-1 the embeddings and similarity scores used to rank candidates, and M21 the fine-tuning this module only decides on. M23 applies grounding to fields on documents, and M30 puts screenshots and clicks inside agents that run whole tasks for hours.

Checkpoint · recall · 5 questions

What the module said

  1. 01

    What does visual grounding add to what a VLM could already do?

  2. 02

    What does the projector in a VLM do?

  3. 03

    In what order does Gemini's box_2d give a box's numbers?

  4. 04

    Why does a Mac with a Retina display cause clicks to miss when an agent sends the raw screenshot?

  5. 05

    What does OmniParser return for a screenshot?

0 / 5 answered

Checkpoint · understanding · 5 questions

Reason it through

  1. 01

    The four eye icons in Blender's outliner embed to almost identical vectors. Why can't an image encoder pick the one for Light?

  2. 02

    Northlight's clicks missed by more the further the target was toward the bottom right. What does that pattern point to?

  3. 03

    The embedding ranker scores the top two candidates 0.91 and 0.90. What should the agent do?

  4. 04

    Each step of an agent succeeds 98% of the time. Why does a 30-step task still fail often?

  5. 05

    When is candidate generation and ranking unnecessary complexity?

0 / 5 answered

Checkpoint · debugging · 4 questions

Debug it

  1. 01

    After switching providers, every box Northlight draws back onto the screenshot looks mirrored across the diagonal, landing near where the target would be if x and y were swapped. What's wrong?

  2. 02

    The agent's box for "the Render menu" is correct, the click lands on it, and the next screenshot shows no menu open. Which failure is it, and what should the agent do?

  3. 03

    On a 4K monitor, grounding accuracy on small icons drops sharply compared with the same app on a laptop screen. What's the likely failure type, and the fix?

  4. 04

    Asked to hide the light, the agent clicks the eye icon in the Camera row. Drawn on the image the model saw, its box sits exactly on the Camera eye. Which failure is it, and what fixes it?

0 / 4 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.