MLGuerrillaStart with M1 →
Architectures·13 min read·Updated 24 September 2026

JEPA

JEPA is a family of self-supervised models from Meta that learn by predicting the features of the parts of an image or video they can't see. The prediction happens in feature space, and never in pixels.

JEPA stands for joint-embedding predictive architecture, a training idea Yann LeCun's group at Meta has been building on since 2023. The model hides part of an image or a video, and then predicts the features of the hidden part, as produced by a copy of the model itself. Nothing predicts pixels.

Pixel prediction spends capacity on detail that can't be predicted, like the exact texture of a face in a crowd, and predicting features lets the model leave that detail out. The family now covers images with I-JEPA and video with V-JEPA, and VL-JEPA extends it to vision-language.

What it is

It predicts the features of what it can't see

Take one image. Keep a large visible region as the context block, then pick a few other regions as targets. One encoder reads the context. A second encoder, a slowly updated copy of the first, reads the full image and produces the features of each target block. A small predictor then has to guess those target features from the context, given only the position of each block.

A figure titled "JEPA predicts the missing features, and never the missing pixels". On the left, a street photo is shown with four blanked rectangles outlined in terracotta, labeled visible as the context block and blanked as the target blocks the model must predict. Arrows lead to two boxes: a Context encoder that sees the visible part, and a dark Target encoder described as a moving average copy, with a note that no gradients reach the target encoder. Both feed a Predictor box, which is given the context features and the position of each target block and predicts that block's features. A loss box measures the distance between two feature vectors, with the target encoder's output routed in as the target the predictor is judged against. A strip at the bottom compares three approaches: MAE predicts the missing pixels so capacity goes into texture and color, DINO matches two views of the whole image so the crops do the work, and JEPA predicts the missing features so the target is what the model itself encodes.
The target encoder never receives gradients. It only changes by averaging in the context encoder's weights, the same trick DINO uses for its teacher.

All three approaches in that bottom strip learn without labels, and what separates them is what the model is asked to reproduce. MAE reconstructs pixels. DINO matches the outputs of two views of the same image. JEPA predicts features of a part it hasn't seen, which the paper describes as learning "without relying on hand-crafted data-augmentations".

In a test run for this page, the released I-JEPA model gave a mirrored copy of a street photo a cosine similarity of 0.917 with the original, against 0.93 for DINOv2 on the same pair. The features are good in the same broad way, and the training recipe is what differs.

Where it shows up

Video is where the family is aimed

Video understanding

V-JEPA applies the same idea over time, hiding parts of a video and predicting their features. Its paper describes models "trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision", on 2 million videos.

Prediction and planning for robots

V-JEPA 2 (2025) pre-trained on "over 1 million hours of internet video", then added a small amount of robot interaction data to train an action-conditioned version called V-JEPA 2-AC. A model that can predict what its features will look like after an action can search over actions, which is planning in feature space. Section 3 walks through one planning step, and the world models entry on this index covers the wider use.

Image features with no labels

I-JEPA gives you a backbone in the same way DINOv2 does, for retrieval, clustering or a small head trained on it.

Vision-language without generating tokens

VL-JEPA (2025) applies the idea to text, predicting "continuous embeddings of the target texts" in place of generating them token by token.

How it works

The target it predicts is produced by a copy of itself

Two encoders and a predictor

The context encoder is trained by gradient descent. The target encoder is an exponential moving average of it, and gets no gradients, the same arrangement DINO uses for its teacher. The predictor is small, and it's thrown away after training, leaving the encoder as the thing you keep.

The masking is the design

I-JEPA samples several target blocks covering a sizeable part of the image and gives the context encoder what's left. The paper calls the masking strategy "a core design choice to guide I-JEPA towards producing semantic representations", because targets that are too small or too scattered can be filled in from local texture, which teaches the model nothing.

V-JEPA masks regions across frames, so predicting them needs some sense of how the scene moves, where a neighboring pixel would be enough for a still image.

Why the objective lives in feature space

A pixel-level objective has to reproduce everything, including detail that's genuinely unpredictable. A feature-level objective lets the model discard that detail, as long as both encoders agree to discard it. The risk is the opposite failure, where both sides agree on a constant and the loss goes to zero, which is collapse. The moving-average target encoder and the asymmetry between the two paths are what hold it off.

A worked example, moving a robot arm toward a goal image

V-JEPA 2-AC shows what feature prediction is for once actions are in the loop. The frozen V-JEPA 2 encoder stays as it is, and a 300M-parameter predictor is trained on top of it to take the features of the current video frame, together with an action, and predict the features of the next frame. The action is a change to the arm's end-effector, the gripper at the end of the arm, with three numbers for position, three for orientation and one for the gripper. Training used under 62 hours of unlabeled clips from the public Droid robot dataset, with no rewards and no labels saying what task each clip was.

Take the paper's grasp task, where a Franka arm has to pick up a cup and the only instruction is one goal image of the finished grasp. One step of planning goes like this.

  1. 01The encoder turns the current camera frame into features, and turns the goal photo into features too.
  2. 02The planner samples 800 candidate actions and asks the predictor what the features would look like after each one.
  3. 03Each candidate gets an energy, the L1 distance between its predicted features and the goal features, and a lower energy means the imagined scene looks more like the goal.
  4. 04The cross-entropy method refits the sampling to the lowest-energy candidates and samples again, 10 times in all.
  5. 05The robot runs only the first action of the best candidate, takes a new camera frame, and plans again from there.

Nothing in that loop draws a picture. The comparison happens between feature vectors, which is why a step took 16 seconds on one RTX 4090 in the paper's setup, where the same search with Cosmos, a video-generation world model, took 4 minutes per action with a tenth as many samples. Averaged over two labs the robots were never trained in, the arm reached a target position in 100% of trials, grasped the cup in 65% and grasped a box in 25%. Pick-and-place worked in 80% of trials with the cup and 65% with the box, but only when the planner was given two intermediate goal photos along the way, because the paper plans one action ahead and a longer search grows exponentially with the horizon. The paper also reports that results depended on where the camera was placed, and the authors tried several positions by hand before settling on one.

V-JEPA 2.1 went back for the dense features

The 2026 report, V-JEPA 2.1, addresses the same weakness DINOv3 did, which is that features good for whole-video questions aren't automatically good per patch. Its recipe adds "a dense predictive loss" where both visible and masked tokens contribute, and applies the objective at several depths of the network.

Versions

The versions you'll see, as of September 2026

The JEPA family (papers and Hugging Face)
ModelReleasedWhat it coversWeights
I-JEPAJanuary 2023Imagesfacebook/ijepa_vith14_1k, CC BY-NC 4.0, so research use only
V-JEPAFebruary 2024VideoReleased on GitHub by Meta
V-JEPA 2June 2025Video, plus an action-conditioned model for planningfacebook/vjepa2-vitl-fpc64-256 and larger, MIT
V-JEPA 2.1March 2026Video with stronger per-patch featuresReported in the paper
VL-JEPADecember 2025Vision-language, predicting text embeddingsReported in the paper

Check the license per checkpoint. The I-JEPA weights on Hugging Face are non-commercial, while the V-JEPA 2 checkpoints are MIT, which is a large practical difference for anything you plan to ship.

Choosing

Pick it for video, and DINO for images today

Which self-supervised model to reach for
The jobReach forWhy
Image features for search or clustering, or a small headDINOv2 or DINOv3More checkpoints and a larger ecosystem, with permissive licenses for DINOv2
Video understanding with no labelsV-JEPA 2It's trained on video from the start, where image models see single frames
Predicting how a scene will change, or planning actionsV-JEPA 2's action-conditioned modelIt predicts future features, which is what planning searches over
Matching images to textCLIP or SigLIPJEPA models have no text tower, apart from VL-JEPA
Answering questions about an imageA vision-language modelJEPA returns features and writes nothing

For images specifically, the honest summary is that DINOv2 and DINOv3 are the practical default, and I-JEPA is the design more papers build on. For video, V-JEPA 2 is both.

Try it

How to try it

I-JEPA loads through transformers. It returns one vector per patch with no pooled output, so averaging the patch tokens is the usual way to get a single vector for an image. This is the code behind the 0.917 above.

ijepa_embed.py — one vector per image from I-JEPA
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor

model_id = "facebook/ijepa_vith14_1k"          # CC BY-NC 4.0: research use
model = AutoModel.from_pretrained(model_id).eval()
proc = AutoProcessor.from_pretrained(model_id)

def embed(path):
    inputs = proc(images=Image.open(path).convert("RGB"), return_tensors="pt")
    with torch.no_grad():
        patches = model(**inputs).last_hidden_state   # one vector per patch
    return patches.mean(dim=1)[0]                     # average them

a, b = embed("photo.jpg"), embed("photo_mirrored.jpg")
print(torch.nn.functional.cosine_similarity(a, b, dim=0).item())

For video, facebook/vjepa2-vitl-fpc64-256 loads the same way with VJEPA2Model and a video processor, and it takes a clip of frames in place of one image.

Ask your AI coding tool

Compare I-JEPA and DINOv2 as frozen backbones on my own images. Load facebook/ijepa_vith14_1k through transformers and dinov2_vits14 through torch.hub, and embed every image in a labeled folder with each model, averaging patch tokens for I-JEPA and using the CLS vector for DINOv2. Then train a logistic regression on each model's frozen features with 5-fold cross-validation, and report accuracy per model and per class. Also report how long embedding took per image for each model, and save a CSV of every image's prediction from both so I can look at where they disagree.

Limits

What it can't do

  • It produces features and nothing else. No labels, no boxes, no text, so every task needs a head or a nearest-neighbor lookup on top.
  • Its features can't be turned back into an image. A pixel-prediction model like MAE has a decoder, and a JEPA has no way to show you what it thinks the hidden part looks like.
  • The image models trail the image specialists. DINOv2 and DINOv3 have more sizes and more permissive licenses, with far more downstream code around them.
  • The planning results are early. V-JEPA 2's robot results come from a specific setup with interaction data, and they aren't a general planner you can drop into a product.
  • Licenses differ sharply within the family. The published I-JEPA weights are non-commercial, so check each checkpoint before building on it.
  • Evaluating it is indirect. Because the objective lives in feature space, comparing two JEPA models means training probes on top of them, which is slower than reading a single benchmark number.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.