What it is
It predicts the features of what it can't see
Take one image. Keep a large visible region as the context block, then pick a few other regions as targets. One encoder reads the context. A second encoder, a slowly updated copy of the first, reads the full image and produces the features of each target block. A small predictor then has to guess those target features from the context, given only the position of each block.

All three approaches in that bottom strip learn without labels, and what separates them is what the model is asked to reproduce. MAE reconstructs pixels. DINO matches the outputs of two views of the same image. JEPA predicts features of a part it hasn't seen, which the paper describes as learning "without relying on hand-crafted data-augmentations".
In a test run for this page, the released I-JEPA model gave a mirrored copy of a street photo a cosine similarity of 0.917 with the original, against 0.93 for DINOv2 on the same pair. The features are good in the same broad way, and the training recipe is what differs.
Where it shows up
Video is where the family is aimed
Video understanding
V-JEPA applies the same idea over time, hiding parts of a video and predicting their features. Its paper describes models "trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision", on 2 million videos.
Prediction and planning for robots
V-JEPA 2 (2025) pre-trained on "over 1 million hours of internet video", then added a small amount of robot interaction data to train an action-conditioned version called V-JEPA 2-AC. A model that can predict what its features will look like after an action can search over actions, which is planning in feature space. Section 3 walks through one planning step, and the world models entry on this index covers the wider use.
Image features with no labels
I-JEPA gives you a backbone in the same way DINOv2 does, for retrieval, clustering or a small head trained on it.
Vision-language without generating tokens
VL-JEPA (2025) applies the idea to text, predicting "continuous embeddings of the target texts" in place of generating them token by token.
How it works
The target it predicts is produced by a copy of itself
Two encoders and a predictor
The context encoder is trained by gradient descent. The target encoder is an exponential moving average of it, and gets no gradients, the same arrangement DINO uses for its teacher. The predictor is small, and it's thrown away after training, leaving the encoder as the thing you keep.
The masking is the design
I-JEPA samples several target blocks covering a sizeable part of the image and gives the context encoder what's left. The paper calls the masking strategy "a core design choice to guide I-JEPA towards producing semantic representations", because targets that are too small or too scattered can be filled in from local texture, which teaches the model nothing.
V-JEPA masks regions across frames, so predicting them needs some sense of how the scene moves, where a neighboring pixel would be enough for a still image.
Why the objective lives in feature space
A pixel-level objective has to reproduce everything, including detail that's genuinely unpredictable. A feature-level objective lets the model discard that detail, as long as both encoders agree to discard it. The risk is the opposite failure, where both sides agree on a constant and the loss goes to zero, which is collapse. The moving-average target encoder and the asymmetry between the two paths are what hold it off.
A worked example, moving a robot arm toward a goal image
V-JEPA 2-AC shows what feature prediction is for once actions are in the loop. The frozen V-JEPA 2 encoder stays as it is, and a 300M-parameter predictor is trained on top of it to take the features of the current video frame, together with an action, and predict the features of the next frame. The action is a change to the arm's end-effector, the gripper at the end of the arm, with three numbers for position, three for orientation and one for the gripper. Training used under 62 hours of unlabeled clips from the public Droid robot dataset, with no rewards and no labels saying what task each clip was.
Take the paper's grasp task, where a Franka arm has to pick up a cup and the only instruction is one goal image of the finished grasp. One step of planning goes like this.
- 01The encoder turns the current camera frame into features, and turns the goal photo into features too.
- 02The planner samples 800 candidate actions and asks the predictor what the features would look like after each one.
- 03Each candidate gets an energy, the L1 distance between its predicted features and the goal features, and a lower energy means the imagined scene looks more like the goal.
- 04The cross-entropy method refits the sampling to the lowest-energy candidates and samples again, 10 times in all.
- 05The robot runs only the first action of the best candidate, takes a new camera frame, and plans again from there.
Nothing in that loop draws a picture. The comparison happens between feature vectors, which is why a step took 16 seconds on one RTX 4090 in the paper's setup, where the same search with Cosmos, a video-generation world model, took 4 minutes per action with a tenth as many samples. Averaged over two labs the robots were never trained in, the arm reached a target position in 100% of trials, grasped the cup in 65% and grasped a box in 25%. Pick-and-place worked in 80% of trials with the cup and 65% with the box, but only when the planner was given two intermediate goal photos along the way, because the paper plans one action ahead and a longer search grows exponentially with the horizon. The paper also reports that results depended on where the camera was placed, and the authors tried several positions by hand before settling on one.
V-JEPA 2.1 went back for the dense features
The 2026 report, V-JEPA 2.1, addresses the same weakness DINOv3 did, which is that features good for whole-video questions aren't automatically good per patch. Its recipe adds "a dense predictive loss" where both visible and masked tokens contribute, and applies the objective at several depths of the network.
Versions
The versions you'll see, as of September 2026
| Model | Released | What it covers | Weights |
|---|---|---|---|
| I-JEPA | January 2023 | Images | facebook/ijepa_vith14_1k, CC BY-NC 4.0, so research use only |
| V-JEPA | February 2024 | Video | Released on GitHub by Meta |
| V-JEPA 2 | June 2025 | Video, plus an action-conditioned model for planning | facebook/vjepa2-vitl-fpc64-256 and larger, MIT |
| V-JEPA 2.1 | March 2026 | Video with stronger per-patch features | Reported in the paper |
| VL-JEPA | December 2025 | Vision-language, predicting text embeddings | Reported in the paper |
Check the license per checkpoint. The I-JEPA weights on Hugging Face are non-commercial, while the V-JEPA 2 checkpoints are MIT, which is a large practical difference for anything you plan to ship.
Choosing
Pick it for video, and DINO for images today
| The job | Reach for | Why |
|---|---|---|
| Image features for search or clustering, or a small head | DINOv2 or DINOv3 | More checkpoints and a larger ecosystem, with permissive licenses for DINOv2 |
| Video understanding with no labels | V-JEPA 2 | It's trained on video from the start, where image models see single frames |
| Predicting how a scene will change, or planning actions | V-JEPA 2's action-conditioned model | It predicts future features, which is what planning searches over |
| Matching images to text | CLIP or SigLIP | JEPA models have no text tower, apart from VL-JEPA |
| Answering questions about an image | A vision-language model | JEPA returns features and writes nothing |
For images specifically, the honest summary is that DINOv2 and DINOv3 are the practical default, and I-JEPA is the design more papers build on. For video, V-JEPA 2 is both.
Try it
How to try it
I-JEPA loads through transformers. It returns one vector per patch with no pooled output, so averaging the patch tokens is the usual way to get a single vector for an image. This is the code behind the 0.917 above.
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
model_id = "facebook/ijepa_vith14_1k" # CC BY-NC 4.0: research use
model = AutoModel.from_pretrained(model_id).eval()
proc = AutoProcessor.from_pretrained(model_id)
def embed(path):
inputs = proc(images=Image.open(path).convert("RGB"), return_tensors="pt")
with torch.no_grad():
patches = model(**inputs).last_hidden_state # one vector per patch
return patches.mean(dim=1)[0] # average them
a, b = embed("photo.jpg"), embed("photo_mirrored.jpg")
print(torch.nn.functional.cosine_similarity(a, b, dim=0).item())For video, facebook/vjepa2-vitl-fpc64-256 loads the same way with VJEPA2Model and a video processor, and it takes a clip of frames in place of one image.
Compare I-JEPA and DINOv2 as frozen backbones on my own images. Load facebook/ijepa_vith14_1k through transformers and dinov2_vits14 through torch.hub, and embed every image in a labeled folder with each model, averaging patch tokens for I-JEPA and using the CLS vector for DINOv2. Then train a logistic regression on each model's frozen features with 5-fold cross-validation, and report accuracy per model and per class. Also report how long embedding took per image for each model, and save a CSV of every image's prediction from both so I can look at where they disagree.
Limits
What it can't do
- It produces features and nothing else. No labels, no boxes, no text, so every task needs a head or a nearest-neighbor lookup on top.
- Its features can't be turned back into an image. A pixel-prediction model like MAE has a decoder, and a JEPA has no way to show you what it thinks the hidden part looks like.
- The image models trail the image specialists. DINOv2 and DINOv3 have more sizes and more permissive licenses, with far more downstream code around them.
- The planning results are early. V-JEPA 2's robot results come from a specific setup with interaction data, and they aren't a general planner you can drop into a product.
- Licenses differ sharply within the family. The published I-JEPA weights are non-commercial, so check each checkpoint before building on it.
- Evaluating it is indirect. Because the objective lives in feature space, comparing two JEPA models means training probes on top of them, which is slower than reading a single benchmark number.
Go deeper
Assran et al. (2023): Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA) · Bardes et al. (2024): Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA) · Assran et al. (2025): V-JEPA 2, Self-Supervised Video Models Enable Understanding, Prediction and Planning · V-JEPA 2.1 (2026): Unlocking Dense Features in Video Self-Supervised Learning · VL-JEPA (2025): Joint Embedding Predictive Architecture for Vision-language · He et al. (2021): Masked Autoencoders Are Scalable Vision Learners (the pixel-prediction alternative)
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
