Reference
Models and architectures
These are the models you run into when you build AI systems, grouped by the kind of data they work with. Each one has a line on what it's used for, so when a problem comes up you can remember which model handles it. The architectures at the end explain how the models work inside.
83 entries · checked 18 September 2026 · 24 written in full, more each week
By modality
Text and language
Models that read and write text. Most AI products start here.
- Large language models (LLMs)Read →Chat, writing, summarizing, coding and most agent work.GPT-6, Claude, Gemini, Grok, Qwen, DeepSeek, Kimi
- Reasoning modelsRead →Multi-step problems in math, code and planning, where the model works through the problem before it answers.OpenAI reasoning models, DeepSeek-R1, Claude and Gemini thinking modes
- Open-weight LLMsRunning a model on your own hardware for privacy, cost control or fine-tuning.Qwen3.8, Kimi K3, DeepSeek V4, GLM-5.3, Gemma 4, gpt-oss, Nemotron 3
- Small language modelsOn-device and low-latency text tasks where a frontier model is too slow or too expensive.Gemma 4, Phi-4, gpt-oss-20b, Claude Haiku 4.5
- Jev and System One modelsRead →Classifying, routing and scoring at high volume. It returns a typed decision with a probability and never writes free text.Jev by TypeSafe AI (September 2026)
- BERT and encoder modelsRead →Fast, cheap text classification and entity tagging, fine-tuned on your own labels.BERT, RoBERTa, DeBERTa-v3, ModernBERT, mmBERT
- Encoder-decoder modelsRead →Translation, summarization and older fine-tuned text-to-text pipelines.T5, Flan-T5, BART, mT5
- Diffusion language modelsVery fast text generation, produced by refining all the tokens in parallel over a few steps.Mercury 2.5, DiffusionGemma, LLaDA 2
Embeddings and retrieval
Models that turn content into vectors so you can search by meaning. They are the core of RAG.
- Text embedding modelsRead →Semantic search, RAG, clustering and finding near-duplicates.OpenAI text-embedding-3, Gemini Embedding 2, Voyage 4, Qwen3-Embedding, BGE-M3
- RerankersRead →Re-scoring the top search results so the most relevant passages reach the LLM.Cohere Rerank 4, Qwen3-Reranker, jina-reranker-v3.5, bge-reranker
- BM25 and sparse retrievalRead →Exact keyword matching for names, codes and IDs, usually paired with embeddings in hybrid search.BM25, SPLADE
- ColBERT and late interactionRead →More precise retrieval by comparing a query and a document token by token, at a higher storage cost.ColBERTv2
- ColPaliRead →Searching PDFs and slides as page images, with no OCR step.ColPali, ColQwen
- Multimodal embeddingsRead →One search index across text, images, audio and video.Gemini Embedding 2, jina-embeddings-v5-omni, Qwen3-VL-Embedding
Vision
Models that look at images and video frames and report what is there and where.
- Image classifiersPutting one label on a whole image, such as defect or no defect, or which product is in the photo.ResNet, EfficientNet, ConvNeXt V2, ViT, MobileNetV5
- YOLORead →Real-time object detection with a box around every object, on GPUs and edge devices.YOLO26, YOLO11, YOLOv12
- DETR-family detectorsAccurate real-time detection that fine-tunes well on your own data.RF-DETR, RT-DETR, D-FINE, DEIMv2
- Grounding DINO and open-vocabulary detectionRead →Finding objects from a text prompt, like “red forklift”, with no training.Grounding DINO, DINO-X, OWLv2, YOLO-World, YOLOE, Rex-Omni
- SAM (Segment Anything)Pixel-exact masks from a click, a box or a text prompt, tracked through video.SAM 3.1, SAM 3, SAM 2
- DINORead →General image features for similarity search, clustering, and training a small classifier with few labels.DINOv3, DINOv2
- Depth estimationEstimating the distance to every pixel from a single camera, for robotics, AR and measurement.Depth Anything 3, Depth Pro, MiDaS
- Pose estimationTracking body, hand and face keypoints for fitness, sports and gesture apps.MediaPipe, RTMPose, YOLO26-pose, OpenPose
- Multi-object trackingFollowing the same object across video frames, for counting and traffic analysis.ByteTrack, BoT-SORT, DeepSORT
- Face detection and recognitionFinding faces and matching identities, for verification and photo search.RetinaFace, ArcFace, InsightFace
Vision-language
Models that connect images and text, so you can search images with words or ask questions about a picture.
- CLIPRead →Zero-shot image classification and searching images by text.CLIP, OpenCLIP, MetaCLIP 2
- SigLIPRead →A CLIP-style model that is cheaper to train, and the image encoder inside many open VLMs.SigLIP 2
- Vision-language models (VLMs)Read →Answering questions about images, screenshots, charts and documents.Qwen3-VL, Gemma 4, InternVL3.5, Molmo 2, Moondream, plus GPT, Claude and Gemini
- Florence-2One small model for captioning, detection, grounding and OCR, picked by a task prompt.Florence-2 base, Florence-2 large
Documents and OCR
Models that turn scans, PDFs and forms into text and structure you can use.
- OCR enginesReading printed text from scans and photos.Tesseract, PaddleOCR, docTR
- Document parsing modelsConverting PDFs with tables, formulas and complex layouts into clean Markdown or JSON.PaddleOCR-VL, DeepSeek-OCR 2, MinerU, olmOCR 2, Docling, Mistral OCR
- Layout-aware modelsPulling fields out of forms and invoices using both the text and where it sits on the page.LayoutLMv3, Donut
Image generation and editing
Models that create or edit images from text and reference images.
- Frontier image modelsTop-quality generation and editing through an API, including readable text inside images.GPT Image 2.5, Nano Banana 2, Seedream 5.0, Midjourney, Imagen
- Stable DiffusionOpen image generation you can run locally and customize.SD 1.5, SDXL, SD 3.5
- FLUXHigh-quality image generation and editing, open or through an API.FLUX.2, FLUX.1 Kontext, FLUX 3
- Other open image modelsSelf-hosted generation and editing beyond Stable Diffusion and FLUX.Qwen-Image, Z-Image, HunyuanImage 3.0, Ideogram 4.0
- ControlNet, LoRA and adaptersSteering generation with a pose, edge or depth map, and teaching a model a new style or subject from a few images.ControlNet, LoRA, IP-Adapter, DreamBooth
- Upscaling, inpainting and background removalThe cleanup steps in an image pipeline.Real-ESRGAN, LaMa, BiRefNet
Video
Models that understand video or generate it.
- Video understanding modelsRecognizing actions, searching footage, and answering questions about clips.V-JEPA 2.1, VideoMAE, Molmo 2, Qwen3-VL, Gemini
- Video generation modelsGenerating clips from text or images, with synchronized sound in the newest models.Veo 3.1, Gemini Omni, Kling 3.0, Seedance 2.5, Runway Gen-4.5
- Open video modelsSelf-hosted video generation and fine-tuning.Wan 2.2, LTX-2.5, HunyuanVideo 1.5
- Talking avatars and lip-syncMaking a face speak a given audio track, for dubbing and video presenters.OmniHuman 1.5, HeyGen Avatar V, InfiniteTalk, LatentSync
Speech and audio
Models that listen, speak and make sound. Together they make up a voice agent.
- WhisperRead →The default open model for transcribing speech in many languages.Whisper large-v3, large-v3-turbo, faster-whisper
- Speech-to-text (ASR)Transcribing calls, meetings and voice commands, in batch or live.Parakeet, Qwen3-ASR, Cohere Transcribe, Deepgram Nova-3, AssemblyAI
- Text-to-speech (TTS)Giving an app a voice, from narration to cloned voices.ElevenLabs v3, Cartesia Sonic, Kokoro, Qwen3-TTS, Chatterbox
- Realtime speech-to-speechVoice agents that listen and answer in one model, with low latency.gpt-realtime, Gemini Live, Qwen3-Omni, Moshi
- Voice activity and turn detectionKnowing when someone is speaking and when they have finished, so a voice agent doesn't talk over them.Silero VAD, Pipecat Smart Turn, LiveKit turn detector
- Speaker diarizationLabeling who spoke when in a recording.pyannote
- Audio classification and embeddingsRecognizing sounds and searching audio by a text description.CLAP, BEATs, AST
- Music generationGenerating songs and background music from a text prompt.Suno, Lyria 3, ElevenLabs Music, Stable Audio 3, ACE-Step
Tabular data
Models for rows and columns, where most business prediction still happens.
- Gradient-boosted treesChurn, fraud, pricing and risk predictions on spreadsheet-style data. Still the baseline to beat.XGBoost, LightGBM, CatBoost
- Tabular foundation modelsPredictions on small and medium tables without a training run.TabPFN, TabICL, Google TabFM, Mitra
- Anomaly detectionFlagging unusual transactions, sensor readings or log events.Isolation Forest, autoencoders, One-Class SVM
Time series
Models that forecast values over time.
- Classical forecastingSales, demand and traffic forecasts with trend and seasonality.ARIMA, ETS, Prophet
- Time series foundation modelsForecasting a new series with no training on it.TimesFM, Chronos-2, Moirai, Toto, TimeGPT
Recommendation and ranking
Models that decide what each user sees in a feed, a store or a search page.
- Two-tower retrievalPicking a few hundred candidates out of millions of items for a feed or search page.Two-tower models, matrix factorization
- Learning to rankOrdering search results and feed items by predicted relevance or clicks.LambdaMART, DLRM, Wide & Deep
- Generative recommendersTreating a user's history as a sequence and predicting the next item, the way an LLM predicts the next token.Meta HSTU, Kuaishou OneRec
Graphs
Models for connected data, where the links between things matter as much as the things.
- Graph neural networksPredictions on connected data such as fraud rings, molecules and social networks.GCN, GraphSAGE, GAT
- Knowledge graphs and GraphRAGAnswering questions that depend on relationships spread across many documents.Microsoft GraphRAG, Neo4j
3D and spatial
Models that reconstruct, understand or generate 3D scenes and objects.
- Gaussian Splatting and NeRFTurning photos or video of a real place into a 3D scene you can view from any angle.3D Gaussian Splatting, NeRF
- Feed-forward 3D reconstructionCamera positions, depth and point clouds from a handful of photos in one pass.VGGT, Depth Anything 3, MapAnything, DUSt3R
- 3D generationCreating 3D assets from an image or a text prompt.TRELLIS.2, Hunyuan3D, SAM 3D
- Point cloud modelsDetecting objects in lidar scans for self-driving and robotics.PointNet++, PointPillars
Robotics and world models
Models that act in the physical world or simulate it.
- Vision-language-action models (VLAs)Controlling a robot from camera input and a written or spoken instruction.π0.5, GR00T N1.6, Gemini Robotics 2, SmolVLA, OpenVLA
- World modelsSimulating how a scene changes over time, to train and test agents and robots.Genie 3, NVIDIA Cosmos 3, V-JEPA 2.1, World Labs Marble
How they work
Architectures
How the models above work inside. Read one when a model page mentions it.
- Transformer and attentionRead →The architecture behind LLMs, vision transformers, Whisper and most modern models.
- Encoder-only, decoder-only and encoder-decoderRead →Why BERT, GPT and T5 are built differently, and what each shape is good at.
- Mixture of Experts (MoE)How very large models run only a fraction of their weights for each token.
- State-space models and hybridsHandling long inputs with less memory than attention needs.Mamba, Mamba-2, Nemotron 3, Jamba
- Convolutional neural networks (CNNs)The classic image architecture, still inside most fast vision models.
- Vision Transformer (ViT)Read →A transformer that reads an image as a sequence of patches.
- U-NetPixel-level output for segmentation, and the denoiser in early diffusion models.
- RNNs and LSTMsThe sequence models that came before transformers, still used for small time-series jobs.
- Diffusion modelsRead →How image and video generators turn random noise into a picture, step by step.
- Flow matching and diffusion transformers (DiT)The training recipe and backbone behind current image and video models like FLUX.
- Variational autoencoders (VAEs)Compressing images into the smaller latent space that diffusion models work in.
- GANsA generator trained against a discriminator, still used for upscaling and fast generation.
- Autoregressive image generationGenerating an image token by token with the same machinery as an LLM.GPT Image, HunyuanImage 3.0
- Contrastive learningTraining two encoders so matching pairs land close together, the idea behind CLIP and embedding models.
- Self-supervised learningLearning from unlabeled data by predicting hidden parts of it.DINO, MAE, JEPA
- JEPARead →Meta's approach of predicting the embeddings of hidden parts of an image or video, used for video understanding and robotics.I-JEPA, V-JEPA 2.1, VL-JEPA
- Bi-encoders and cross-encodersRead →Why embedding models are fast and rerankers are accurate, and how search systems use both.
- Omni modelsOne model that takes in and produces any mix of text, images, audio and video.Gemini Omni, Qwen3-Omni, FLUX 3, NVIDIA Cosmos 3
The modules cover how to build production systems around these models, from evals and retrieval to agents and guardrails.
