MLGuerrillaStart with M1 →

Reference

Models and architectures

These are the models you run into when you build AI systems, grouped by the kind of data they work with. Each one has a line on what it's used for, so when a problem comes up you can remember which model handles it. The architectures at the end explain how the models work inside.

83 entries · checked 18 September 2026 · 24 written in full, more each week

By modality

Text and language

Models that read and write text. Most AI products start here.

Embeddings and retrieval

Models that turn content into vectors so you can search by meaning. They are the core of RAG.

Vision

Models that look at images and video frames and report what is there and where.

Vision-language

Models that connect images and text, so you can search images with words or ask questions about a picture.

Documents and OCR

Models that turn scans, PDFs and forms into text and structure you can use.

  • OCR enginesReading printed text from scans and photos.Tesseract, PaddleOCR, docTR
  • Document parsing modelsConverting PDFs with tables, formulas and complex layouts into clean Markdown or JSON.PaddleOCR-VL, DeepSeek-OCR 2, MinerU, olmOCR 2, Docling, Mistral OCR
  • Layout-aware modelsPulling fields out of forms and invoices using both the text and where it sits on the page.LayoutLMv3, Donut

Image generation and editing

Models that create or edit images from text and reference images.

  • Frontier image modelsTop-quality generation and editing through an API, including readable text inside images.GPT Image 2.5, Nano Banana 2, Seedream 5.0, Midjourney, Imagen
  • Stable DiffusionOpen image generation you can run locally and customize.SD 1.5, SDXL, SD 3.5
  • FLUXHigh-quality image generation and editing, open or through an API.FLUX.2, FLUX.1 Kontext, FLUX 3
  • Other open image modelsSelf-hosted generation and editing beyond Stable Diffusion and FLUX.Qwen-Image, Z-Image, HunyuanImage 3.0, Ideogram 4.0
  • ControlNet, LoRA and adaptersSteering generation with a pose, edge or depth map, and teaching a model a new style or subject from a few images.ControlNet, LoRA, IP-Adapter, DreamBooth
  • Upscaling, inpainting and background removalThe cleanup steps in an image pipeline.Real-ESRGAN, LaMa, BiRefNet

Video

Models that understand video or generate it.

  • Video understanding modelsRecognizing actions, searching footage, and answering questions about clips.V-JEPA 2.1, VideoMAE, Molmo 2, Qwen3-VL, Gemini
  • Video generation modelsGenerating clips from text or images, with synchronized sound in the newest models.Veo 3.1, Gemini Omni, Kling 3.0, Seedance 2.5, Runway Gen-4.5
  • Open video modelsSelf-hosted video generation and fine-tuning.Wan 2.2, LTX-2.5, HunyuanVideo 1.5
  • Talking avatars and lip-syncMaking a face speak a given audio track, for dubbing and video presenters.OmniHuman 1.5, HeyGen Avatar V, InfiniteTalk, LatentSync

Speech and audio

Models that listen, speak and make sound. Together they make up a voice agent.

  • WhisperRead →The default open model for transcribing speech in many languages.Whisper large-v3, large-v3-turbo, faster-whisper
  • Speech-to-text (ASR)Transcribing calls, meetings and voice commands, in batch or live.Parakeet, Qwen3-ASR, Cohere Transcribe, Deepgram Nova-3, AssemblyAI
  • Text-to-speech (TTS)Giving an app a voice, from narration to cloned voices.ElevenLabs v3, Cartesia Sonic, Kokoro, Qwen3-TTS, Chatterbox
  • Realtime speech-to-speechVoice agents that listen and answer in one model, with low latency.gpt-realtime, Gemini Live, Qwen3-Omni, Moshi
  • Voice activity and turn detectionKnowing when someone is speaking and when they have finished, so a voice agent doesn't talk over them.Silero VAD, Pipecat Smart Turn, LiveKit turn detector
  • Speaker diarizationLabeling who spoke when in a recording.pyannote
  • Audio classification and embeddingsRecognizing sounds and searching audio by a text description.CLAP, BEATs, AST
  • Music generationGenerating songs and background music from a text prompt.Suno, Lyria 3, ElevenLabs Music, Stable Audio 3, ACE-Step

Tabular data

Models for rows and columns, where most business prediction still happens.

  • Gradient-boosted treesChurn, fraud, pricing and risk predictions on spreadsheet-style data. Still the baseline to beat.XGBoost, LightGBM, CatBoost
  • Tabular foundation modelsPredictions on small and medium tables without a training run.TabPFN, TabICL, Google TabFM, Mitra
  • Anomaly detectionFlagging unusual transactions, sensor readings or log events.Isolation Forest, autoencoders, One-Class SVM

Time series

Models that forecast values over time.

  • Classical forecastingSales, demand and traffic forecasts with trend and seasonality.ARIMA, ETS, Prophet
  • Time series foundation modelsForecasting a new series with no training on it.TimesFM, Chronos-2, Moirai, Toto, TimeGPT

Recommendation and ranking

Models that decide what each user sees in a feed, a store or a search page.

  • Two-tower retrievalPicking a few hundred candidates out of millions of items for a feed or search page.Two-tower models, matrix factorization
  • Learning to rankOrdering search results and feed items by predicted relevance or clicks.LambdaMART, DLRM, Wide & Deep
  • Generative recommendersTreating a user's history as a sequence and predicting the next item, the way an LLM predicts the next token.Meta HSTU, Kuaishou OneRec

Graphs

Models for connected data, where the links between things matter as much as the things.

  • Graph neural networksPredictions on connected data such as fraud rings, molecules and social networks.GCN, GraphSAGE, GAT
  • Knowledge graphs and GraphRAGAnswering questions that depend on relationships spread across many documents.Microsoft GraphRAG, Neo4j

3D and spatial

Models that reconstruct, understand or generate 3D scenes and objects.

  • Gaussian Splatting and NeRFTurning photos or video of a real place into a 3D scene you can view from any angle.3D Gaussian Splatting, NeRF
  • Feed-forward 3D reconstructionCamera positions, depth and point clouds from a handful of photos in one pass.VGGT, Depth Anything 3, MapAnything, DUSt3R
  • 3D generationCreating 3D assets from an image or a text prompt.TRELLIS.2, Hunyuan3D, SAM 3D
  • Point cloud modelsDetecting objects in lidar scans for self-driving and robotics.PointNet++, PointPillars

Robotics and world models

Models that act in the physical world or simulate it.

  • Vision-language-action models (VLAs)Controlling a robot from camera input and a written or spoken instruction.π0.5, GR00T N1.6, Gemini Robotics 2, SmolVLA, OpenVLA
  • World modelsSimulating how a scene changes over time, to train and test agents and robots.Genie 3, NVIDIA Cosmos 3, V-JEPA 2.1, World Labs Marble

How they work

Architectures

How the models above work inside. Read one when a model page mentions it.

The modules cover how to build production systems around these models, from evals and retrieval to agents and guardrails.

Browse the modulesBack to the overview