MLGuerrillaStart with M1 →
Text and language·20 min read·Updated 23 September 2026

Large language models (LLMs)

An LLM takes text as input and writes text back one token at a time. It is the model behind chat assistants, coding agents and most of the AI features you use.

A large language model, or LLM, is a neural network trained to predict the next token of text. A token is a chunk of text, usually a word or part of a word. The model predicts a token, adds it to the end of the input and predicts again, and everything it produces, from a chat reply to a tool call, comes out of that same loop.

The "large" refers to the number of parameters, the learned numbers inside the model. GPT-3 had 175 billion of them in 2020, and Moonshot AI's Kimi K3, released with open weights in 2026, has about 2.8 trillion. This page covers how an LLM works inside and the papers the design comes from. It also lists the models you'll see as of September 2026, and when a different kind of model fits the job better.

What it is

An LLM turns text into a probability for every possible next token

The clearest way to see what an LLM does is to run a small one and look at what comes out. Alibaba's Qwen2.5-3B-Instruct runs on a laptop through Ollama, and the input below is the raw start of a sentence with no chat formatting around it, "The customer was charged twice for the same order, so this support ticket should go to the". What comes back is a probability for every entry in the model's vocabulary, which is 151,936 tokens for this one, and those probabilities add up to 1.

A figure titled "An LLM's output is a probability for every token in its vocabulary". The top panel shows the prompt "The customer was charged twice for the same order, so this support ticket should go to the" going into Qwen2.5-3B, and a horizontal bar chart of the ten most likely next tokens with their probabilities: " billing" 30.2% in terracotta, then " ____" 7.8%, " __" 4.0%, " ___" 3.7%, " complaint" 3.6%, " Billing" 3.6%, " [" 2.6%, " issue" 2.5%, " fraud" 2.5% and " customer" 2.3%. A note says these ten add up to 62.9%, and the other 37.1% is spread across the remaining 151,926 tokens. The bottom panel shows eight generation steps at temperature 0, each picking the top token and adding it to the input: " billing" 30.2%, " department" 65.4%, "." 47.5%, " A" 24.7%, "." 94.2%, " Correct" 65.5%, " B" 99.8% and "." 100%. A bracket under the first three steps says the model completes the sentence, and a bracket under the last five says it is now writing a quiz. The caption says that with no chat template around the input, nothing marks where an answer should end, so it keeps going.
The underscore tokens and the quiz most likely come from the same place in the training data, fill-in-the-blank exercises with sentences shaped like this one.

The top answer, " billing", gets 30.2%. " Billing" with a capital B is a separate token and gets another 3.6%, and three of the top ten are runs of underscores. The ten most likely tokens only add up to 62.9%, and the other 37.1% is spread across the rest of the vocabulary.

To produce text, the model picks one token from that distribution, adds it to the input and runs again. When it takes the most likely token every time, which is what a temperature of 0 does, this model writes " billing department. A. Correct B." and keeps going. It completes the sentence and then writes what looks like a multiple-choice quiz, and within 60 tokens it has moved on to an unrelated quiz question about power supply systems, because it's continuing the text the way its training data would.

This model has been post-trained to answer and stop, and that behavior depends on its chat template, the special tokens that mark where a user's message ends and the assistant's reply begins. Sent through the template, the same sentence gets a 72-token answer that starts "This support ticket should be directed to the billing or customer service team", and then the model writes its end-of-turn token, which tells the software to stop. The section on how it works covers post-training, and How LLMs Actually Run Things walks through tokens and this loop step by step.

Most frontier LLMs now take images as well as text as input. As of September 18, 2026, Anthropic's and OpenAI's model docs list text and image input with text output for all their current models. The amount of input a model can take at once is its context window. That's about 1 million tokens for OpenAI's and Anthropic's current models, apart from Claude Haiku 4.5 at 200,000.

Where it shows up

LLMs do most of the text work in AI products, from chat to coding agents

Most AI features you'll build call an LLM somewhere. The same model handles all of the jobs below, and what changes between them is what goes into the input and what your code does with the output.

Chat assistants

A chat assistant is the loop from the last section with a conversation around it. On every turn, your code sends the whole conversation so far and the model writes the next reply.

Writing and summarizing

The model drafts a reply to the ticket, or summarizes a 40-message thread for whoever picks it up next. Translation works the same way, since it's text in and text out.

Writing code

Coding assistants and coding agents, like Claude Code and OpenAI's Codex, are LLMs that read and edit code, usually with tools that let them run it and see the result.

Pulling structured data out of text

An LLM can read an invoice email and write back JSON with the amount and the due date. Providers offer structured outputs, a mode that makes the model's JSON match a schema you give it. The Model as a Component covers how to get JSON your code can trust.

Answering from your own documents

The model only knows what was in its training data, which stops at a cutoff date. Retrieval-augmented generation, or RAG, from Lewis et al. (2020), fetches the passages from your documents that match the question and puts them in the input, so the model answers from them.

Agents that use tools

The model can write a tool call, like a request to look up order A-104, and your code runs it and sends the result back as more input. ReAct (Yao et al., 2022) showed the pattern of alternating reasoning steps with actions, and today's agents run on that loop.

How it works

Inside, an LLM is a stack of transformer layers trained on trillions of tokens

Text becomes a list of vectors

The tokenizer splits the input into tokens and swaps each one for its ID in the vocabulary. Most LLMs build their vocabulary with byte-pair encoding, which Sennrich et al. (2015) adapted for neural models. It starts from single characters and keeps merging the most frequent pair of neighbors into a new token until the vocabulary is the size it wants. The 90-character prompt above came to 18 tokens for Qwen2.5-3B.

Each token ID then picks one row from a table of learned vectors, called embeddings. In Qwen2.5-3B each row holds 2,048 numbers, so the 18 tokens become 18 vectors of 2,048 numbers each.

Layers of attention let each token read the ones before it

The vectors then pass through a stack of identical layers, 36 of them in Qwen2.5-3B. Each layer has two parts:

  • Attention, the mechanism from Attention Is All You Need (Vaswani et al., 2017), lets each token's vector take in information from the tokens before it, weighted by how relevant each one is. In an LLM, attention is causal, meaning each token can only read the tokens before it, because during generation the later ones don't exist yet.
  • A feed-forward network then transforms each token's vector on its own.

The original transformer had two halves, an encoder that read the input and a decoder that wrote the output, because it was built for translation. GPT-1 (Radford et al., 2018) kept only the decoder and trained it to predict the next token, and almost every LLM since is built that way. The transformer article covers attention step by step.

After the last layer, the vector at the final position goes through one more matrix that gives one score per vocabulary entry, so 151,936 scores for Qwen2.5-3B. Softmax, the function that turns scores into probabilities that sum to 1, gives the distribution in the first figure.

A left-to-right figure titled "One step inside Qwen2.5-3B, with its real sizes", made of six cards joined by arrows. TEXT shows 90 characters, the prompt from above. TOKENIZER shows 18 token IDs, one ID per chunk of text. EMBEDDINGS shows 18 × 2,048 numbers, one row of numbers per token. A dark stacked card, 36 LAYERS, contains an attention block and a feed-forward block, with the note that each token reads only earlier tokens. OUTPUT shows 151,936 scores, one score per vocabulary entry. SOFTMAX, outlined in terracotta, shows the token " billing" at 30.2%, from probabilities that sum to 1. A terracotta arrow runs from SOFTMAX back to EMBEDDINGS, labeled "pick one token, add it to the input, and run it through all 36 layers". The caption says every new token goes through the whole stack again, so a 500-token answer is 500 passes.
The sizes come from Qwen2.5-3B's published config. Frontier models follow the same steps at a much larger size, and the closed ones don't publish their exact numbers.

Sampling picks the next token

With a temperature of 0 the model always takes the most likely token. On one machine answering one request at a time, like the Ollama run above, the same input then usually gives the same output. A hosted API can still return different outputs at temperature 0, because the server batches your request with other people's, and a different batch size changes the order of the floating-point arithmetic enough to flip a close choice between two tokens. He and Thinking Machines Lab (2025) sent the same prompt to Qwen3-235B 1,000 times at temperature 0 and got 80 different completions, and all 1,000 matched once they switched to kernels that give the same result at any batch size.

A higher temperature flattens the distribution before picking, so less likely tokens get chosen more often. Top-p, or nucleus sampling from Holtzman et al. (2019), samples only from the smallest group of tokens whose probabilities add up to p, which cuts off the long tail, like the 37.1% spread across the rest of Qwen's vocabulary.

Providers have started taking these settings away on their newest models. On Claude Sonnet 5 and Claude Opus 4.7 and later, setting temperature, top_p or top_k to anything but the default returns an error, according to Anthropic's release notes. Google deprecated the same three for its latest Gemini models in July 2026. The Model as a Component covers the settings that are left.

Pretraining teaches the model language by predicting the next token

An LLM starts as a transformer with random weights. Pretraining shows it trillions of tokens of text and, after every prediction, nudges the weights so the real next token gets a higher probability. It needs no labels, since the next token of real text is the answer.

GPT-2 (Radford et al., 2019), with 1.5 billion parameters, showed that a model trained this way begins to learn tasks like translation and summarization "without any explicit supervision". GPT-3 (Brown et al., 2020), with 175 billion parameters, showed it could pick up a new task from a few examples written into the prompt, with no retraining.

Kaplan et al. (2020) found that the loss, a measure of how wrong the model's predictions are, "scales as a power-law with model size, dataset size, and the amount of compute used for training". Hoffmann et al. (2022) then found that the large models of the time were undertrained, and that model size and training data should grow together. Their 70-billion-parameter Chinchilla, trained on 1.4 trillion tokens with the same compute budget as the 280-billion-parameter Gopher, beat Gopher across a large range of tasks. Training data has kept growing since, and the Qwen2.5 technical report says the family behind the figures on this page was pretrained on 18 trillion tokens.

Post-training turns a text predictor into an assistant

A pretrained model only continues text. Post-training teaches it to answer a request and stop, and it usually happens in stages.

The usual post-training stages
StageWhat the model trains onWhat changesWhere it comes from
Supervised fine-tuning (SFT)Written examples of good answers to promptsIt learns the request-and-answer format, and when to stopOuyang et al. (2022), the InstructGPT paper
Preference tuning (RLHF or DPO)Pairs of answers where a person marked the better oneIts answers move toward what people preferRLHF in the InstructGPT paper, and DPO in Rafailov et al. (2023)
Reinforcement learning with verifiable rewards (RLVR)Problems a program can check, like math answers and code with testsIt learns to work through long problems before it answersDeepSeek-R1 (DeepSeek-AI, 2025)

In the InstructGPT paper, people preferred the answers of a 1.3-billion-parameter model after post-training to those of the 175-billion-parameter GPT-3, "despite having 100x fewer parameters". Post-training also teaches the model its chat template, which is how a chat API turns your list of messages into one sequence of tokens. The Qwen run at the top of this page skipped the template, so the model had no user message to answer and no place to end its turn, and it wrote a quiz.

Reasoning models come out of the last stage. Wei et al. (2022) showed that prompting a model to write out intermediate steps, called chain of thought, improves its answers on math and logic problems. DeepSeek-R1 then showed that reinforcement learning alone can train a model to do that without being prompted. The reasoning models entry on this index covers them.

Serving reads the input in one pass and writes the output one token at a time

Your whole prompt goes through the model in one parallel pass, called prefill. The output comes one token per pass, called decoding. Each decoding pass reuses the stored attention results for every earlier token, a store called the KV cache, so the model doesn't recompute the prompt for every new token. vLLM (Kwon et al., 2023) is a serving system built around managing that cache across many requests at once.

That split is why output costs more. As of September 18, 2026, output tokens cost 5 to 6 times as much as input tokens across OpenAI's and Anthropic's current models. The range runs from GPT-5.6 Luna at $0.20 per million input tokens and $1.20 per million output tokens up to GPT-6 Astra and Claude Fable 5.1 at $10 and $50.

Many of the largest open-weight models are also mixtures of experts. Each layer holds many feed-forward networks, called experts, and a small router picks a few of them for each token. DeepSeek-V4.1-Flash's config lists 384 experts with 6 chosen per token, so each token runs through a small share of the model's 763 billion parameters. The idea goes back to Shazeer et al. (2017), and the mixture-of-experts entry on this index covers it.

Versions

The LLMs you'll see, as of September 18, 2026

Closed models run only on their maker's API and the clouds it partners with. Open-weight models publish their weights, so you can run them on your own hardware or through a hosting provider.

Closed frontier LLMs (vendor docs and changelogs, as of September 18, 2026)
MakerCurrent modelsNewest release
OpenAIGPT-6 Astra, and GPT-5.6 Sol, Terra and LunaGPT-6 Astra, September 3, 2026
AnthropicClaude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5Claude Fable 5.1, September 1, 2026
GoogleGemini 3.8 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.1 Pro PreviewGemini 3.8 Flash, September 2, 2026
xAIGrok 4.6Grok 4.6, August 12, 2026
Open-weight LLMs (Hugging Face model pages, as of September 18, 2026)
ModelMakerParametersLicense
DeepSeek-V4.1-FlashDeepSeek763B, mixture of expertsMIT
Kimi K3Moonshot AI2.78T, mixture of expertsMoonshot's own license
GLM-5.3Z.ai753B, mixture of expertsZ.ai's own license
Gemma 4 31BGoogle31.3BApache 2.0
Qwen3.8-27BAlibaba27.8BApache 2.0
gpt-oss-120bOpenAI117BApache 2.0

The parameter counts are the totals Hugging Face lists for each model's weights. Rankings between these models change every few weeks, so the only comparison that settles a choice is running your own test cases on the ones you're considering. The open-weight LLMs entry on this index covers running them yourself.

Choosing

Pick an LLM when the output is language, and a neighbor when the answer is narrower

An LLM is the most general tool on this index. Next to the other text models here it's also slow and expensive per decision, and its output is text your code has to check. So the question for any job is whether you need its generality.

Which model to reach for
The jobReach forWhy
Write a reply or write codeAn LLMThe output is language, and only a generative model writes it
Answer from your own documentsAn LLM with retrievalThe model writes the answer, and retrieval supplies the facts
Route or score text into a list you can write down, at high volumeJevIt returns a probability for each of your options in one pass, with no text to parse
Classify into fixed labels you have thousands of examples forA fine-tuned encoderIt's small and fast, and it runs on your own hardware
Find similar documents or search by meaningA text embedding modelIt turns text into vectors you can compare, with no generation
Solve a multi-step math or coding problemA reasoning model, or an LLM's reasoning modeIt works through the problem before it answers
Keep data on your own servers, or cut the cost of high volumeAn open-weight or small LLMYou run it yourself and pay for hardware, with no per-token bill

Routing a ticket with an LLM works, and the Qwen example shows what you're working against, a distribution over 151,936 tokens with " billing" and " Billing" counted separately. Structured outputs can force the answer into your options, and you still pay for generating it. Jev and a fine-tuned encoder both give you a probability over your own labels directly.

Which of them is most accurate on your task is something only a test shows. A fine-tuned encoder's accuracy comes from your labeled examples, and Jev's speed and quality numbers are TypeSafe's own so far. So the table says where each one is worth testing, and the choice comes from running the same labeled test cases through each candidate.

Model size is a choice of its own. A small model is cheaper and faster, and it usually fails more often on hard inputs. The usual way to find out is to prototype with the strongest model first, because if it can't do the job a smaller general model prompted the same way almost certainly can't either, then test smaller ones against the same test cases. A small model fine-tuned on your own examples is a separate case, and on a narrow task it can beat a much larger prompted one, like the 149M encoder that beat a prompted 3B model on the BERT article's 77-way task.

Try it

How to try it

Every major provider has an SDK that looks roughly like this. This is Anthropic's Python SDK (pip install anthropic) calling Claude Opus 5. The client reads your API key from the ANTHROPIC_API_KEY environment variable.

ask.py — one call to a hosted LLM
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=16000,
    messages=[{
        "role": "user",
        "content": "In one sentence, which team should handle this ticket? "
                   "'I was charged twice for order A-104.'",
    }],
)

# Opus 5 thinks before it answers by default, so print only the text blocks.
for block in response.content:
    if block.type == "text":
        print(block.text)
print(response.usage.input_tokens, "tokens in,", response.usage.output_tokens, "out")

To see the probabilities from the first figure yourself, install Ollama, run ollama pull qwen2.5:3b, and run this script. It asks for the next token only, with the ten most likely candidates. The numbers can come out slightly different on other hardware or Ollama versions.

next_token.py — the top-10 next tokens from a local model
import math

import requests

prompt = ("The customer was charged twice for the same order, "
          "so this support ticket should go to the")

r = requests.post("http://localhost:11434/api/generate", json={
    "model": "qwen2.5:3b",
    "prompt": prompt,
    "raw": True,          # no chat template, just the text
    "stream": False,
    "options": {"num_predict": 1, "temperature": 0},
    "logprobs": True,
    "top_logprobs": 10,
})

for t in r.json()["logprobs"][0]["top_logprobs"]:
    print(f"{t['token']!r:>14}  {math.exp(t['logprob']):.1%}")
Ask your AI coding tool

Write a Python script that shows how temperature changes an LLM's next-token choice. Use Ollama's /api/generate endpoint with qwen2.5:3b, raw: true, logprobs: true and top_logprobs: 20. For the prompt I pass on the command line, get the top 20 next tokens once, then recompute their distribution at temperatures 0.2, 0.7, 1.0 and 1.5 by dividing the log-probabilities by the temperature and renormalizing over those 20 tokens. Print a table with one row per token and one column per temperature. Then sample 200 next tokens from Ollama at temperature 1.0 and print how often each token came up next to its probability.

Limits

What it can't do

  • It makes things up. An LLM can write a fluent, confident statement that's false, including invented citations and quotes. Kalai et al. (2025) argue that this happens because "the training and evaluation procedures reward guessing over acknowledging uncertainty".
  • Its stated confidence isn't a probability you can use. Xiong et al. (2023) found that LLMs "tend to be overconfident" when they state their confidence in words, and the GPT-4 Technical Report found that post-training made GPT-4's own token probabilities less calibrated.
  • It can't count letters or do exact arithmetic reliably. It reads tokens, and counting or adding up many small pieces is work a single pass doesn't do reliably. Use code for anything code can compute exactly.
  • It doesn't know anything after its training cutoff. As of September 18, 2026, Anthropic lists a reliable knowledge cutoff of May 2026 for Claude Opus 5, and OpenAI lists April 30, 2026 for GPT-6 Astra.
  • It keeps nothing between calls. Every call starts from only what's in its input, so memory is something your code builds by resending history or retrieving facts.
  • It reads long inputs unevenly. Liu et al. (2023) found that accuracy was often highest when the relevant information sat at the start or end of the input, and dropped when it sat in the middle. Those were 2023 models and newer ones may do better, but a long input is still a reason to test where your key facts land.
  • It can't reliably tell your instructions from instructions inside the text it reads. A web page or a ticket that says "ignore your previous instructions" can change what it does. AI Security & Guardrails covers defending against this.
  • The same prompt can give different answers. Sampling adds randomness, and on a hosted API batching can change the answer even at temperature 0. An alias like gpt-5.6 can also move to a new model underneath you.
  • Long outputs are slow and expensive. Every output token is one more pass through the whole model, and output tokens cost 5 to 6 times as much as input.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.