What it is
An LLM turns text into a probability for every possible next token
The clearest way to see what an LLM does is to run a small one and look at what comes out. Alibaba's Qwen2.5-3B-Instruct runs on a laptop through Ollama, and the input below is the raw start of a sentence with no chat formatting around it, "The customer was charged twice for the same order, so this support ticket should go to the". What comes back is a probability for every entry in the model's vocabulary, which is 151,936 tokens for this one, and those probabilities add up to 1.

The top answer, " billing", gets 30.2%. " Billing" with a capital B is a separate token and gets another 3.6%, and three of the top ten are runs of underscores. The ten most likely tokens only add up to 62.9%, and the other 37.1% is spread across the rest of the vocabulary.
To produce text, the model picks one token from that distribution, adds it to the input and runs again. When it takes the most likely token every time, which is what a temperature of 0 does, this model writes " billing department. A. Correct B." and keeps going. It completes the sentence and then writes what looks like a multiple-choice quiz, and within 60 tokens it has moved on to an unrelated quiz question about power supply systems, because it's continuing the text the way its training data would.
This model has been post-trained to answer and stop, and that behavior depends on its chat template, the special tokens that mark where a user's message ends and the assistant's reply begins. Sent through the template, the same sentence gets a 72-token answer that starts "This support ticket should be directed to the billing or customer service team", and then the model writes its end-of-turn token, which tells the software to stop. The section on how it works covers post-training, and How LLMs Actually Run Things walks through tokens and this loop step by step.
Most frontier LLMs now take images as well as text as input. As of September 18, 2026, Anthropic's and OpenAI's model docs list text and image input with text output for all their current models. The amount of input a model can take at once is its context window. That's about 1 million tokens for OpenAI's and Anthropic's current models, apart from Claude Haiku 4.5 at 200,000.
Where it shows up
LLMs do most of the text work in AI products, from chat to coding agents
Most AI features you'll build call an LLM somewhere. The same model handles all of the jobs below, and what changes between them is what goes into the input and what your code does with the output.
Chat assistants
A chat assistant is the loop from the last section with a conversation around it. On every turn, your code sends the whole conversation so far and the model writes the next reply.
Writing and summarizing
The model drafts a reply to the ticket, or summarizes a 40-message thread for whoever picks it up next. Translation works the same way, since it's text in and text out.
Writing code
Coding assistants and coding agents, like Claude Code and OpenAI's Codex, are LLMs that read and edit code, usually with tools that let them run it and see the result.
Pulling structured data out of text
An LLM can read an invoice email and write back JSON with the amount and the due date. Providers offer structured outputs, a mode that makes the model's JSON match a schema you give it. The Model as a Component covers how to get JSON your code can trust.
Answering from your own documents
The model only knows what was in its training data, which stops at a cutoff date. Retrieval-augmented generation, or RAG, from Lewis et al. (2020), fetches the passages from your documents that match the question and puts them in the input, so the model answers from them.
Agents that use tools
The model can write a tool call, like a request to look up order A-104, and your code runs it and sends the result back as more input. ReAct (Yao et al., 2022) showed the pattern of alternating reasoning steps with actions, and today's agents run on that loop.
How it works
Inside, an LLM is a stack of transformer layers trained on trillions of tokens
Text becomes a list of vectors
The tokenizer splits the input into tokens and swaps each one for its ID in the vocabulary. Most LLMs build their vocabulary with byte-pair encoding, which Sennrich et al. (2015) adapted for neural models. It starts from single characters and keeps merging the most frequent pair of neighbors into a new token until the vocabulary is the size it wants. The 90-character prompt above came to 18 tokens for Qwen2.5-3B.
Each token ID then picks one row from a table of learned vectors, called embeddings. In Qwen2.5-3B each row holds 2,048 numbers, so the 18 tokens become 18 vectors of 2,048 numbers each.
Layers of attention let each token read the ones before it
The vectors then pass through a stack of identical layers, 36 of them in Qwen2.5-3B. Each layer has two parts:
- Attention, the mechanism from Attention Is All You Need (Vaswani et al., 2017), lets each token's vector take in information from the tokens before it, weighted by how relevant each one is. In an LLM, attention is causal, meaning each token can only read the tokens before it, because during generation the later ones don't exist yet.
- A feed-forward network then transforms each token's vector on its own.
The original transformer had two halves, an encoder that read the input and a decoder that wrote the output, because it was built for translation. GPT-1 (Radford et al., 2018) kept only the decoder and trained it to predict the next token, and almost every LLM since is built that way. The transformer article covers attention step by step.
After the last layer, the vector at the final position goes through one more matrix that gives one score per vocabulary entry, so 151,936 scores for Qwen2.5-3B. Softmax, the function that turns scores into probabilities that sum to 1, gives the distribution in the first figure.

Sampling picks the next token
With a temperature of 0 the model always takes the most likely token. On one machine answering one request at a time, like the Ollama run above, the same input then usually gives the same output. A hosted API can still return different outputs at temperature 0, because the server batches your request with other people's, and a different batch size changes the order of the floating-point arithmetic enough to flip a close choice between two tokens. He and Thinking Machines Lab (2025) sent the same prompt to Qwen3-235B 1,000 times at temperature 0 and got 80 different completions, and all 1,000 matched once they switched to kernels that give the same result at any batch size.
A higher temperature flattens the distribution before picking, so less likely tokens get chosen more often. Top-p, or nucleus sampling from Holtzman et al. (2019), samples only from the smallest group of tokens whose probabilities add up to p, which cuts off the long tail, like the 37.1% spread across the rest of Qwen's vocabulary.
Providers have started taking these settings away on their newest models. On Claude Sonnet 5 and Claude Opus 4.7 and later, setting temperature, top_p or top_k to anything but the default returns an error, according to Anthropic's release notes. Google deprecated the same three for its latest Gemini models in July 2026. The Model as a Component covers the settings that are left.
Pretraining teaches the model language by predicting the next token
An LLM starts as a transformer with random weights. Pretraining shows it trillions of tokens of text and, after every prediction, nudges the weights so the real next token gets a higher probability. It needs no labels, since the next token of real text is the answer.
GPT-2 (Radford et al., 2019), with 1.5 billion parameters, showed that a model trained this way begins to learn tasks like translation and summarization "without any explicit supervision". GPT-3 (Brown et al., 2020), with 175 billion parameters, showed it could pick up a new task from a few examples written into the prompt, with no retraining.
Kaplan et al. (2020) found that the loss, a measure of how wrong the model's predictions are, "scales as a power-law with model size, dataset size, and the amount of compute used for training". Hoffmann et al. (2022) then found that the large models of the time were undertrained, and that model size and training data should grow together. Their 70-billion-parameter Chinchilla, trained on 1.4 trillion tokens with the same compute budget as the 280-billion-parameter Gopher, beat Gopher across a large range of tasks. Training data has kept growing since, and the Qwen2.5 technical report says the family behind the figures on this page was pretrained on 18 trillion tokens.
Post-training turns a text predictor into an assistant
A pretrained model only continues text. Post-training teaches it to answer a request and stop, and it usually happens in stages.
| Stage | What the model trains on | What changes | Where it comes from |
|---|---|---|---|
| Supervised fine-tuning (SFT) | Written examples of good answers to prompts | It learns the request-and-answer format, and when to stop | Ouyang et al. (2022), the InstructGPT paper |
| Preference tuning (RLHF or DPO) | Pairs of answers where a person marked the better one | Its answers move toward what people prefer | RLHF in the InstructGPT paper, and DPO in Rafailov et al. (2023) |
| Reinforcement learning with verifiable rewards (RLVR) | Problems a program can check, like math answers and code with tests | It learns to work through long problems before it answers | DeepSeek-R1 (DeepSeek-AI, 2025) |
In the InstructGPT paper, people preferred the answers of a 1.3-billion-parameter model after post-training to those of the 175-billion-parameter GPT-3, "despite having 100x fewer parameters". Post-training also teaches the model its chat template, which is how a chat API turns your list of messages into one sequence of tokens. The Qwen run at the top of this page skipped the template, so the model had no user message to answer and no place to end its turn, and it wrote a quiz.
Reasoning models come out of the last stage. Wei et al. (2022) showed that prompting a model to write out intermediate steps, called chain of thought, improves its answers on math and logic problems. DeepSeek-R1 then showed that reinforcement learning alone can train a model to do that without being prompted. The reasoning models entry on this index covers them.
Serving reads the input in one pass and writes the output one token at a time
Your whole prompt goes through the model in one parallel pass, called prefill. The output comes one token per pass, called decoding. Each decoding pass reuses the stored attention results for every earlier token, a store called the KV cache, so the model doesn't recompute the prompt for every new token. vLLM (Kwon et al., 2023) is a serving system built around managing that cache across many requests at once.
That split is why output costs more. As of September 18, 2026, output tokens cost 5 to 6 times as much as input tokens across OpenAI's and Anthropic's current models. The range runs from GPT-5.6 Luna at $0.20 per million input tokens and $1.20 per million output tokens up to GPT-6 Astra and Claude Fable 5.1 at $10 and $50.
Many of the largest open-weight models are also mixtures of experts. Each layer holds many feed-forward networks, called experts, and a small router picks a few of them for each token. DeepSeek-V4.1-Flash's config lists 384 experts with 6 chosen per token, so each token runs through a small share of the model's 763 billion parameters. The idea goes back to Shazeer et al. (2017), and the mixture-of-experts entry on this index covers it.
Versions
The LLMs you'll see, as of September 18, 2026
Closed models run only on their maker's API and the clouds it partners with. Open-weight models publish their weights, so you can run them on your own hardware or through a hosting provider.
| Maker | Current models | Newest release |
|---|---|---|
| OpenAI | GPT-6 Astra, and GPT-5.6 Sol, Terra and Luna | GPT-6 Astra, September 3, 2026 |
| Anthropic | Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 | Claude Fable 5.1, September 1, 2026 |
| Gemini 3.8 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.1 Pro Preview | Gemini 3.8 Flash, September 2, 2026 | |
| xAI | Grok 4.6 | Grok 4.6, August 12, 2026 |
| Model | Maker | Parameters | License |
|---|---|---|---|
| DeepSeek-V4.1-Flash | DeepSeek | 763B, mixture of experts | MIT |
| Kimi K3 | Moonshot AI | 2.78T, mixture of experts | Moonshot's own license |
| GLM-5.3 | Z.ai | 753B, mixture of experts | Z.ai's own license |
| Gemma 4 31B | 31.3B | Apache 2.0 | |
| Qwen3.8-27B | Alibaba | 27.8B | Apache 2.0 |
| gpt-oss-120b | OpenAI | 117B | Apache 2.0 |
The parameter counts are the totals Hugging Face lists for each model's weights. Rankings between these models change every few weeks, so the only comparison that settles a choice is running your own test cases on the ones you're considering. The open-weight LLMs entry on this index covers running them yourself.
Choosing
Pick an LLM when the output is language, and a neighbor when the answer is narrower
An LLM is the most general tool on this index. Next to the other text models here it's also slow and expensive per decision, and its output is text your code has to check. So the question for any job is whether you need its generality.
| The job | Reach for | Why |
|---|---|---|
| Write a reply or write code | An LLM | The output is language, and only a generative model writes it |
| Answer from your own documents | An LLM with retrieval | The model writes the answer, and retrieval supplies the facts |
| Route or score text into a list you can write down, at high volume | Jev | It returns a probability for each of your options in one pass, with no text to parse |
| Classify into fixed labels you have thousands of examples for | A fine-tuned encoder | It's small and fast, and it runs on your own hardware |
| Find similar documents or search by meaning | A text embedding model | It turns text into vectors you can compare, with no generation |
| Solve a multi-step math or coding problem | A reasoning model, or an LLM's reasoning mode | It works through the problem before it answers |
| Keep data on your own servers, or cut the cost of high volume | An open-weight or small LLM | You run it yourself and pay for hardware, with no per-token bill |
Routing a ticket with an LLM works, and the Qwen example shows what you're working against, a distribution over 151,936 tokens with " billing" and " Billing" counted separately. Structured outputs can force the answer into your options, and you still pay for generating it. Jev and a fine-tuned encoder both give you a probability over your own labels directly.
Which of them is most accurate on your task is something only a test shows. A fine-tuned encoder's accuracy comes from your labeled examples, and Jev's speed and quality numbers are TypeSafe's own so far. So the table says where each one is worth testing, and the choice comes from running the same labeled test cases through each candidate.
Model size is a choice of its own. A small model is cheaper and faster, and it usually fails more often on hard inputs. The usual way to find out is to prototype with the strongest model first, because if it can't do the job a smaller general model prompted the same way almost certainly can't either, then test smaller ones against the same test cases. A small model fine-tuned on your own examples is a separate case, and on a narrow task it can beat a much larger prompted one, like the 149M encoder that beat a prompted 3B model on the BERT article's 77-way task.
Try it
How to try it
Every major provider has an SDK that looks roughly like this. This is Anthropic's Python SDK (pip install anthropic) calling Claude Opus 5. The client reads your API key from the ANTHROPIC_API_KEY environment variable.
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
messages=[{
"role": "user",
"content": "In one sentence, which team should handle this ticket? "
"'I was charged twice for order A-104.'",
}],
)
# Opus 5 thinks before it answers by default, so print only the text blocks.
for block in response.content:
if block.type == "text":
print(block.text)
print(response.usage.input_tokens, "tokens in,", response.usage.output_tokens, "out")To see the probabilities from the first figure yourself, install Ollama, run ollama pull qwen2.5:3b, and run this script. It asks for the next token only, with the ten most likely candidates. The numbers can come out slightly different on other hardware or Ollama versions.
import math
import requests
prompt = ("The customer was charged twice for the same order, "
"so this support ticket should go to the")
r = requests.post("http://localhost:11434/api/generate", json={
"model": "qwen2.5:3b",
"prompt": prompt,
"raw": True, # no chat template, just the text
"stream": False,
"options": {"num_predict": 1, "temperature": 0},
"logprobs": True,
"top_logprobs": 10,
})
for t in r.json()["logprobs"][0]["top_logprobs"]:
print(f"{t['token']!r:>14} {math.exp(t['logprob']):.1%}")Write a Python script that shows how temperature changes an LLM's next-token choice. Use Ollama's /api/generate endpoint with qwen2.5:3b, raw: true, logprobs: true and top_logprobs: 20. For the prompt I pass on the command line, get the top 20 next tokens once, then recompute their distribution at temperatures 0.2, 0.7, 1.0 and 1.5 by dividing the log-probabilities by the temperature and renormalizing over those 20 tokens. Print a table with one row per token and one column per temperature. Then sample 200 next tokens from Ollama at temperature 1.0 and print how often each token came up next to its probability.
Limits
What it can't do
- It makes things up. An LLM can write a fluent, confident statement that's false, including invented citations and quotes. Kalai et al. (2025) argue that this happens because "the training and evaluation procedures reward guessing over acknowledging uncertainty".
- Its stated confidence isn't a probability you can use. Xiong et al. (2023) found that LLMs "tend to be overconfident" when they state their confidence in words, and the GPT-4 Technical Report found that post-training made GPT-4's own token probabilities less calibrated.
- It can't count letters or do exact arithmetic reliably. It reads tokens, and counting or adding up many small pieces is work a single pass doesn't do reliably. Use code for anything code can compute exactly.
- It doesn't know anything after its training cutoff. As of September 18, 2026, Anthropic lists a reliable knowledge cutoff of May 2026 for Claude Opus 5, and OpenAI lists April 30, 2026 for GPT-6 Astra.
- It keeps nothing between calls. Every call starts from only what's in its input, so memory is something your code builds by resending history or retrieving facts.
- It reads long inputs unevenly. Liu et al. (2023) found that accuracy was often highest when the relevant information sat at the start or end of the input, and dropped when it sat in the middle. Those were 2023 models and newer ones may do better, but a long input is still a reason to test where your key facts land.
- It can't reliably tell your instructions from instructions inside the text it reads. A web page or a ticket that says "ignore your previous instructions" can change what it does. AI Security & Guardrails covers defending against this.
- The same prompt can give different answers. Sampling adds randomness, and on a hosted API batching can change the answer even at temperature 0. An alias like
gpt-5.6can also move to a new model underneath you. - Long outputs are slow and expensive. Every output token is one more pass through the whole model, and output tokens cost 5 to 6 times as much as input.
Go deeper
Vaswani et al. (2017): Attention Is All You Need · Radford et al. (2018): Improving Language Understanding by Generative Pre-Training (GPT-1) · Brown et al. (2020): Language Models are Few-Shot Learners (GPT-3) · Kaplan et al. (2020): Scaling Laws for Neural Language Models · Hoffmann et al. (2022): Training Compute-Optimal Large Language Models (Chinchilla) · Ouyang et al. (2022): Training language models to follow instructions with human feedback · DeepSeek-AI (2025): DeepSeek-R1 · Kalai et al. (2025): Why Language Models Hallucinate
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
