MLGuerrillaStart with M1 →
Text and language·20 min read·Updated 24 September 2026

Reasoning models

A reasoning model writes out its working before it answers, and you pay for every token of it. On twenty questions run through the same model twice, thinking cost eleven times the tokens and got one fewer right. On harder benchmarks the same model gains a lot from thinking, so the budget is something to measure per task.

A reasoning model is a language model trained to write out a long stretch of working before it gives you an answer. The working is called the thinking, it usually arrives in its own block or behind its own field in the API, and it is made of ordinary generated tokens that you are billed for and that you wait for.

The category started as a separate kind of model and is turning into a setting. OpenAI's o1 and DeepSeek-R1 shipped as models that always think. Qwen3, Claude and Gemini now take a flag, so the same weights answer either way, which makes the cost of thinking something you can measure directly.

What it is

The thinking is a second budget, spent before the answer

Running the same model over the same questions twice, with nothing changed but that flag, is the cleanest version of the measurement. These are twenty questions in four bands, from arithmetic a child could do up to counting the trailing zeros of 100 factorial.

A figure titled "Thinking mode cost eleven times the tokens and got one fewer right", from twenty questions asked twice through qwen3:8b with the same weights, the same seed and one flag changed. The main panel breaks the twenty into four bands of five. For questions needing one step or none, such as 7 times 8 and the capital of Australia, thinking off used 21 tokens and got 5 of 5 right, and thinking on used 1,210 tokens and also got 5 of 5. For questions needing a few steps, such as arrival times, counting letters and a rectangle, thinking off used 438 tokens for 5 of 5 and thinking on used 10,221 tokens for 4 of 5. For questions needing several steps, such as modular arithmetic, coin counting and algebra, thinking off used 1,664 tokens for 5 of 5 and thinking on used 13,324 tokens for 5 of 5. For questions needing more steps still, such as the trailing zeros of 100 factorial and units digits, thinking off used 1,355 tokens for 5 of 5 and thinking on used 12,729 tokens for 5 of 5. Across all twenty, thinking off used 3,478 tokens in 150 seconds and got 20 of 20 right, and thinking on used 37,484 tokens in 1,561 seconds and got 19 of 20, which is 10.8 times the tokens and 10.4 times the wall clock. Two worked examples follow. For what is 7 times 8, thinking off took 3 tokens and 0.2 seconds and thinking on took 204 tokens and 8.3 seconds, and both answered 56. For the socks question, outlined in terracotta, thinking off took 69 tokens and 3.0 seconds and answered 3, while thinking on took 4,096 tokens and 173.8 seconds and answered nothing, because it wrote 15,290 characters of thinking, hit its 4,096-token output limit and returned an empty string. A final panel, headed where it does pay and labelled as coming from the paper rather than this run, notes that DeepSeek trained a model to think with reinforcement learning alone and its score on the 2024 AIME competition went from 15.6 percent to 77.9 percent as it learned to write longer working, with nobody telling it to write more, and that competition mathematics is where the extra tokens buy something while a laptop-sized model on ordinary questions is not. The caption reads: the thinking is a second budget spent before the answer, and it is charged whether it helps or not.
Both columns are the same 8B model at temperature 0 with the same seed. The only difference between them is a boolean in the request.

The one-step band is the part to look at first. Twenty-one tokens became 1,210, which is fifty-eight times as many, to produce the same five answers. "What is 7 times 8" took 3 tokens and 0.2 seconds with the flag off, and 204 tokens and 8.3 seconds with it on.

The harder bands do not rescue it. On questions about modular arithmetic and factorials, thinking off answered all five correctly using 1,664 tokens, because the model writes out its working anyway when the question needs it. Thinking on used eight times that and got the same five right.

The failure is the one that should worry you. Asked how many socks you must draw from two red and three blue to be sure of a pair, the model with thinking on produced 15,290 characters of deliberation, hit its token limit, and returned an empty answer. With thinking off it answered "3" in 69 tokens. A model that talks itself past its output limit and returns nothing is a production incident, and it is a failure mode that only exists once thinking is switched on.

Every run here used temperature 0, which always picks the most likely next token, so the same question gets the same answer each time. The Qwen3-8B model card says "DO NOT use greedy decoding" in thinking mode, and warns that it can cause worse answers and endless repetition, which is what the socks run looks like. The card recommends temperature 0.6 with a larger output limit, and the Qwen team's own evaluations allowed 32,768 tokens where this run allowed 4,096 or 6,144. So the socks failure shows what happens with settings you might pick for reproducibility. It doesn't show how often thinking fails at the recommended settings, and that run hasn't been done here.

Where it pays

Where thinking changes the answer

The twenty questions above could only show the cost, because thinking off already got all twenty right and there was nothing left for thinking to fix. To see the benefit you need questions the model gets wrong without thinking. The Qwen team measured exactly that for the same 8B model, on benchmarks much harder than anything in the figure.

Qwen3-8B with thinking off and on, from the Qwen3 technical report, Tables 17 and 18
BenchmarkWhat it testsThinking offThinking on
MATH-500Competition-style maths problems87.497.4
AIME'24A 2024 US invitational maths exam29.176.0
AIME'25The 2025 exam of the same kind20.967.3
ZebraLogicLogic-grid puzzles with many constraints26.784.8
LiveCodeBench v5Recent competitive-programming problems22.857.5

These are the report's numbers, and they haven't been reproduced on this laptop. The report ran them with sampled decoding and allowed up to 32,768 output tokens, or 38,912 on AIME, which is roughly five to ten times the caps in the figure. So the same weights that wasted tokens on "7 times 8" more than doubled their AIME score when the problems were hard enough and the budget was large enough.

A budget sweep on questions the model gets wrong

The local version of that measurement starts by finding questions the model misses with thinking off, then gives it several thinking budgets on each one. The budget here works the way the Qwen3 report describes. The model thinks until it closes the block itself or reaches the budget, and at the budget the early-stop sentence from the report is inserted so an answer still comes out. Budget 0 means an empty thinking block. The prompts asked for the number only, so at budget 0 the model wrote no working at all.

The partial sweep on qwen3:8b, temperature 0, 2026-09-24. Each cell is the answer, then total tokens and seconds
QuestionCorrect answerBudget 0Budget 512Budget 2,048Budget 4,096
4817 times 392618,911,54218,923,642, 9 tok, 0.6 s18,914,682, 521 tok, 24 s, cap hit18,914,682, 2,057 tok, 88 s, cap hit18,923,642, 2,988 tok, 149 s, finished early
The letter e in "nevertheless the referee presented seventeen excellent sentences"2214, 3 tok, 0.4 s14, 515 tok, 26 s, cap hit14, 2,051 tok, 96 s, cap hit14, 3,142 tok, 161 s, finished early
7389 times 624746,159,08346,064,013, 9 tok, 0.9 s46,134,033, 521 tok, 43 s, cap hit46,134,033, 2,057 tok, 101 s, cap hitnot run

None of the three went from wrong to right. On the first multiplication, 2,979 tokens of thinking that ended on their own came back to the exact wrong answer the model gave with no thinking at all. The letter count stayed at 14 at every budget. That is three questions, and the run stopped before the rest of the planned questions finished, so it can't say how often a budget helps. It does show that more thinking doesn't help on every question the model gets wrong. A likely reason, which this run didn't test, is that long multiplication and counting letters depend on digit-by-digit and letter-by-letter work that an 8B model does poorly whether it thinks or not.

That leaves this article without a worked example, run on this machine, of thinking turning a wrong answer into a right one. The table above from the Qwen3 report says those questions exist for this model at contest difficulty. Finding one locally would take the recommended sampling settings, larger budgets and a question pool built from contest-style problems, which is what the sweep below is designed to do.

Choosing a budget on your own test cases

The budget is a trade between three numbers you can measure, which are the share of test cases answered correctly, the latency, and the tokens you pay for. The only way to know where your workload sits is to run it.

  • Take 50 to 100 test cases from real traffic, each with an answer you can check automatically.
  • Run every test case at budgets 0, 512, 2,048 and 8,192, and with no cap below the provider's maximum, using the provider's recommended sampling settings. Sampled answers vary, so run each test case three times or more.
  • For each budget, record accuracy, median and 95th-percentile latency, total tokens, and how often the budget was reached.
  • List every test case whose answer changed between budgets, in both directions. A budget that fixes five answers and breaks two nets three, and the two it broke are worth reading.
  • Count an empty answer or a reached cap with no answer as a failure, never as a skipped test case.
  • Pick the smallest budget whose accuracy is close enough to the best one for your product and whose 95th-percentile latency fits the wait you can afford.

Where it shows up

Where reasoning models show up

Competition mathematics and hard proofs

The job the category was built for, and where the published gains are largest. DeepSeek-AI (2025) reports its model going from 15.6% to 77.9% pass@1 on the 2024 AIME competition through reinforcement learning alone. The gain is not limited to huge models. The Qwen3 report scores the same 8B model measured in this article at 29.1 on AIME'24 with thinking off and 76.0 with it on, with up to 38,912 output tokens allowed per problem and sampled decoding.

Code that has to be right the first time

Writing out the edge cases before writing the function can catch the bug you would have shipped. On LiveCodeBench v5, a benchmark of recent competitive-programming problems, the Qwen3 report scores Qwen3-8B at 22.8 with thinking off and 57.5 with it on. Those are contest problems with tests, so they say more about algorithmic code than about the glue code most applications are made of.

Planning inside agents

An agent choosing which tool to call, in what order, with what arguments, is doing multi-step work where a wrong first move costs a whole loop. The thinking budget buys a check before the action.

Scientific and legal analysis with many constraints

Problems where a dozen conditions all have to hold at once are the shape where thinking has the clearest published gains. The nearest benchmark is ZebraLogic, a dataset of logic-grid puzzles, where the Qwen3 report scores the same 8B model at 26.7 with thinking off and 84.8 with it on. A model without thinking still generates its answer one token at a time, so the difference is how much working it writes before it commits to an answer. Legal analysis has no comparable public number, so treat it as a place to test.

Distilling into something smaller

DeepSeek-R1-Distill-Qwen-7B is a 7B model trained on a large model's thinking traces. The paper's own framing is that the patterns a big model discovers "can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models".

How it works

Reinforcement learning, think tags, and why the working got longer

It is trained, and it is not a prompt

Asking a model to think step by step is chain-of-thought prompting, and it has been around since 2022. A reasoning model has the behaviour trained in, so it does it without being asked and does it more thoroughly.

DeepSeek's route was reinforcement learning against answers that can be checked. The model produces working and an answer, the answer is marked right or wrong, and the reward flows back. The paper's claim is that reasoning can be "incentivized through pure reinforcement learning, obviating the need for human-labeled reasoning trajectories", so nobody had to write out worked solutions for it to imitate.

The thinking block is a trained output format

The visible thinking is not a separate system. The paper describes a format reward where "the model is incentivized to encapsulate its reasoning process within designated tags, specifically <think> and </think>", so the block you see is the model emitting two learned strings around a stretch of ordinary generation.

That is why it shows up as tokens on your bill and as seconds on your clock. It is the same decoding loop the LLM article describes, run for longer before anything is shown to you.

Nobody told it to write more

The length came out of the training. The paper reports that "DeepSeek-R1-Zero naturally learns to solve reasoning tasks with more thinking time", with average response length climbing through the RL run. Longer working scored better, so the model wrote longer working.

The paper also records an aha moment, where an intermediate model stops mid-solution and reconsiders its approach in a conversational tone. That behaviour was never demonstrated to it.

Hidden thinking, and what you are charged for

Some hosted models return their thinking and some summarise or withhold it, and in both cases the tokens are billed. Read the provider's pricing page for whether reasoning tokens count as output, and what the reasoning-effort or thinking-budget parameter does to them, and check the date on whatever you read.

Budget caps are the control worth setting, and the socks failure shows why the cap has to sit on the thinking. That run did have a limit, 4,096 tokens across everything the model generated, and the thinking used all of it before the answer started.

What a thinking budget does to the model

A thinking budget is a cap on the thinking tokens alone, and what happens at the cap decides whether you get an answer. A plain output limit like Ollama's num_predict counts the thinking and the answer together and stops generation outright, so when the thinking runs long the answer is what gets cut off. The Qwen3 report describes a gentler version. When the thinking reaches the budget, the serving code stops it and inserts the sentence "Considering the limited time by the user, I have to give the solution based on the thinking directly now." followed by the closing </think> tag, and the model then writes an answer from the working it has so far. The budget ends the thinking early enough to leave room for that answer.

The report says this ability was not trained directly. It came out of training one model on both modes, so a half-finished thinking block followed by an answer is something the model already knows how to handle. The report's own scaling curves, for its largest model on mathematics, coding and science benchmarks, show scores rising smoothly as the budget grows. It also reports one place where thinking made the model worse, which is long-context retrieval on the RULER benchmark, and the authors' guess is that the thinking interferes with finding the right passage.

Versions

The models you'll see, as of September 2026

Open models, with monthly downloads read 2026-09-22
ModelReleasedLicenseParamsDownloads a month
deepseek-ai/DeepSeek-R1January 2025MIT684.5B794K
deepseek-ai/DeepSeek-R1-Distill-Qwen-7BJanuary 2025MIT7.6B245K
Qwen/QwQ-32BMarch 2025Apache-2.032.8B74K
Qwen/Qwen3-8B, which carries a thinking flagApril 2025Apache-2.08.2B12.7M

That last row is the story of the category. A general model with a thinking switch is downloaded sixteen times as often as the most famous dedicated reasoning model and 170 times as often as QwQ. The separate-model era lasted about a year.

Among hosted models, OpenAI, Anthropic and Google all expose reasoning as a parameter on their general models now. Read each provider's docs for the parameter name and what it costs, and date what you read, because this is the fastest-moving row on this index.

Choosing

Choosing, which means measuring where to start

The table below gives a starting point for each kind of job and how strong the evidence behind it is. The twenty-question run is one 8B model at temperature 0 on short questions, so it can show that thinking wasted tokens on those questions. It can't show that a whole category of work never benefits, and a larger model or a harder version of the same job could come out the other way. Treat every row as the first setting to test with the sweep above.

Where to start, and what the starting point rests on
The jobStart withWhat that rests on
Short classification, extraction, routingOff, then test a small budgetThis run, where one-step questions cost fifty-eight times the tokens for identical answers. That is one model on five easy questions
Rewriting, summarising, translatingOff, then testNo measurement here. A reasonable guess, since these jobs have few steps to get wrong
Anything user-facing and interactiveOff or a tight cap, then test8.3 seconds to answer 7 times 8 on a laptop CPU. Hosted models are faster, so measure your own latency
Competition-style mathematics and proofsOn, with a large budgetThe Qwen3 report, where the same 8B model went from 29.1 to 76.0 on AIME'24, and DeepSeek-R1's 15.6% to 77.9%
Algorithmic code with testsOn, then test the budgetThe Qwen3 report's LiveCodeBench v5 scores, 22.8 off and 57.5 on. Everyday application code hasn't been measured
Puzzles and analysis with many constraintsOn, then test the budgetZebraLogic in the Qwen3 report, 26.7 off and 84.8 on
Agent planning between tool callsTest both, with a capNo measurement here. A wrong first action costs a whole loop, which is the argument for thinking, and each step's latency adds up, which is the argument against
Structured output against a schemaTest, and check the formatA long thinking block ahead of JSON needs parsing that keeps the two apart
Long-document retrievalOff, then testThe Qwen3 report found thinking mode slightly worse on the RULER long-context benchmark
Anything with a hard latency budgetOff or a tight capThe wall clock in this run was 10.4 times longer with thinking on

Turn it on for a measured reason. A sensible default is off or a small budget, with larger budgets for the calls where you have seen thinking change an answer on your own test cases, and a cap on every call so a runaway block ends in an answer.

Try it

How to try it

The comparison behind the figure is one boolean, and running it on your own prompts takes a few minutes.

compare.py — the same question, asked both ways
import json, time, urllib.request

def ask(prompt, think):
    body = json.dumps({"model": "qwen3:8b", "prompt": prompt, "think": think,
                       "stream": False,
                       "options": {"temperature": 0, "seed": 0,
                                   "num_predict": 4096}}).encode()
    req = urllib.request.Request("http://localhost:11434/api/generate", body,
                                 {"Content-Type": "application/json"})
    t0 = time.time()
    out = json.load(urllib.request.urlopen(req, timeout=1800))
    return {"secs": round(time.time() - t0, 1),
            "tokens": out.get("eval_count", 0),          # thinking included
            "thinking": len(out.get("thinking") or ""),
            "answer": (out.get("response") or "").strip(),
            "stopped_early": out.get("done_reason") == "length"}

for flag in (False, True):
    print(flag, ask("What is 7 times 8? Answer with the number only.", flag))

Watch stopped_early in particular. That flag is how the socks failure shows up in a log, and a run that returns an empty answer will otherwise look like a model that had nothing to say.

Ask your AI coding tool

Measure what thinking mode costs on my own prompts before I turn it on. Take a CSV of prompts with the expected answer for each, and send every prompt twice to a local qwen3:8b through the Ollama generate API at temperature 0 and seed 0, once with think false and once with think true, using the same num_predict for both. For each prompt record the wall time, the eval_count tokens, the length of the thinking field, whether done_reason was length, and whether the answer matches the expected one. Then report, per prompt and in total: the token ratio, the wall-clock ratio, the accuracy with and without thinking, and a list of every prompt where thinking changed the answer in either direction. Flag any prompt where thinking mode hit the token cap and returned an empty answer.

The on-and-off comparison shows the cost. To find the budget, run the sweep from earlier on your own test cases.

Ask your AI coding tool

Help me pick a thinking budget for qwen3:8b on my own test cases. Take a CSV of prompts with a checkable expected answer for each. Call the Ollama generate API with raw set to true and build the Qwen3 chat template myself, so I control the thinking block. For budget 0, send an empty think block. For each budget B in 512, 2048 and 8192, generate thinking with num_predict B and stop on the closing think tag. If it stops because of the length limit, append the Qwen3 report's early-stop sentence and the closing think tag, then generate the answer. Use temperature 0.6, top_p 0.95 and top_k 20, and run every prompt three times per budget with different seeds. For each budget report accuracy, median and 95th-percentile latency, total tokens, and how often the budget was reached. List every prompt whose answer changed between budgets, in either direction, and count any empty answer as wrong.

Limits

What it can't do

  • The thinking is not free and is often wasted. On twenty easy questions here it cost 10.8 times the tokens and 10.4 times the wall clock, and changed no answer from wrong to right, because thinking off already had all twenty right.
  • It can talk itself out of an answer. One question produced 15,290 characters of deliberation, hit the token cap and returned nothing. Cap the budget and log when the cap is reached.
  • The thinking is not an explanation you can trust. It is generated text that was rewarded for leading to correct answers, and a model can reach the right answer through working that does not hold up.
  • It does not fix a model that lacks the knowledge. More steps over a fact the model never learned produces a longer wrong answer.
  • A bigger budget doesn't fix every wrong answer. In the partial sweep, a letter count stayed wrong at every budget up to 4,096, and a multiplication came back to the same wrong answer after thinking finished on its own.
  • Latency makes it unusable in some products. 8.3 seconds for an arithmetic answer rules out anything a person is waiting on.
  • This run is one 8B model on twenty questions, decoded greedily. It shows the cost clearly and it is nowhere near a benchmark. The gains in this article come from the DeepSeek-R1 and Qwen3 reports, on far harder problems with far larger budgets, and they haven't been reproduced here. The one local attempt to find a gain, the partial budget sweep, finished three questions and found none.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.