What it is
The thinking is a second budget, spent before the answer
Running the same model over the same questions twice, with nothing changed but that flag, is the cleanest version of the measurement. These are twenty questions in four bands, from arithmetic a child could do up to counting the trailing zeros of 100 factorial.

The one-step band is the part to look at first. Twenty-one tokens became 1,210, which is fifty-eight times as many, to produce the same five answers. "What is 7 times 8" took 3 tokens and 0.2 seconds with the flag off, and 204 tokens and 8.3 seconds with it on.
The harder bands do not rescue it. On questions about modular arithmetic and factorials, thinking off answered all five correctly using 1,664 tokens, because the model writes out its working anyway when the question needs it. Thinking on used eight times that and got the same five right.
The failure is the one that should worry you. Asked how many socks you must draw from two red and three blue to be sure of a pair, the model with thinking on produced 15,290 characters of deliberation, hit its token limit, and returned an empty answer. With thinking off it answered "3" in 69 tokens. A model that talks itself past its output limit and returns nothing is a production incident, and it is a failure mode that only exists once thinking is switched on.
Every run here used temperature 0, which always picks the most likely next token, so the same question gets the same answer each time. The Qwen3-8B model card says "DO NOT use greedy decoding" in thinking mode, and warns that it can cause worse answers and endless repetition, which is what the socks run looks like. The card recommends temperature 0.6 with a larger output limit, and the Qwen team's own evaluations allowed 32,768 tokens where this run allowed 4,096 or 6,144. So the socks failure shows what happens with settings you might pick for reproducibility. It doesn't show how often thinking fails at the recommended settings, and that run hasn't been done here.
Where it pays
Where thinking changes the answer
The twenty questions above could only show the cost, because thinking off already got all twenty right and there was nothing left for thinking to fix. To see the benefit you need questions the model gets wrong without thinking. The Qwen team measured exactly that for the same 8B model, on benchmarks much harder than anything in the figure.
| Benchmark | What it tests | Thinking off | Thinking on |
|---|---|---|---|
| MATH-500 | Competition-style maths problems | 87.4 | 97.4 |
| AIME'24 | A 2024 US invitational maths exam | 29.1 | 76.0 |
| AIME'25 | The 2025 exam of the same kind | 20.9 | 67.3 |
| ZebraLogic | Logic-grid puzzles with many constraints | 26.7 | 84.8 |
| LiveCodeBench v5 | Recent competitive-programming problems | 22.8 | 57.5 |
These are the report's numbers, and they haven't been reproduced on this laptop. The report ran them with sampled decoding and allowed up to 32,768 output tokens, or 38,912 on AIME, which is roughly five to ten times the caps in the figure. So the same weights that wasted tokens on "7 times 8" more than doubled their AIME score when the problems were hard enough and the budget was large enough.
A budget sweep on questions the model gets wrong
The local version of that measurement starts by finding questions the model misses with thinking off, then gives it several thinking budgets on each one. The budget here works the way the Qwen3 report describes. The model thinks until it closes the block itself or reaches the budget, and at the budget the early-stop sentence from the report is inserted so an answer still comes out. Budget 0 means an empty thinking block. The prompts asked for the number only, so at budget 0 the model wrote no working at all.
| Question | Correct answer | Budget 0 | Budget 512 | Budget 2,048 | Budget 4,096 |
|---|---|---|---|---|---|
| 4817 times 3926 | 18,911,542 | 18,923,642, 9 tok, 0.6 s | 18,914,682, 521 tok, 24 s, cap hit | 18,914,682, 2,057 tok, 88 s, cap hit | 18,923,642, 2,988 tok, 149 s, finished early |
| The letter e in "nevertheless the referee presented seventeen excellent sentences" | 22 | 14, 3 tok, 0.4 s | 14, 515 tok, 26 s, cap hit | 14, 2,051 tok, 96 s, cap hit | 14, 3,142 tok, 161 s, finished early |
| 7389 times 6247 | 46,159,083 | 46,064,013, 9 tok, 0.9 s | 46,134,033, 521 tok, 43 s, cap hit | 46,134,033, 2,057 tok, 101 s, cap hit | not run |
None of the three went from wrong to right. On the first multiplication, 2,979 tokens of thinking that ended on their own came back to the exact wrong answer the model gave with no thinking at all. The letter count stayed at 14 at every budget. That is three questions, and the run stopped before the rest of the planned questions finished, so it can't say how often a budget helps. It does show that more thinking doesn't help on every question the model gets wrong. A likely reason, which this run didn't test, is that long multiplication and counting letters depend on digit-by-digit and letter-by-letter work that an 8B model does poorly whether it thinks or not.
That leaves this article without a worked example, run on this machine, of thinking turning a wrong answer into a right one. The table above from the Qwen3 report says those questions exist for this model at contest difficulty. Finding one locally would take the recommended sampling settings, larger budgets and a question pool built from contest-style problems, which is what the sweep below is designed to do.
Choosing a budget on your own test cases
The budget is a trade between three numbers you can measure, which are the share of test cases answered correctly, the latency, and the tokens you pay for. The only way to know where your workload sits is to run it.
- Take 50 to 100 test cases from real traffic, each with an answer you can check automatically.
- Run every test case at budgets 0, 512, 2,048 and 8,192, and with no cap below the provider's maximum, using the provider's recommended sampling settings. Sampled answers vary, so run each test case three times or more.
- For each budget, record accuracy, median and 95th-percentile latency, total tokens, and how often the budget was reached.
- List every test case whose answer changed between budgets, in both directions. A budget that fixes five answers and breaks two nets three, and the two it broke are worth reading.
- Count an empty answer or a reached cap with no answer as a failure, never as a skipped test case.
- Pick the smallest budget whose accuracy is close enough to the best one for your product and whose 95th-percentile latency fits the wait you can afford.
Where it shows up
Where reasoning models show up
Competition mathematics and hard proofs
The job the category was built for, and where the published gains are largest. DeepSeek-AI (2025) reports its model going from 15.6% to 77.9% pass@1 on the 2024 AIME competition through reinforcement learning alone. The gain is not limited to huge models. The Qwen3 report scores the same 8B model measured in this article at 29.1 on AIME'24 with thinking off and 76.0 with it on, with up to 38,912 output tokens allowed per problem and sampled decoding.
Code that has to be right the first time
Writing out the edge cases before writing the function can catch the bug you would have shipped. On LiveCodeBench v5, a benchmark of recent competitive-programming problems, the Qwen3 report scores Qwen3-8B at 22.8 with thinking off and 57.5 with it on. Those are contest problems with tests, so they say more about algorithmic code than about the glue code most applications are made of.
Planning inside agents
An agent choosing which tool to call, in what order, with what arguments, is doing multi-step work where a wrong first move costs a whole loop. The thinking budget buys a check before the action.
Scientific and legal analysis with many constraints
Problems where a dozen conditions all have to hold at once are the shape where thinking has the clearest published gains. The nearest benchmark is ZebraLogic, a dataset of logic-grid puzzles, where the Qwen3 report scores the same 8B model at 26.7 with thinking off and 84.8 with it on. A model without thinking still generates its answer one token at a time, so the difference is how much working it writes before it commits to an answer. Legal analysis has no comparable public number, so treat it as a place to test.
Distilling into something smaller
DeepSeek-R1-Distill-Qwen-7B is a 7B model trained on a large model's thinking traces. The paper's own framing is that the patterns a big model discovers "can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models".
Versions
The models you'll see, as of September 2026
| Model | Released | License | Params | Downloads a month |
|---|---|---|---|---|
deepseek-ai/DeepSeek-R1 | January 2025 | MIT | 684.5B | 794K |
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | January 2025 | MIT | 7.6B | 245K |
Qwen/QwQ-32B | March 2025 | Apache-2.0 | 32.8B | 74K |
Qwen/Qwen3-8B, which carries a thinking flag | April 2025 | Apache-2.0 | 8.2B | 12.7M |
That last row is the story of the category. A general model with a thinking switch is downloaded sixteen times as often as the most famous dedicated reasoning model and 170 times as often as QwQ. The separate-model era lasted about a year.
Among hosted models, OpenAI, Anthropic and Google all expose reasoning as a parameter on their general models now. Read each provider's docs for the parameter name and what it costs, and date what you read, because this is the fastest-moving row on this index.
Choosing
Choosing, which means measuring where to start
The table below gives a starting point for each kind of job and how strong the evidence behind it is. The twenty-question run is one 8B model at temperature 0 on short questions, so it can show that thinking wasted tokens on those questions. It can't show that a whole category of work never benefits, and a larger model or a harder version of the same job could come out the other way. Treat every row as the first setting to test with the sweep above.
| The job | Start with | What that rests on |
|---|---|---|
| Short classification, extraction, routing | Off, then test a small budget | This run, where one-step questions cost fifty-eight times the tokens for identical answers. That is one model on five easy questions |
| Rewriting, summarising, translating | Off, then test | No measurement here. A reasonable guess, since these jobs have few steps to get wrong |
| Anything user-facing and interactive | Off or a tight cap, then test | 8.3 seconds to answer 7 times 8 on a laptop CPU. Hosted models are faster, so measure your own latency |
| Competition-style mathematics and proofs | On, with a large budget | The Qwen3 report, where the same 8B model went from 29.1 to 76.0 on AIME'24, and DeepSeek-R1's 15.6% to 77.9% |
| Algorithmic code with tests | On, then test the budget | The Qwen3 report's LiveCodeBench v5 scores, 22.8 off and 57.5 on. Everyday application code hasn't been measured |
| Puzzles and analysis with many constraints | On, then test the budget | ZebraLogic in the Qwen3 report, 26.7 off and 84.8 on |
| Agent planning between tool calls | Test both, with a cap | No measurement here. A wrong first action costs a whole loop, which is the argument for thinking, and each step's latency adds up, which is the argument against |
| Structured output against a schema | Test, and check the format | A long thinking block ahead of JSON needs parsing that keeps the two apart |
| Long-document retrieval | Off, then test | The Qwen3 report found thinking mode slightly worse on the RULER long-context benchmark |
| Anything with a hard latency budget | Off or a tight cap | The wall clock in this run was 10.4 times longer with thinking on |
Turn it on for a measured reason. A sensible default is off or a small budget, with larger budgets for the calls where you have seen thinking change an answer on your own test cases, and a cap on every call so a runaway block ends in an answer.
Try it
How to try it
The comparison behind the figure is one boolean, and running it on your own prompts takes a few minutes.
import json, time, urllib.request
def ask(prompt, think):
body = json.dumps({"model": "qwen3:8b", "prompt": prompt, "think": think,
"stream": False,
"options": {"temperature": 0, "seed": 0,
"num_predict": 4096}}).encode()
req = urllib.request.Request("http://localhost:11434/api/generate", body,
{"Content-Type": "application/json"})
t0 = time.time()
out = json.load(urllib.request.urlopen(req, timeout=1800))
return {"secs": round(time.time() - t0, 1),
"tokens": out.get("eval_count", 0), # thinking included
"thinking": len(out.get("thinking") or ""),
"answer": (out.get("response") or "").strip(),
"stopped_early": out.get("done_reason") == "length"}
for flag in (False, True):
print(flag, ask("What is 7 times 8? Answer with the number only.", flag))Watch stopped_early in particular. That flag is how the socks failure shows up in a log, and a run that returns an empty answer will otherwise look like a model that had nothing to say.
Measure what thinking mode costs on my own prompts before I turn it on. Take a CSV of prompts with the expected answer for each, and send every prompt twice to a local qwen3:8b through the Ollama generate API at temperature 0 and seed 0, once with think false and once with think true, using the same num_predict for both. For each prompt record the wall time, the eval_count tokens, the length of the thinking field, whether done_reason was length, and whether the answer matches the expected one. Then report, per prompt and in total: the token ratio, the wall-clock ratio, the accuracy with and without thinking, and a list of every prompt where thinking changed the answer in either direction. Flag any prompt where thinking mode hit the token cap and returned an empty answer.
The on-and-off comparison shows the cost. To find the budget, run the sweep from earlier on your own test cases.
Help me pick a thinking budget for qwen3:8b on my own test cases. Take a CSV of prompts with a checkable expected answer for each. Call the Ollama generate API with raw set to true and build the Qwen3 chat template myself, so I control the thinking block. For budget 0, send an empty think block. For each budget B in 512, 2048 and 8192, generate thinking with num_predict B and stop on the closing think tag. If it stops because of the length limit, append the Qwen3 report's early-stop sentence and the closing think tag, then generate the answer. Use temperature 0.6, top_p 0.95 and top_k 20, and run every prompt three times per budget with different seeds. For each budget report accuracy, median and 95th-percentile latency, total tokens, and how often the budget was reached. List every prompt whose answer changed between budgets, in either direction, and count any empty answer as wrong.
Limits
What it can't do
- The thinking is not free and is often wasted. On twenty easy questions here it cost 10.8 times the tokens and 10.4 times the wall clock, and changed no answer from wrong to right, because thinking off already had all twenty right.
- It can talk itself out of an answer. One question produced 15,290 characters of deliberation, hit the token cap and returned nothing. Cap the budget and log when the cap is reached.
- The thinking is not an explanation you can trust. It is generated text that was rewarded for leading to correct answers, and a model can reach the right answer through working that does not hold up.
- It does not fix a model that lacks the knowledge. More steps over a fact the model never learned produces a longer wrong answer.
- A bigger budget doesn't fix every wrong answer. In the partial sweep, a letter count stayed wrong at every budget up to 4,096, and a multiplication came back to the same wrong answer after thinking finished on its own.
- Latency makes it unusable in some products. 8.3 seconds for an arithmetic answer rules out anything a person is waiting on.
- This run is one 8B model on twenty questions, decoded greedily. It shows the cost clearly and it is nowhere near a benchmark. The gains in this article come from the DeepSeek-R1 and Qwen3 reports, on far harder problems with far larger budgets, and they haven't been reproduced here. The one local attempt to find a gain, the partial budget sweep, finished three questions and found none.
Go deeper
DeepSeek-AI (2025): DeepSeek-R1, incentivizing reasoning capability in LLMs via reinforcement learning · Wei et al. (2022): Chain-of-Thought Prompting, where writing the working down began · OpenAI (2024): the o1 System Card · Qwen3 on Hugging Face, whose models carry the thinking flag measured here · Qwen Team (2025): Qwen3 Technical Report, with the thinking budget and the 8B model's scores in both modes
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
