Reasoning effort, thinking budget and sampling
Local LLM stackNote
The symptom
An agent on Qwen3.8 thought at length, never wrote the file, and stopped with:
I ran out of output-token budget twice in a row before completing a response, so I could not finish this step.
Three separate defects caused it. None of them was model quality.
Defect 1: the template defaults to xhigh
R2 / Stock template
Effort is a nudge.
It is not a token cap.
long ‘think carefully’ instruction
↑ silently rewritten to xhigh
injects nothing
‘keep thinking brief’
raise_exception → HTTP 500
--reasoning-budget.Thinking off: use --reasoning-budget 0 or enable_thinking: false.
The default lives in the GGUF's own chat template, not in the server or the client:
{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
{%- if resolved_reasoning_effort == 'high' %}
{%- set resolved_reasoning_effort = 'xhigh' %}
{%- endif %}
What each value really does in the stock template:
| Value | What the template injects |
|---|---|
xhigh | A long "think carefully, validate, consider alternatives" instruction. The default. |
medium | Nothing. A neutral baseline, not "half of xhigh". |
low | "Keep your thinking brief and move to the conclusion." |
high | Silently rewritten to xhigh. |
none | raise_exception, so the server returns 500. |
The real ladder is only xhigh, medium, low. No value limits the token count.
To turn thinking off with the stock template, use --reasoning-budget 0 or enable_thinking: false, not reasoning_effort: none.
Defect 2: --reasoning-preserve replayed every thought
This flag keeps every earlier thinking trace in the history. That is fine for chat. In a 30 to 60 step agent loop, every step replays every earlier trace. It was removed.
Defect 3: the budget equalled the client's max_tokens
R3 / One shared output budget
Leave room for the answer.
BEFORE · max_tokens 8192
8192 thinking ceiling · answer: 0 tokens
I ran out of output-token budget twice in a row
AFTER · maxTokens 32768
12288
thinking cap
20480
available for the answer
At the cap: budget message injected.
max_tokens as the answer. Each bar shows its own budget. The answer segment is available capacity, not a measured answer length.Thinking tokens are generated tokens. They count against the same max_tokens as the answer. The client hardcoded max_tokens = 8192, and a first fix set --reasoning-budget 8192. A turn could think to the cap and leave zero tokens for the reply. The retry failed the same way, hence "twice in a row".
The fix on llama-swap
--chat-template-kwargs '{"reasoning_effort":"medium"}'
--reasoning-budget 2048
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer."
And delete --reasoning-preserve.
Checked with a hard combinatorics prompt at max_tokens: 8192: thinking was cut at the cap in mid-sentence, the message was injected, and the full 11516-character answer came back with finish_reason: stop. Before the fix this turn returned nothing.
\n in the budget messagellama-swap and the Studio launcher pass arguments without a shell. \n arrives in the model's context as a literal backslash and n. Use plain text.
The same bug on the second server
The fix went into llama-swap only. Unsloth Studio, the server the agent harness actually uses, ran for two more weeks with no cap. One session died at step 57 with outputTokens exactly equal to maxTokens (16384), all of it spent thinking.
preserve_thinking: false did not help. It drops thinking from turns before the last user message, and an agent session is one long turn. In that session, earlier thinking was 63 % of the step-57 prompt.
How 12288 was chosen
R5 / One measured session
Cut the runaway step.
Leave the ordinary ones alone.
57 blocks, 66,142 tokens. Most steps barely think. A few explode.
Token axis is logarithmic. Coral line = selected cap.
1 / 57
steps clipped
6.2 %
thinking cut
All 57 thinking blocks of that session were tokenized with the model's own vocabulary:
57 blocks, 66,142 real tokens, 3.79 chars per token
median 168, mean 1,160, p90 3,528, max 16,384
Most steps barely think. A few explode.
| Cap | Steps clipped | Thinking cut |
|---|---|---|
| 4096 | 5 of 57 | 31.0 % |
| 8192 | 2 of 57 | 12.7 % |
| 12288 | 1 of 57 | 6.2 % |
| 16384 | 0 of 57 | 0 % |
8192 was tried first. It clipped a normal step by 206 tokens. 12288 cuts only the runaway step. With maxTokens: 32768 that leaves 20480 tokens for the answer.
Characters divided by 4 was 95 % right on the total, and five times wrong on the median. Use POST /tokenize on the raw llama-server port. The idea comes from llama.cpp issue #27971.
The rule of "budget at about one third of max_tokens" from the first fix was dropped. It was a guess. 12288 is measured, from one session. Re-measure before you lower it.
Limits of --reasoning-budget
- It is a startup flag. Changing it means a model reload.
- Setting it disables the per-request
thinking_budget_tokensfield. See llama.cpp discussion #21445. --jinjais required for any template kwarg to work.- The cut is hard. Softer options are in open pull requests: #27578 adds a soft wrap-up hint and a grace period, and names "thought hungry" models like Qwen3.8 27B as the reason. #27592 adds a fractional budget for #27571. #27514 fixes a crash in the budget sampler.
Low or medium: measured
The first fix switched to low, based on one prompt. A proper test later:
| Setting | 10 coding tasks | Agent loop, 20 turns |
|---|---|---|
low + AGENTS.md | 2,256 chars, 12.7 s, 10 of 10 | median 1,053, max 3,398 |
medium + AGENTS.md | 1,809 chars, 11.1 s, 10 of 10 | median 669, max 2,945 |
medium used 33 % less thinking at the same quality. The difference is not formally significant (p = 0.26 on 20 pairs), but low won nowhere. xhigh used 8.1 times as much as medium, and lost quality by running out of budget.
The instruction matters more than the dial
A ## Thinking section in the harness's AGENTS.md cut thinking by 43 % at full quality. How you phrase it matters:
| Instruction | Thinking chars |
|---|---|
| "Reason once and commit. Do not second-guess, reconsider, or re-verify." | 8,997 |
| "Plan once, in at most five short lines, then act." | 1,809 |
Naming the loop feeds it. All three invented "medium" instructions made thinking 2.5 to 2.8 times longer than injecting nothing. Pair every "do not" with a concrete budget, or leave it out.
The froggeric chat template
Every Unsloth Qwen3.8 GGUF embeds Qwen's official template. It defaults to xhigh and crashes on the stringified-JSON tool arguments that agent harnesses send. Unsloth Studio uses the community froggeric v22.4 template instead:
- default
medium, with zero injected tokens - aliases:
high,maxandultracodemap toxhigh,minimaltolow,noneandoffdisable thinking - inline tags, last one wins:
<|think_off|>,<|think_low|>,<|think_medium|>,<|think_xhigh|>
Rendering both templates offline and comparing bytes showed they are identical at every explicit effort level. That turned a sweep that needed model reloads into one that needed none.
--reasoning-format deepseek moves thinking into the reasoning_content field. That stops thinking tokens from leaking into a tool call and stalling the agent loop.
Sampling
From the model card, confirmed against Unsloth's Qwen3.8 docs:
| Mode | temp | top_p | top_k | min_p | presence | repeat |
|---|---|---|---|---|---|---|
| Thinking (what I run) | 1.0 | 0.95 | 20 | 0.0 | 0.0 | 1.0 |
| Instruct, thinking off | 0.7 | 0.80 | 20 | 0.0 | 1.5 | 1.0 |
Do not mix the two rows.
llama-server's own defaults are top_k 40 and min_p 0.05. A raw curl to the server port gets those unless the server sets them. Put the sampler flags in the server config, so every client inherits them.
Two more rules from model cards of other MTP models (Defiant Fable by DavidAU): keep temperature at 1 or less, and keep repeat penalty at 1. Both break MTP otherwise.
max_tokens is not free. On Studio, a max_tokens of 90176 on a 193280 window let one answer use half the context. 32768 is plenty.
Next step: MTP speculative decoding
- Leads to
- MTP speculative decoding