Skip to main content

Reasoning effort, thinking budget and sampling

Local LLM stackNote

The symptom​

An agent on Qwen3.8 thought at length, never wrote the file, and stopped with:

I ran out of output-token budget twice in a row before completing a response, so I could not finish this step.

Three separate defects caused it. None of them was model quality.

Defect 1: the template defaults to xhigh​

R2 / Stock template

Effort is a nudge.
It is not a token cap.

DEFAULTxhigh

long ‘think carefully’ instruction

GHOST RUNGhigh

↑ silently rewritten to xhigh

NEUTRAL, NOT HALF OF XHIGHmedium

injects nothing

BRIEF INSTRUCTIONlow

‘keep thinking brief’

STOCK TEMPLATEnone

raise_exception → HTTP 500

No value limits the token count. The cap is --reasoning-budget.

Thinking off: use --reasoning-budget 0 or enable_thinking: false.

The default lives in the GGUF's own chat template, not in the server or the client:

{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
{%- if resolved_reasoning_effort == 'high' %}
{%- set resolved_reasoning_effort = 'xhigh' %}
{%- endif %}

What each value really does in the stock template:

ValueWhat the template injects
xhighA long "think carefully, validate, consider alternatives" instruction. The default.
mediumNothing. A neutral baseline, not "half of xhigh".
low"Keep your thinking brief and move to the conclusion."
highSilently rewritten to xhigh.
noneraise_exception, so the server returns 500.

The real ladder is only xhigh, medium, low. No value limits the token count.

To turn thinking off with the stock template, use --reasoning-budget 0 or enable_thinking: false, not reasoning_effort: none.

Defect 2: --reasoning-preserve replayed every thought​

This flag keeps every earlier thinking trace in the history. That is fine for chat. In a 30 to 60 step agent loop, every step replays every earlier trace. It was removed.

Defect 3: the budget equalled the client's max_tokens​

R3 / One shared output budget

Leave room for the answer.

BEFORE · max_tokens 8192

8192 thinking ceiling · answer: 0 tokens

I ran out of output-token budget twice in a row

AFTER · maxTokens 32768

12288
thinking cap

20480
available for the answer

At the cap: budget message injected.

Thinking tokens count against the same max_tokens as the answer. Each bar shows its own budget. The answer segment is available capacity, not a measured answer length.

Thinking tokens are generated tokens. They count against the same max_tokens as the answer. The client hardcoded max_tokens = 8192, and a first fix set --reasoning-budget 8192. A turn could think to the cap and leave zero tokens for the reply. The retry failed the same way, hence "twice in a row".

The fix on llama-swap​

--chat-template-kwargs '{"reasoning_effort":"medium"}'
--reasoning-budget 2048
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer."

And delete --reasoning-preserve.

Checked with a hard combinatorics prompt at max_tokens: 8192: thinking was cut at the cap in mid-sentence, the message was injected, and the full 11516-character answer came back with finish_reason: stop. Before the fix this turn returned nothing.

No \n in the budget message

llama-swap and the Studio launcher pass arguments without a shell. \n arrives in the model's context as a literal backslash and n. Use plain text.

The same bug on the second server​

The fix went into llama-swap only. Unsloth Studio, the server the agent harness actually uses, ran for two more weeks with no cap. One session died at step 57 with outputTokens exactly equal to maxTokens (16384), all of it spent thinking.

preserve_thinking: false did not help. It drops thinking from turns before the last user message, and an agent session is one long turn. In that session, earlier thinking was 63 % of the step-57 prompt.

How 12288 was chosen​

R5 / One measured session

Cut the runaway step.
Leave the ordinary ones alone.

57 blocks, 66,142 tokens. Most steps barely think. A few explode.

10010001000020000median 168mean 1,160p90 3,528max 16,384

Token axis is logarithmic. Coral line = selected cap.

1 / 57

steps clipped

6.2 %

thinking cut

chosen: cuts only the runaway step No per-step distribution is inferred.

All 57 thinking blocks of that session were tokenized with the model's own vocabulary:

57 blocks, 66,142 real tokens, 3.79 chars per token
median 168, mean 1,160, p90 3,528, max 16,384

Most steps barely think. A few explode.

CapSteps clippedThinking cut
40965 of 5731.0 %
81922 of 5712.7 %
122881 of 576.2 %
163840 of 570 %

8192 was tried first. It clipped a normal step by 206 tokens. 12288 cuts only the runaway step. With maxTokens: 32768 that leaves 20480 tokens for the answer.

Measure, do not divide by 4

Characters divided by 4 was 95 % right on the total, and five times wrong on the median. Use POST /tokenize on the raw llama-server port. The idea comes from llama.cpp issue #27971.

The rule of "budget at about one third of max_tokens" from the first fix was dropped. It was a guess. 12288 is measured, from one session. Re-measure before you lower it.

Limits of --reasoning-budget​

  • It is a startup flag. Changing it means a model reload.
  • Setting it disables the per-request thinking_budget_tokens field. See llama.cpp discussion #21445.
  • --jinja is required for any template kwarg to work.
  • The cut is hard. Softer options are in open pull requests: #27578 adds a soft wrap-up hint and a grace period, and names "thought hungry" models like Qwen3.8 27B as the reason. #27592 adds a fractional budget for #27571. #27514 fixes a crash in the budget sampler.

Low or medium: measured​

The first fix switched to low, based on one prompt. A proper test later:

Setting10 coding tasksAgent loop, 20 turns
low + AGENTS.md2,256 chars, 12.7 s, 10 of 10median 1,053, max 3,398
medium + AGENTS.md1,809 chars, 11.1 s, 10 of 10median 669, max 2,945

medium used 33 % less thinking at the same quality. The difference is not formally significant (p = 0.26 on 20 pairs), but low won nowhere. xhigh used 8.1 times as much as medium, and lost quality by running out of budget.

The instruction matters more than the dial​

A ## Thinking section in the harness's AGENTS.md cut thinking by 43 % at full quality. How you phrase it matters:

InstructionThinking chars
"Reason once and commit. Do not second-guess, reconsider, or re-verify."8,997
"Plan once, in at most five short lines, then act."1,809
Never write a bare prohibition

Naming the loop feeds it. All three invented "medium" instructions made thinking 2.5 to 2.8 times longer than injecting nothing. Pair every "do not" with a concrete budget, or leave it out.

The froggeric chat template​

Every Unsloth Qwen3.8 GGUF embeds Qwen's official template. It defaults to xhigh and crashes on the stringified-JSON tool arguments that agent harnesses send. Unsloth Studio uses the community froggeric v22.4 template instead:

  • default medium, with zero injected tokens
  • aliases: high, max and ultracode map to xhigh, minimal to low, none and off disable thinking
  • inline tags, last one wins: <|think_off|>, <|think_low|>, <|think_medium|>, <|think_xhigh|>

Rendering both templates offline and comparing bytes showed they are identical at every explicit effort level. That turned a sweep that needed model reloads into one that needed none.

--reasoning-format deepseek moves thinking into the reasoning_content field. That stops thinking tokens from leaking into a tool call and stalling the agent loop.

Sampling​

From the model card, confirmed against Unsloth's Qwen3.8 docs:

Modetemptop_ptop_kmin_ppresencerepeat
Thinking (what I run)1.00.95200.00.01.0
Instruct, thinking off0.70.80200.01.51.0

Do not mix the two rows.

Server defaults are different

llama-server's own defaults are top_k 40 and min_p 0.05. A raw curl to the server port gets those unless the server sets them. Put the sampler flags in the server config, so every client inherits them.

Two more rules from model cards of other MTP models (Defiant Fable by DavidAU): keep temperature at 1 or less, and keep repeat penalty at 1. Both break MTP otherwise.

max_tokens is not free. On Studio, a max_tokens of 90176 on a 193280 window let one answer use half the context. 32768 is plenty.

Next step: MTP speculative decoding