Skip to main content

Pitfalls and fixes

Local LLM stackRunbook

Symptom table​

SymptomLikely causeFix
Decode under 20 t/s, VRAM looks fullLayers spilled to the CPUCount CPU layers. Lower -c or set -fitt 128
About 75 t/s prefill, GPU use 0 %, CPU at 1500 %KV type is q5_1 or q5_0Use q8_0 or q4_0
Loads, passes /health, crashes on first replyContext too big for compute buffersLower -c, re-tune -fitt, test with a long prompt
Q5 model will not load at any tensor split-ngl and --tensor-split pinnedRemove both. Let auto-fit place layers
"MTP could not start" on Unsloth StudioGPU Memory set to ManualSet Auto, add --fit-target 128
failed to create MTP context on any contextThe quant stripped the MTP tensorsUse a quant that keeps MTP
Replies come back emptyText is in reasoning_content, or stock template at xhighRead reasoning_content. Set effort to medium
"Ran out of output-token budget twice"Thinking used all of max_tokensSet --reasoning-budget well below max_tokens
Server 500 on a request with reasoning_effort: noneThe stock template raises on noneUse --reasoning-budget 0 or enable_thinking: false
Fast 404 from llama-swapWrong or wrong-case model nameUse the id or an alias
500 or 502 "upstream exited prematurely"llama-server crashedRun the command by hand to see the real error
Client 502 with no server errorHost out of memory, oom-killCheck journalctl. 262144 on Studio used about 25 GiB of RAM
Speed halves when ComfyUI runsGPU1 has no margin leftcurl :8081/unload first, or raise -fitt to 256
dsh reports context overflow, or truncatesIts settings drifted from the serveructl sync, then restart the desktop app
invalid argument: -fitt 128zsh passed two words as one argumentUse an array or bash

What did not work​

TriedResultWhy
-ngl 48 on one GPULoaded at 15659 MiB, crashed on first decodeCompute buffers allocate on first use. 42 was the real ceiling
--tensor-split 18,10 or 17,11FailedThe split must match the cards, and should not be set at all
Pinned -ngl 99 and split with UD-Q5_K_XLWould not load at any of six ratiosPinning disables auto-fit
Default --fit-target (1024) at 26214410 CPU layers, 25.8 t/sMargin too large for this box
-sm row41.2 t/s against 52.3PCIe traffic costs more than one CPU layer
-sm tensorOut of memory on loadExperimental upstream
No MTP28.8 t/s against 51+MTP is worth +78 % on llama-swap
GGML_CUDA_DISABLE_GRAPHS=1Breaks the dual-GPU setupIt was a single-GPU workaround
q5_1 or q5_0 KV on Unsloth Studio13.8 t/s decode, GPU idleNo CUDA flash-attention kernel
Q6_K at 131072 with q8_0 KV27.3 t/s, silent spillWeights, KV and MTP reserve exceed free VRAM
262144 context on Unsloth StudioAbout 25 % slower, fragileThe fork's fit is marginal at full context
--reasoning-preserve in agent loopsEvery step replays every earlier thoughtFine for chat only
--reasoning-budget equal to max_tokensNo tokens left for the answerThinking counts against max_tokens
reasoning_effort: lowSame quality, more thinking than mediumThe first test used one prompt
reasoning_effort: xhigh8.1 times the thinking of mediumRuns out of budget mid-thought
Bare "do not second-guess" rules5 times more thinkingNaming the loop feeds it
Three models for three agent roles15 s reload, 30 to 60 times per taskOnly one model fits at a time
--spec-draft-p-min 0.75 (on a 9B MTP model)116.5 t/s against 144.2 at 0The gate throws away free tokens

How the config got here​

WhenConfigDecode
Start, one GPUQ4_K_M, -ngl 42, 6553622.7 t/s
One GPU, FFN of 34 layers on the CPU6553623.8 t/s
Second GPU added, -ot offload line deleted-ngl 99 --tensor-split 16,12, q8_0, 9830464 to 66 t/s
16 AugUD-Q5_K_XL, auto-fit, 98304, q8_053.9 to 56.9 t/s
16 AugEffort medium, budget 2048, preserve removedThinking fixed on llama-swap
18 Aug163840 with q5_1 KV52.3 t/s
19 AugUD-Q4_K_XL, 262144, q4_0, -fitt 12852.3 t/s
27 AugStudio: GPU Memory Auto, --fit-target 12817 to 52.5 t/s
28 AugStudio: Q6_K knee found at 9830448.6 t/s
31 AugStudio q4-agent, budget 12288, effort medium53.8 t/s

The biggest single win was deleting one line: the -ot rule that kept feed-forward weights on the CPU. On one 16 GB card it was needed. With 28 GB it only cost speed.

Removing unnecessary flags produced most speedups. Fit-target 128 was the useful addition.
Removed: -ot, -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve.
Kept: -fitt 128.
When in doubt, remove a flag

Almost every speedup in this table came from removing a setting or using a default: the offload line, -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve. The one flag that had to be added was -fitt 128.

Next step: Open questions and sources