Pitfalls and fixes
Local LLM stackRunbook
Symptom table
| Symptom | Likely cause | Fix |
|---|---|---|
| Decode under 20 t/s, VRAM looks full | Layers spilled to the CPU | Count CPU layers. Lower -c or set -fitt 128 |
| About 75 t/s prefill, GPU use 0 %, CPU at 1500 % | KV type is q5_1 or q5_0 | Use q8_0 or q4_0 |
Loads, passes /health, crashes on first reply | Context too big for compute buffers | Lower -c, re-tune -fitt, test with a long prompt |
| Q5 model will not load at any tensor split | -ngl and --tensor-split pinned | Remove both. Let auto-fit place layers |
| "MTP could not start" on Unsloth Studio | GPU Memory set to Manual | Set Auto, add --fit-target 128 |
failed to create MTP context on any context | The quant stripped the MTP tensors | Use a quant that keeps MTP |
| Replies come back empty | Text is in reasoning_content, or stock template at xhigh | Read reasoning_content. Set effort to medium |
| "Ran out of output-token budget twice" | Thinking used all of max_tokens | Set --reasoning-budget well below max_tokens |
Server 500 on a request with reasoning_effort: none | The stock template raises on none | Use --reasoning-budget 0 or enable_thinking: false |
| Fast 404 from llama-swap | Wrong or wrong-case model name | Use the id or an alias |
500 or 502 "upstream exited prematurely" | llama-server crashed | Run the command by hand to see the real error |
Client 502 with no server error | Host out of memory, oom-kill | Check journalctl. 262144 on Studio used about 25 GiB of RAM |
| Speed halves when ComfyUI runs | GPU1 has no margin left | curl :8081/unload first, or raise -fitt to 256 |
| dsh reports context overflow, or truncates | Its settings drifted from the server | uctl sync, then restart the desktop app |
invalid argument: -fitt 128 | zsh passed two words as one argument | Use an array or bash |
What did not work
| Tried | Result | Why |
|---|---|---|
-ngl 48 on one GPU | Loaded at 15659 MiB, crashed on first decode | Compute buffers allocate on first use. 42 was the real ceiling |
--tensor-split 18,10 or 17,11 | Failed | The split must match the cards, and should not be set at all |
Pinned -ngl 99 and split with UD-Q5_K_XL | Would not load at any of six ratios | Pinning disables auto-fit |
Default --fit-target (1024) at 262144 | 10 CPU layers, 25.8 t/s | Margin too large for this box |
-sm row | 41.2 t/s against 52.3 | PCIe traffic costs more than one CPU layer |
-sm tensor | Out of memory on load | Experimental upstream |
| No MTP | 28.8 t/s against 51+ | MTP is worth +78 % on llama-swap |
GGML_CUDA_DISABLE_GRAPHS=1 | Breaks the dual-GPU setup | It was a single-GPU workaround |
| q5_1 or q5_0 KV on Unsloth Studio | 13.8 t/s decode, GPU idle | No CUDA flash-attention kernel |
| Q6_K at 131072 with q8_0 KV | 27.3 t/s, silent spill | Weights, KV and MTP reserve exceed free VRAM |
| 262144 context on Unsloth Studio | About 25 % slower, fragile | The fork's fit is marginal at full context |
--reasoning-preserve in agent loops | Every step replays every earlier thought | Fine for chat only |
--reasoning-budget equal to max_tokens | No tokens left for the answer | Thinking counts against max_tokens |
reasoning_effort: low | Same quality, more thinking than medium | The first test used one prompt |
reasoning_effort: xhigh | 8.1 times the thinking of medium | Runs out of budget mid-thought |
| Bare "do not second-guess" rules | 5 times more thinking | Naming the loop feeds it |
| Three models for three agent roles | 15 s reload, 30 to 60 times per task | Only one model fits at a time |
--spec-draft-p-min 0.75 (on a 9B MTP model) | 116.5 t/s against 144.2 at 0 | The gate throws away free tokens |
How the config got here
| When | Config | Decode |
|---|---|---|
| Start, one GPU | Q4_K_M, -ngl 42, 65536 | 22.7 t/s |
| One GPU, FFN of 34 layers on the CPU | 65536 | 23.8 t/s |
Second GPU added, -ot offload line deleted | -ngl 99 --tensor-split 16,12, q8_0, 98304 | 64 to 66 t/s |
| 16 Aug | UD-Q5_K_XL, auto-fit, 98304, q8_0 | 53.9 to 56.9 t/s |
| 16 Aug | Effort medium, budget 2048, preserve removed | Thinking fixed on llama-swap |
| 18 Aug | 163840 with q5_1 KV | 52.3 t/s |
| 19 Aug | UD-Q4_K_XL, 262144, q4_0, -fitt 128 | 52.3 t/s |
| 27 Aug | Studio: GPU Memory Auto, --fit-target 128 | 17 to 52.5 t/s |
| 28 Aug | Studio: Q6_K knee found at 98304 | 48.6 t/s |
| 31 Aug | Studio q4-agent, budget 12288, effort medium | 53.8 t/s |
The biggest single win was deleting one line: the -ot rule that kept feed-forward weights on the CPU. On one 16 GB card it was needed. With 28 GB it only cost speed.

-ot, -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve.Kept:
-fitt 128.When in doubt, remove a flag
Almost every speedup in this table came from removing a setting or using a default: the offload line, -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve. The one flag that had to be added was -fitt 128.
Next step: Open questions and sources
- Leads to
- Open questions and sources