llama-swap setup
Local LLM stackNote
llama-swap is a small proxy. It starts the right llama-server for the model name in each request and stops the previous one. Every client gets an OpenAI-compatible API on one port.
What it serves
S1 / Route, unload, load
Many aliases.
One resident model.
:8081main + coder → qwen3.8-27b
uncensored → qwen3.8-27b-uncensored
qwen3.8-27bunloaded
qwen3.8-27b-uncensoredresident
Shared VRAM pool
The uncensored server answers. The original server is unloaded.
Qwen3.8-27b
wrong or wrong-case name → fast 404 (~13 ms)
| Model id | Aliases | Use |
|---|---|---|
qwen3.8-27b | main default coder fast-coder vision-chat qwen3.8 qwen38 | Everything by default |
qwen3.8-27b-uncensored | uncensored heavy deep heavy-uncensored rvn | Security and red-team work |
Aliases let old client configs keep working after a model swap. Model ids are case-sensitive, so Qwen3.8-27B only works because it is in the alias list.
The config
~/.config/llama-swap/config.yaml. The rationale lives in a separate TUNING.md, so the config stays short. It went from 161 lines to 44.
macros:
common: >-
--host 127.0.0.1 -fa 1 --jinja --metrics --no-warmup --parallel 1
--mmproj ~/models/mmproj-Qwen3.8-27B-F16.gguf
--no-mmproj-offload --image-min-tokens 1024
--chat-template-kwargs '{"reasoning_effort":"medium"}'
--reasoning-budget 2048
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer."
-t 8 -tb 16
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0
--presence-penalty 0 --repeat-penalty 1.0
models:
"qwen3.8-27b":
cmd: >-
${bin} --port ${PORT} ${common}
-m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf
-c 262144 --cache-type-k q4_0 --cache-type-v q4_0 -fitt 128
--spec-type draft-mtp --spec-draft-n-max 3
aliases: [main, default, coder, fast-coder, vision-chat, qwen3.8, qwen38]
The uncensored entry is the same with its own weights and --spec-draft-n-max 2. The config is a sketch of the real file. Paths are shortened.
Every flag, in plain words
| Flag | Why it is there |
|---|---|
--host 127.0.0.1 | Each llama-server listens on localhost only. Clients talk to llama-swap, never to the backend. |
-fa 1 | Flash attention. Needed for quantized KV on CUDA. |
--jinja | Use the chat template. Required for the reasoning kwargs to work at all. |
--parallel 1 | One slot. Required with MTP. |
--no-warmup | Skips the warmup run, so a model swap finishes sooner. |
--mmproj, --no-mmproj-offload | Vision on, but the projector stays on the CPU and costs no VRAM until an image arrives. |
--image-min-tokens 1024 | The loader warns that Qwen-VL needs this for accurate grounding. |
--chat-template-kwargs | Sets reasoning effort to medium. The GGUF default is xhigh. |
--reasoning-budget 2048 | Hard cap on thinking tokens. See Reasoning and sampling. |
-t 8 -tb 16 | Threads. Speed does not change, see below. |
| Sampler flags | The model card's thinking-mode row. |
-c 262144 | Full native context. |
q4_0 KV | Buys the room for 262144. |
-fitt 128 | Reserve 128 MiB per GPU instead of 1024. This one flag unlocked full context. |
--spec-type draft-mtp | MTP speculative decoding. About +78 % here. |
Not set on purpose: -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve. Each one made things worse. See Pitfalls and fixes.
Threads do not matter here
Measured at 262144 context on a 12712-token prompt:
| Threads | Prefill | Decode |
|---|---|---|
-t 8 -tb 24 | 835 t/s | 51.6 t/s |
-t 8 -tb 16 | 833 t/s | 50.8 t/s |
-t 6 -tb 16 | 834 t/s | 50.7 t/s |
All within noise. The work is on the GPUs, and the CPU threads mostly wait at the CUDA barrier. -tb 16 keeps the E-cores out of that waiting pool at no cost. -t 8 stays because image encoding runs on the CPU, and that path was not benchmarked.
This matches what others report only partly. A LocalLLaMA thread found +80 % from tuning threads, and another found P-cores slower than E-cores with MoE offload. Both involve CPU offload. With everything on the GPU, threads stop mattering.
Eight spinning decode threads out of 24 is about 33 % of the machine. That is the peak you see. 40 to 60 °C is idle to light load for this chip.
Check it after a change
BeforeYou edited config.yaml. With -watch-config, llama-swap reloads it without a restart.
Steps
cd ~
systemctl --user restart llama-swap.service
curl -s http://127.0.0.1:8081/v1/models \
| python3 -c "import json,sys; print([m['id'] for m in json.load(sys.stdin)['data']])"
curl -s http://127.0.0.1:8081/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"main","messages":[{"role":"user","content":"Reply with exactly: CONFIG OK"}],"max_tokens":2000}'
curl -s http://127.0.0.1:8081/upstream/qwen3.8-27b/props \
| python3 -c "import json,sys; print(json.load(sys.stdin)['default_generation_settings']['n_ctx'])"
The model list shows both ids, the reply contains CONFIG OK, and n_ctx prints 262144.
Error messages that mislead
- A fast 404 (about 13 ms) means a wrong model name. A slow failure is a real load problem.
- llama-swap reports crashes as a generic
500or502"upstream exited prematurely". The real error is only in llama-server's own output. Run the command by hand to see it. - An "empty" reply is often in
reasoning_content, notcontent.
Before deleting a model
Search every client config for the model name first. An old model was the default in an opencode project and held the heavy and deep aliases. Deleting the GGUF alone would have broken that project without an error.
cd ~
grep -rn "old-model-name" ~/.config/opencode ~/*/opencode.json 2>/dev/null
Next step: Context and fit-target