Skip to main content

llama-swap setup

Local LLM stackNote

llama-swap is a small proxy. It starts the right llama-server for the model name in each request and stops the previous one. Every client gets an OpenAI-compatible API on one port.

What it serves​

S1 / Route, unload, load

Many aliases.
One resident model.

llama-swap :8081

main + coder → qwen3.8-27b
uncensored → qwen3.8-27b-uncensored

qwen3.8-27b

unloaded

qwen3.8-27b-uncensored

resident

Shared VRAM pool

The uncensored server answers. The original server is unloaded.

Qwen3.8-27b
wrong or wrong-case name → fast 404 (~13 ms)

One model resident at a time. Aliases keep old configs working.
Model idAliasesUse
qwen3.8-27bmain default coder fast-coder vision-chat qwen3.8 qwen38Everything by default
qwen3.8-27b-uncensoreduncensored heavy deep heavy-uncensored rvnSecurity and red-team work

Aliases let old client configs keep working after a model swap. Model ids are case-sensitive, so Qwen3.8-27B only works because it is in the alias list.

The config​

~/.config/llama-swap/config.yaml. The rationale lives in a separate TUNING.md, so the config stays short. It went from 161 lines to 44.

macros:
common: >-
--host 127.0.0.1 -fa 1 --jinja --metrics --no-warmup --parallel 1
--mmproj ~/models/mmproj-Qwen3.8-27B-F16.gguf
--no-mmproj-offload --image-min-tokens 1024
--chat-template-kwargs '{"reasoning_effort":"medium"}'
--reasoning-budget 2048
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer."
-t 8 -tb 16
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0
--presence-penalty 0 --repeat-penalty 1.0

models:
"qwen3.8-27b":
cmd: >-
${bin} --port ${PORT} ${common}
-m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf
-c 262144 --cache-type-k q4_0 --cache-type-v q4_0 -fitt 128
--spec-type draft-mtp --spec-draft-n-max 3
aliases: [main, default, coder, fast-coder, vision-chat, qwen3.8, qwen38]

The uncensored entry is the same with its own weights and --spec-draft-n-max 2. The config is a sketch of the real file. Paths are shortened.

Every flag, in plain words​

FlagWhy it is there
--host 127.0.0.1Each llama-server listens on localhost only. Clients talk to llama-swap, never to the backend.
-fa 1Flash attention. Needed for quantized KV on CUDA.
--jinjaUse the chat template. Required for the reasoning kwargs to work at all.
--parallel 1One slot. Required with MTP.
--no-warmupSkips the warmup run, so a model swap finishes sooner.
--mmproj, --no-mmproj-offloadVision on, but the projector stays on the CPU and costs no VRAM until an image arrives.
--image-min-tokens 1024The loader warns that Qwen-VL needs this for accurate grounding.
--chat-template-kwargsSets reasoning effort to medium. The GGUF default is xhigh.
--reasoning-budget 2048Hard cap on thinking tokens. See Reasoning and sampling.
-t 8 -tb 16Threads. Speed does not change, see below.
Sampler flagsThe model card's thinking-mode row.
-c 262144Full native context.
q4_0 KVBuys the room for 262144.
-fitt 128Reserve 128 MiB per GPU instead of 1024. This one flag unlocked full context.
--spec-type draft-mtpMTP speculative decoding. About +78 % here.

Not set on purpose: -ngl, --tensor-split, -sm, GGML_CUDA_DISABLE_GRAPHS, --reasoning-preserve. Each one made things worse. See Pitfalls and fixes.

Threads do not matter here​

Measured at 262144 context on a 12712-token prompt:

ThreadsPrefillDecode
-t 8 -tb 24835 t/s51.6 t/s
-t 8 -tb 16833 t/s50.8 t/s
-t 6 -tb 16834 t/s50.7 t/s

All within noise. The work is on the GPUs, and the CPU threads mostly wait at the CUDA barrier. -tb 16 keeps the E-cores out of that waiting pool at no cost. -t 8 stays because image encoding runs on the CPU, and that path was not benchmarked.

This matches what others report only partly. A LocalLLaMA thread found +80 % from tuning threads, and another found P-cores slower than E-cores with MoE offload. Both involve CPU offload. With everything on the GPU, threads stop mattering.

Why the CPU shows 3 to 40 %

Eight spinning decode threads out of 24 is about 33 % of the machine. That is the peak you see. 40 to 60 °C is idle to light load for this chip.

Check it after a change​

BeforeYou edited config.yaml. With -watch-config, llama-swap reloads it without a restart.

Steps

cd ~
systemctl --user restart llama-swap.service

curl -s http://127.0.0.1:8081/v1/models \
| python3 -c "import json,sys; print([m['id'] for m in json.load(sys.stdin)['data']])"

curl -s http://127.0.0.1:8081/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"main","messages":[{"role":"user","content":"Reply with exactly: CONFIG OK"}],"max_tokens":2000}'

curl -s http://127.0.0.1:8081/upstream/qwen3.8-27b/props \
| python3 -c "import json,sys; print(json.load(sys.stdin)['default_generation_settings']['n_ctx'])"
Done when

The model list shows both ids, the reply contains CONFIG OK, and n_ctx prints 262144.

Error messages that mislead​

  • A fast 404 (about 13 ms) means a wrong model name. A slow failure is a real load problem.
  • llama-swap reports crashes as a generic 500 or 502 "upstream exited prematurely". The real error is only in llama-server's own output. Run the command by hand to see it.
  • An "empty" reply is often in reasoning_content, not content.

Before deleting a model​

Search every client config for the model name first. An old model was the default in an opencode project and held the heavy and deep aliases. Deleting the GGUF alone would have broken that project without an error.

cd ~
grep -rn "old-model-name" ~/.config/opencode ~/*/opencode.json 2>/dev/null

Next step: Context and fit-target