Qwen3.8-27B on two GPUs, start here
Local LLM stackGuide
This section is the cleaned-up result of about six weeks of notes, benchmarks and wrong turns. It covers two servers that share one GPU box, the clients that talk to them, and every setting that was measured. Pages are short. Read the one you need, then stop.
This section covers one tested stack: a Qwen3.8-27B GGUF on two NVIDIA GPUs, served by two llama.cpp-based servers, with large contexts and MTP speculative decoding.
The principles may transfer to other setups. The measured numbers and configuration rules belong to this tested stack unless a page says otherwise.
Choose what you are trying to do
- I want to change a setting. Read Context and fit-target first. It explains the one metric that tells you if a change helped. Then read the page for the server you are changing: llama-swap or Unsloth Studio.
- Something broke. Go to Pitfalls and fixes. It has a symptom table.
- The model thinks forever or returns nothing. Go to Reasoning and sampling.
- I want to know why it is this fast. Read Hardware and model, then MTP speculative decoding.
The stack in one picture
O2 / The shared stack
Same model. Same GPUs.
Two different server stacks.
opencode · small scripts
:8081upstream llama.cpp
general clients
Linux + macOS
:8888Unsloth fork
agent work
Qwen3.8-27B GGUF
UD-Q4_K_XL · built-in MTP head
16 GB
12 GB
i7-13700K · 32 GB RAM
| Layer | What runs there |
|---|---|
| Hardware | RTX 5070 Ti 16 GB (GPU0) + RTX 4070 12 GB (GPU1), i7-13700K, 32 GB RAM |
| Model | Qwen3.8-27B GGUF from Unsloth, mostly UD-Q4_K_XL, with built-in MTP draft head |
| Server A | llama-swap on port 8081, upstream-based llama.cpp build, general clients |
| Server B | Unsloth Studio on port 8888, Unsloth's llama.cpp fork, agent work |
| Clients | Deepseek Harness (dsh) on Linux and macOS, Mindshub Cowork, opencode, small scripts |
The same GGUF on the same GPUs gives different results on the two servers. Context cost and MTP gain both differ. Re-measure before you copy a setting from one to the other.
The numbers that matter
O3 / The speed journey
a bit slower,
2.7x the context
One GPU, 34 FFN layers on CPU
23.8 t/s
context 65536
Two GPUs, first working config
64–66 t/s
context 98304
llama-swap today, UD-Q4_K_XL, q4_0 KV
52.3 t/s
context 262144
Unsloth Studio q4-agent, q8_0 KV
53.8 t/s
context 131072
| Setup | Context | Decode speed |
|---|---|---|
| One GPU, 34 layers of FFN on the CPU | 65536 | 23.8 t/s |
| Two GPUs, first working config | 98304 | 64 to 66 t/s |
llama-swap today, UD-Q4_K_XL, q4_0 KV | 262144 | 52.3 t/s |
Unsloth Studio q4-agent, q8_0 KV | 131072 | 53.8 t/s |
Decode speed is measured at a short prompt depth unless a page says otherwise.
Rules for this stack
These rules survived testing on both documented servers. Treat them as configuration rules for this stack, not as universal llama.cpp defaults. Each rule has its own page with the measurements behind it.
- Do not set
-nglor--tensor-spliton this two-GPU setup. Both switch off auto-fit. - Set
--fit-target 128, and re-check it every time you change-c. - Judge a config by how many layers land on the CPU, not by "it loaded". Portable takeaway: a config that loads and passes
/healthcan still crash on the first real generation, so test with a real, long prompt. - KV cache type is
q8_0, orq4_0when you need the room. Notq5_1orq5_0. - Keep MTP on (
--spec-type draft-mtp) and keep--parallel 1. - Use the thinking sampler row: temp 1.0, top_p 0.95, top_k 20, min_p 0, presence 0, repeat 1.0.
- Reasoning effort is
medium, notxhigh. - Cap thinking with
--reasoning-budget. The effort setting does not cap anything. - Keep GPU1 empty. Nothing else may sit on the 4070.
- Gate every benchmark on HTTP 200 and a fresh output file. Portable takeaway: a failed request can leave the previous run's output in place, so a crashed config can still report a speed.
Next step: Hardware and model