Skip to main content

Qwen3.8-27B on two GPUs, start here

Local LLM stackGuide

Removing unnecessary settings let the shared GPU setup run faster.
Most of the speed came from deleting settings.

This section is the cleaned-up result of about six weeks of notes, benchmarks and wrong turns. It covers two servers that share one GPU box, the clients that talk to them, and every setting that was measured. Pages are short. Read the one you need, then stop.

Scope

This section covers one tested stack: a Qwen3.8-27B GGUF on two NVIDIA GPUs, served by two llama.cpp-based servers, with large contexts and MTP speculative decoding.

The principles may transfer to other setups. The measured numbers and configuration rules belong to this tested stack unless a page says otherwise.

Choose what you are trying to do​

The stack in one picture​

O2 / The shared stack

Same model. Same GPUs.
Two different server stacks.

Mindshub Cowork

opencode · small scripts

llama-swap :8081

upstream llama.cpp
general clients

dsh

Linux + macOS

Unsloth Studio :8888

Unsloth fork
agent work

One shared model file

Qwen3.8-27B GGUF
UD-Q4_K_XL · built-in MTP head

GPU0RTX 5070 Ti

16 GB

GPU1RTX 4070

12 GB

i7-13700K · 32 GB RAM

same GGUF, same GPUs, different results. Re-measure.
LayerWhat runs there
HardwareRTX 5070 Ti 16 GB (GPU0) + RTX 4070 12 GB (GPU1), i7-13700K, 32 GB RAM
ModelQwen3.8-27B GGUF from Unsloth, mostly UD-Q4_K_XL, with built-in MTP draft head
Server Allama-swap on port 8081, upstream-based llama.cpp build, general clients
Server BUnsloth Studio on port 8888, Unsloth's llama.cpp fork, agent work
ClientsDeepseek Harness (dsh) on Linux and macOS, Mindshub Cowork, opencode, small scripts
Two servers, two binaries

The same GGUF on the same GPUs gives different results on the two servers. Context cost and MTP gain both differ. Re-measure before you copy a setting from one to the other.

The numbers that matter​

O3 / The speed journey

a bit slower,
2.7x the context

One GPU, 34 FFN layers on CPU

23.8 t/s

context 65536

Two GPUs, first working config

64–66 t/s

context 98304

llama-swap today, UD-Q4_K_XL, q4_0 KV

52.3 t/s

context 262144

Unsloth Studio q4-agent, q8_0 KV

53.8 t/s

context 131072

Thick bars: decode, 0–70 t/s. Thin bars: context, 0–262144. Decode at short prompt depth. The outlined tip marks the full 64–66 range.
SetupContextDecode speed
One GPU, 34 layers of FFN on the CPU6553623.8 t/s
Two GPUs, first working config9830464 to 66 t/s
llama-swap today, UD-Q4_K_XL, q4_0 KV26214452.3 t/s
Unsloth Studio q4-agent, q8_0 KV13107253.8 t/s

Decode speed is measured at a short prompt depth unless a page says otherwise.

Rules for this stack​

These rules survived testing on both documented servers. Treat them as configuration rules for this stack, not as universal llama.cpp defaults. Each rule has its own page with the measurements behind it.

  1. Do not set -ngl or --tensor-split on this two-GPU setup. Both switch off auto-fit.
  2. Set --fit-target 128, and re-check it every time you change -c.
  3. Judge a config by how many layers land on the CPU, not by "it loaded". Portable takeaway: a config that loads and passes /health can still crash on the first real generation, so test with a real, long prompt.
  4. KV cache type is q8_0, or q4_0 when you need the room. Not q5_1 or q5_0.
  5. Keep MTP on (--spec-type draft-mtp) and keep --parallel 1.
  6. Use the thinking sampler row: temp 1.0, top_p 0.95, top_k 20, min_p 0, presence 0, repeat 1.0.
  7. Reasoning effort is medium, not xhigh.
  8. Cap thinking with --reasoning-budget. The effort setting does not cap anything.
  9. Keep GPU1 empty. Nothing else may sit on the 4070.
  10. Gate every benchmark on HTTP 200 and a fresh output file. Portable takeaway: a failed request can leave the previous run's output in place, so a crashed config can still report a speed.

Next step: Hardware and model