Local LLM stack
Qwen3.8-27B on two GPUs with llama.cpp, llama-swap and Unsloth Studio. What worked, what did not.
12 entries- Benchmarking honestlySix ways my benchmarks produced numbers that never happened, and a small script that avoids them.
- Clients and harnessesWhich tools talk to which server, and the client-side settings that must match the server.
- Context, fit-target and CPU layersThe one flag that unlocked full context, the one number that tells you if a config is good, and why pinning layers breaks dual-GPU setups.
- Hardware and modelThe dual-GPU box, why a 27B model fits with a huge context, and which quant to pick.
- llama-swap setupThe llama-swap config that serves Qwen3.8-27B at full context, flag by flag, and how to check it after a change.
- MTP speculative decodingThe built-in draft head that gives up to +78 percent decode speed, how to tune it, and the settings that looked right but were not.
- Open questions and sourcesWhere my own notes disagree, what is still untested, and every outside source this section draws on.
- OrcaSAQ-2 Cyber and DFlash2An uncensored Qwen3.8-27B twin, and the external DFlash2 drafter that reaches 99 t/s, once you serve it on the right binary.
- Pitfalls and fixesEverything that did not work, a symptom table for when something breaks, and the history of configs that led here.
- Reasoning effort, thinking budget and samplingWhy Qwen3.8 thought until it ran out of tokens, how the effort setting really works, and the sampler values to use.
- Qwen3.8-27B on two GPUs, start hereHow my local Qwen3.8-27B stack is set up, what the numbers are, and which page to read next.
- Unsloth Studio and the agent harnessThe second server, its presets, every Run settings field, and how the agent harness stays in step with it.