Benchmarking honestly
Six ways my benchmarks produced numbers that never happened, and a small script that avoids them.
Six ways my benchmarks produced numbers that never happened, and a small script that avoids them.
Which tools talk to which server, and the client-side settings that must match the server.
The one flag that unlocked full context, the one number that tells you if a config is good, and why pinning layers breaks dual-GPU setups.
The dual-GPU box, why a 27B model fits with a huge context, and which quant to pick.
The llama-swap config that serves Qwen3.8-27B at full context, flag by flag, and how to check it after a change.
The built-in draft head that gives up to +78 percent decode speed, how to tune it, and the settings that looked right but were not.
Where my own notes disagree, what is still untested, and every outside source this section draws on.
A local 27B cyber model, one plain prompt, two malicious Windows binaries fully pulled apart — AES config, C2, IOCs and all.
An uncensored Qwen3.8-27B twin, and the external DFlash2 drafter that reaches 99 t/s, once you serve it on the right binary.
Everything that did not work, a symptom table for when something breaks, and the history of configs that led here.
How my local Qwen3.8-27B stack is set up, what the numbers are, and which page to read next.
Why Qwen3.8 thought until it ran out of tokens, how the effort setting really works, and the sampler values to use.
The second server, its presets, every Run settings field, and how the agent harness stays in step with it.