Skip to main content

13 docs tagged with "Local LLM"

View all tags

Benchmarking honestly

Six ways my benchmarks produced numbers that never happened, and a small script that avoids them.

Clients and harnesses

Which tools talk to which server, and the client-side settings that must match the server.

Context, fit-target and CPU layers

The one flag that unlocked full context, the one number that tells you if a config is good, and why pinning layers breaks dual-GPU setups.

Hardware and model

The dual-GPU box, why a 27B model fits with a huge context, and which quant to pick.

llama-swap setup

The llama-swap config that serves Qwen3.8-27B at full context, flag by flag, and how to check it after a change.

MTP speculative decoding

The built-in draft head that gives up to +78 percent decode speed, how to tune it, and the settings that looked right but were not.

OrcaSAQ-2 Cyber and DFlash2

An uncensored Qwen3.8-27B twin, and the external DFlash2 drafter that reaches 99 t/s, once you serve it on the right binary.

Pitfalls and fixes

Everything that did not work, a symptom table for when something breaks, and the history of configs that led here.