Benchmarking honestly
Local LLM stackNote
The traps
1. Reading an old result file
A script that runs curl -o out.json and then parses out.json reads the previous run's file when the request fails. That is how a config that crashed every time showed "60.2 t/s, reproduced twice".
Fix: check for HTTP 200 and write each run to a new file.
2. A cold page cache
The first run of each model showed prefill of 45 to 74 t/s. Warm, the same config gave 2500 to 4500 t/s. Load the model once and throw away that run.
3. "It loaded" is not "it works"
A context that is too big loads, passes /health, and then crashes on the first real generation:
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 1145.13 MiB on device 0:
cudaMalloc failed: out of memory
graph_reserve: failed to allocate compute buffers
The compute buffer grows with context. Always run a real generation, and one long prompt.
4. Short prompts
Short-prompt prefill is mostly fixed overhead (about 100 t/s here, against 835 t/s on a real 12712-token prompt). Short prompts also hide out-of-memory errors from large batch buffers.
5. Temperature 0 makes seeds useless
With top_k: 1 every seed gives the same bytes. Three "replications" at temperature 0 are one sample measured three times. Bench at the temperature you actually use. At 0.7 the same prompt gave 113, 2,031 and 1,149 thinking characters.
6. Measuring the wrong regime
A single-shot benchmark showed medium and low effort nearly equal and did not show the thinking-burn bug at all. The bug lived in the agent regime: tools defined, many tool results, deep context. Test the way you use the model.
Two shell traps
pkill -f "port 8099"matches its own shell and kills the script that runs it (exit 144). Usepgrep -x llama-server | xargs -r kill.- zsh does not split unquoted variables into words.
${fitt:+-fitt $fitt}arrives as one argument, and llama-server rejects it withinvalid argument: -fitt 128. Use an array, or run the script with bash.
A benchmark that cannot lie
B4 / Trust the current run
A failed request
must not print an old speed.
parse timings
FAILED HTTP
Beforellama-swap is running on port 8081 and the model is warm.
cd ~
f=$(mktemp --suffix=.json)
code=$(curl -s -m 900 http://127.0.0.1:8081/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user",
"content":"Explain in 120 words why copper conducts electricity."}],
"max_tokens":400}' -o "$f" -w '%{http_code}')
[ "$code" = 200 ] && python3 -c "
import json;t=json.load(open('$f'))['timings']
print('decode %.2f t/s | prefill %.2f t/s'%(t['predicted_per_second'],t['prompt_per_second']))" \
|| echo "FAILED HTTP $code"
It prints one decode ... | prefill ... line. A failure prints FAILED HTTP with the code, never an old number.
What to record for every run
Record
Rendering chat templates offline and comparing the bytes answered three test questions for free. See Reasoning and sampling.
Next step: Clients and harnesses
- Leads to
- Clients and harnesses