Skip to main content

Benchmarking honestly

Local LLM stackNote

The traps​

1. Reading an old result file​

A script that runs curl -o out.json and then parses out.json reads the previous run's file when the request fails. That is how a config that crashed every time showed "60.2 t/s, reproduced twice".

Fix: check for HTTP 200 and write each run to a new file.

2. A cold page cache​

The first run of each model showed prefill of 45 to 74 t/s. Warm, the same config gave 2500 to 4500 t/s. Load the model once and throw away that run.

3. "It loaded" is not "it works"​

A context that is too big loads, passes /health, and then crashes on the first real generation:

ggml_backend_cuda_buffer_type_alloc_buffer: allocating 1145.13 MiB on device 0:
cudaMalloc failed: out of memory
graph_reserve: failed to allocate compute buffers

The compute buffer grows with context. Always run a real generation, and one long prompt.

4. Short prompts​

Short-prompt prefill is mostly fixed overhead (about 100 t/s here, against 835 t/s on a real 12712-token prompt). Short prompts also hide out-of-memory errors from large batch buffers.

5. Temperature 0 makes seeds useless​

With top_k: 1 every seed gives the same bytes. Three "replications" at temperature 0 are one sample measured three times. Bench at the temperature you actually use. At 0.7 the same prompt gave 113, 2,031 and 1,149 thinking characters.

6. Measuring the wrong regime​

A single-shot benchmark showed medium and low effort nearly equal and did not show the thinking-burn bug at all. The bug lived in the agent regime: tools defined, many tool results, deep context. Test the way you use the model.

Two shell traps​

  • pkill -f "port 8099" matches its own shell and kills the script that runs it (exit 144). Use pgrep -x llama-server | xargs -r kill.
  • zsh does not split unquoted variables into words. ${fitt:+-fitt $fitt} arrives as one argument, and llama-server rejects it with invalid argument: -fitt 128. Use an array, or run the script with bash.

A benchmark that cannot lie​

B4 / Trust the current run

A failed request
must not print an old speed.

new temp file
curl with -w %{http_code}
HTTP == 200?
Yes

parse timings

No

FAILED HTTP

Fresh file + HTTP 200. Both gates must pass before a number becomes a result.

Beforellama-swap is running on port 8081 and the model is warm.

cd ~
f=$(mktemp --suffix=.json)
code=$(curl -s -m 900 http://127.0.0.1:8081/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user",
"content":"Explain in 120 words why copper conducts electricity."}],
"max_tokens":400}' -o "$f" -w '%{http_code}')
[ "$code" = 200 ] && python3 -c "
import json;t=json.load(open('$f'))['timings']
print('decode %.2f t/s | prefill %.2f t/s'%(t['predicted_per_second'],t['prompt_per_second']))" \
|| echo "FAILED HTTP $code"
Done when

It prints one decode ... | prefill ... line. A failure prints FAILED HTTP with the code, never an old number.

What to record for every run​

Record

Diff templates before you spend GPU time

Rendering chat templates offline and comparing the bytes answered three test questions for free. See Reasoning and sampling.

Next step: Clients and harnesses