Skip to main content

Open questions and sources

Local LLM stackNote

Where sources disagree​

Notes disagree
Is q4_0 KV cache safe for quality?
  • A NotebookLM summary called Q4 KV "catastrophic", with 8.3 % output similarity against 81.6 % for Q8. The same summary says that number came from a 7B coding model, and it is a text-similarity score, not a correctness score.
  • A later note puts q4_0 at about 92 % of f16 quality on code generation, and q8_0 at about 98 %.
  • llama-swap runs q4_0 at 262144 today. Unsloth Studio uses q8_0 and drops to q4_0 only to buy context.
  • YouTube commenters (unverified) say Qwen3.8 resists KV quantization damage until about 100000 tokens. One runs K at q8 and V at a "turbo4" type from a fork at 256k for coding.

My rule: q8_0 when it fits, q4_0 when you need the room. The quality cost at very long context is not measured.

q5_1 KV: worked on llama-swap, unusable on Studio

llama-swap served 163840 with q5_1 KV at 52.3 t/s. On Studio's fork the same KV type ran attention on the CPU at 13.8 t/s. The binaries differ. q5_1 was not re-tested on llama-swap after the Studio finding.

How many layers: 64 or 65?

Some notes say 64 layers. The GGUF metadata says block_count = 65. One command used -ngl 65 for "64 language layers plus the output head". The metadata is the source of truth.

Single GPU at -ngl 48

One run crashed on first decode at -ngl 48. Another reported 19.95 t/s and 15162 MiB with no crash at the same setting. The builds and contexts may differ. It does not matter now that two GPUs are in use.

reasoning_effort: none

With the stock template it raises and the server returns 500. With the froggeric template none and off turn thinking off. A harness config comment says "none disables thinking (verified)", but does not say which template was used.

Vision projector on CPU or GPU

One suggestion was to move the projector to the GPU or the 4070 for faster images. Every measured config keeps --no-mmproj-offload, because GPU1 has no free memory. Image speed on CPU against GPU was not measured.

Not tested yet​

Not tested
  • --spec-draft-p-min 0.60 to 0.75 on Qwen3.8.
  • The effect of -t on CPU image encoding.
  • --flash-attn off as a way to make q5 KV types usable.
  • q4_0 KV quality at very long context.
  • A separate non-thinking profile with the instruct sampler row.
  • Whether the 12288 thinking budget holds for write-heavy agent phases. It was measured on one diagnosis-heavy session.
  • The GPU power limit and its real speed cost.

Sources​

Upstream projects​

llama.cpp issues, pull requests and discussions​

  • #27578: soft wrap-up hint and bounded grace for the reasoning budget, by masterjaso
  • #27592: fractional --reasoning-budget, by blablanla859-wq
  • #27571: budget based on conversation length (feature request)
  • #27514: handle empty forced_tokens in the budget sampler, by arnavahire19
  • #27971: report reasoning token count in usage
  • #27342: DFlash2 speculative decoding, by SubSir
  • Discussion #21445: adjusting reasoning-budget per request
  • Discussion #20574: host-memory prompt caching tutorial

Community​

What counts as evidence here​

Measured

Numbers in this section come from my own measurements on the hardware described, unless a line says otherwise. Claims from YouTube videos, comments and AI summaries are marked unverified. Outside benchmarks are on different hardware and are not comparable to mine.

Next step: Back to Start here