Skip to main content

Hardware and model

Local LLM stackNote

The box​

PartDetail
GPU0RTX 5070 Ti, 16 GB (16303 MiB usable), top slot, PCIe 5.0 x16
GPU1RTX 4070, 12 GB, bottom slot, PCIe 3.0 running at x4
BoardASUS ROG STRIX B660-F GAMING WIFI
CPUi7-13700K, 8 P-cores + 8 E-cores, 24 threads
RAM32 GB
PSU800 W

The x4 slot is fine for inference​

A narrow PCIe link slows the one-time weight transfer but has room for small per-token activations.
Loading: 19 GB weights, about 90 s. Generating: 64–66 t/s measured. Same PCIe 3.0 x4 link, about 4 GB/s.

The 4070 sits in a chipset slot at PCIe 3.0 x4, about 4 GB/s. That makes loading slow: about 90 seconds to move 19 GB. It does not slow generation, because only small activations cross between the cards per token.

A prediction that was wrong

Before the second card went in, a chat assistant predicted the x4 slot would act "like a speed brake" and keep decode at 19 to 22 t/s. The measured result was 64 to 66 t/s. Measure, do not predict.

Power​

Unverified

An 800 W supply is close to the limit with both cards under load. An unverified estimate from the same chat: 5070 Ti 250 to 300 W, 4070 200 to 220 W, CPU and the rest 150 to 200 W, spikes of 650 to 720 W. The suggested fix is a power limit of about 80 % on both cards, which it claimed costs 1 to 2 % of speed. I have not measured that.

# Example: cap GPU power (watts) with nvidia-smi. Pick values for your cards.
sudo nvidia-smi -i 0 -pl 240
sudo nvidia-smi -i 1 -pl 160

Resizable BAR is not the problem​

Out-of-memory crashes with gigabytes free look like a BAR or MMIO problem. They were not. Both cards showed full 16 GB resizable BARs above 4 GB, and Above 4G Decoding was working. The real cause was context size. See Context and fit-target.

The model: Qwen3.8-27B is a hybrid​

The GGUF metadata says general.architecture = qwen35. It is not a plain dense transformer.

Qwen3.8-27B · 65 blocks
  • full attention (≈16 layers): KV cache grows with context
  • Gated DeltaNet linear attention (≈48): fixed-size state
  • separate MTP head: blk.64.nextn.*
dense 65-layer model≈ 4x
Only ~16 of 65 layers pay for context. That is why 262144 fits in 28 GB.
qwen35.block_count = 65
qwen35.full_attention_interval = 4 only every 4th layer is full attention
qwen35.ssm.state_size = 128 the others are linear attention (Gated DeltaNet)
qwen35.attention.head_count_kv = 4
qwen35.attention.key_length = 256
qwen35.context_length = 262144

About 16 layers keep a KV cache that grows with context. About 48 carry a fixed-size recurrent state. -ctk and -ctv only change the KV cache, not the recurrent state.

The MTP draft head is inside the GGUF (blk.64.nextn.*, qwen35.nextn_predict_layers 1). You do not need a separate draft model.

KV cache cost​

H3 / KV cache cost

Full context: 8.5 GiB.
q4_0 saves ~4 GiB, not 12.

4 KV heads × 256 dim × 2 (K+V) × bytes × ~16 layers

at 131072 tokens

q8_0 · 4.25 GiB

q5_1 · 3.0 GiB

no CUDA FA kernel

q4_0 · 2.25 GiB

at 262144 tokens

q8_0 · 8.5 GiB

q5_1 · 6.0 GiB

no CUDA FA kernel

q4_0 · 4.5 GiB

Dense-model comparison: ≈ 4x q8_0 at 262144

All bars share a 0–36 GiB scale. The dotted bar is an estimate, not a measurement.

On Unsloth Studio this estimate ran ~40 % low. Trust the server's fit log near the edge.

Formula: 4 KV heads x 256 head dim x 2 (K and V) x bytes per value x ~16 layers.

KV typePer tokenAt 131072At 262144
q8_034 KiB4.25 GiB8.5 GiB
q5_124 KiB3.0 GiB6.0 GiB
q4_018 KiB2.25 GiB4.5 GiB

A dense 65-layer model would need about four times this. Going from q8_0 to q4_0 saves about 4 GiB at full context, not the 12 GiB you might expect.

The estimate runs low

On Unsloth Studio this formula was about 40 % below what the server actually budgeted. Near the edge, trust the server's own fit log.

Valid KV types in llama.cpp​

f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1. There is no FP8 KV cache in llama.cpp. Unsloth's FP8 advice is for their vLLM path.

On CUDA, flash attention only has kernels for f16, bf16, q8_0 and q4_0. The q5 types fall back to the CPU. That is why rule 4 exists.

Which quant​

Published KL divergence against BF16 (lower is closer to the full model):

QuantKL divergenceTop-1 agreementSize
UD-Q4_K_XL0.0095596.02 %17.9 GB
UD-Q5_K_XL0.0043797.28 %about 20 GB
UD-Q6_K0.0024297.93 %22.4 GB
Q8_00.0006498.93 %28.9 GB

What I use and why:

  • UD-Q4_K_XL is the daily driver on both servers. It leaves the most room for context and MTP.
  • UD-Q6_K runs on Unsloth Studio at 98304 context. It costs 8 to 10 % of speed against Q4.
  • UD-Q5_K_XL was used on llama-swap for three days. It was replaced when Q4 plus full context turned out faster.
  • Only one vision projector is kept, mmproj-Qwen3.8-27B-F16.gguf, renamed from the generic mmproj-F16.gguf so two models could not collide.
Unverified claims about the quants
Unverified

These came from a NotebookLM summary of YouTube videos. They are not checked.

  • UD-Q4_K_XL uses Unsloth Dynamic quantization with an importance matrix and keeps sensitive tensors at 5 bits or more.
  • Unsloth replaced its GGUF builds on 19 August. Pin a Hugging Face revision if you need reproducible results.
  • Plain uniform Q4 weights may damage the gated delta layers and cause silent drift.

Community reading​

Next step: llama-swap setup