Hardware and model
Local LLM stackNote
The box
| Part | Detail |
|---|---|
| GPU0 | RTX 5070 Ti, 16 GB (16303 MiB usable), top slot, PCIe 5.0 x16 |
| GPU1 | RTX 4070, 12 GB, bottom slot, PCIe 3.0 running at x4 |
| Board | ASUS ROG STRIX B660-F GAMING WIFI |
| CPU | i7-13700K, 8 P-cores + 8 E-cores, 24 threads |
| RAM | 32 GB |
| PSU | 800 W |
The x4 slot is fine for inference

The 4070 sits in a chipset slot at PCIe 3.0 x4, about 4 GB/s. That makes loading slow: about 90 seconds to move 19 GB. It does not slow generation, because only small activations cross between the cards per token.
Before the second card went in, a chat assistant predicted the x4 slot would act "like a speed brake" and keep decode at 19 to 22 t/s. The measured result was 64 to 66 t/s. Measure, do not predict.
Power
UnverifiedAn 800 W supply is close to the limit with both cards under load. An unverified estimate from the same chat: 5070 Ti 250 to 300 W, 4070 200 to 220 W, CPU and the rest 150 to 200 W, spikes of 650 to 720 W. The suggested fix is a power limit of about 80 % on both cards, which it claimed costs 1 to 2 % of speed. I have not measured that.
# Example: cap GPU power (watts) with nvidia-smi. Pick values for your cards.
sudo nvidia-smi -i 0 -pl 240
sudo nvidia-smi -i 1 -pl 160
Resizable BAR is not the problem
Out-of-memory crashes with gigabytes free look like a BAR or MMIO problem. They were not. Both cards showed full 16 GB resizable BARs above 4 GB, and Above 4G Decoding was working. The real cause was context size. See Context and fit-target.
The model: Qwen3.8-27B is a hybrid
The GGUF metadata says general.architecture = qwen35. It is not a plain dense transformer.
- full attention (≈16 layers): KV cache grows with context
- Gated DeltaNet linear attention (≈48): fixed-size state
- separate MTP head: blk.64.nextn.*
qwen35.block_count = 65
qwen35.full_attention_interval = 4 only every 4th layer is full attention
qwen35.ssm.state_size = 128 the others are linear attention (Gated DeltaNet)
qwen35.attention.head_count_kv = 4
qwen35.attention.key_length = 256
qwen35.context_length = 262144
About 16 layers keep a KV cache that grows with context. About 48 carry a fixed-size recurrent state. -ctk and -ctv only change the KV cache, not the recurrent state.
The MTP draft head is inside the GGUF (blk.64.nextn.*, qwen35.nextn_predict_layers 1). You do not need a separate draft model.
KV cache cost
H3 / KV cache cost
Full context: 8.5 GiB.
q4_0 saves ~4 GiB, not 12.
4 KV heads × 256 dim × 2 (K+V) × bytes × ~16 layers
at 131072 tokens
q8_0 · 4.25 GiB
q5_1 · 3.0 GiB
no CUDA FA kernel
q4_0 · 2.25 GiB
at 262144 tokens
q8_0 · 8.5 GiB
q5_1 · 6.0 GiB
no CUDA FA kernel
q4_0 · 4.5 GiB
Dense-model comparison: ≈ 4x q8_0 at 262144
On Unsloth Studio this estimate ran ~40 % low. Trust the server's fit log near the edge.
Formula: 4 KV heads x 256 head dim x 2 (K and V) x bytes per value x ~16 layers.
| KV type | Per token | At 131072 | At 262144 |
|---|---|---|---|
| q8_0 | 34 KiB | 4.25 GiB | 8.5 GiB |
| q5_1 | 24 KiB | 3.0 GiB | 6.0 GiB |
| q4_0 | 18 KiB | 2.25 GiB | 4.5 GiB |
A dense 65-layer model would need about four times this. Going from q8_0 to q4_0 saves about 4 GiB at full context, not the 12 GiB you might expect.
On Unsloth Studio this formula was about 40 % below what the server actually budgeted. Near the edge, trust the server's own fit log.
Valid KV types in llama.cpp
f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1. There is no FP8 KV cache in llama.cpp. Unsloth's FP8 advice is for their vLLM path.
On CUDA, flash attention only has kernels for f16, bf16, q8_0 and q4_0. The q5 types fall back to the CPU. That is why rule 4 exists.
Which quant
Published KL divergence against BF16 (lower is closer to the full model):
| Quant | KL divergence | Top-1 agreement | Size |
|---|---|---|---|
| UD-Q4_K_XL | 0.00955 | 96.02 % | 17.9 GB |
| UD-Q5_K_XL | 0.00437 | 97.28 % | about 20 GB |
| UD-Q6_K | 0.00242 | 97.93 % | 22.4 GB |
| Q8_0 | 0.00064 | 98.93 % | 28.9 GB |
What I use and why:
- UD-Q4_K_XL is the daily driver on both servers. It leaves the most room for context and MTP.
- UD-Q6_K runs on Unsloth Studio at 98304 context. It costs 8 to 10 % of speed against Q4.
- UD-Q5_K_XL was used on llama-swap for three days. It was replaced when Q4 plus full context turned out faster.
- Only one vision projector is kept,
mmproj-Qwen3.8-27B-F16.gguf, renamed from the genericmmproj-F16.ggufso two models could not collide.
Unverified claims about the quants
These came from a NotebookLM summary of YouTube videos. They are not checked.
- UD-Q4_K_XL uses Unsloth Dynamic quantization with an importance matrix and keeps sensitive tensors at 5 bits or more.
- Unsloth replaced its GGUF builds on 19 August. Pin a Hugging Face revision if you need reproducible results.
- Plain uniform Q4 weights may damage the gated delta layers and cause silent drift.
Community reading
- How can I run this well on an RTX 5070 Ti?, a Hugging Face discussion on the model page.
- Qwen3.8-27B Local Hardware Guide on kingy.ai.
- Unsloth's Qwen3.8 docs for quants and sampling.
Next step: llama-swap setup