Context, fit-target and CPU layers
Local LLM stackNote
How llama.cpp places the model
C1 / Auto-fit
A smaller reserve.
Fewer layers on the CPU.
context 262144 · KV q4_0
47.2 t/s
2 CPU layers
--fit-target 128 MiB
GPU0 RTX 5070 Ti
GPU1 RTX 4070
On a real 12712-token prompt: 52.3 t/s.
Recent llama.cpp builds have auto-fit (--fit on). It looks at free memory on each GPU and places layers so the weights, KV cache and compute buffers fit. Layers that do not fit go to the CPU. The llama.cpp multi-GPU docs say layer mode splits in proportion to available memory when --tensor-split is not set.
Auto-fit keeps a safety margin on each GPU. That margin is --fit-target (short form -fitt), in MiB.
-fitt, --fit-target MiB0,MiB1,... target margin per device for --fit
single value broadcast to all devices
default: 1024
The metric: CPU layers
Auto-fit never fails loudly. When the context grows, it moves layers to the CPU and keeps going. The server loads, answers correctly, and gets slow. VRAM still looks full.
So after every change, count how many layers landed on the CPU. Run llama-server with -v, save the log, then:
cd ~
grep -E 'load_tensors: layer' srv.log | tail -66 | grep -c 'device CPU'
The last 66 lines are the final placement. Speed tracks this number almost linearly.
Speed tracks CPU layers, not 'it loaded'
UD-Q5_K_XL, default marginUD-Q4_K_XL, q4_0
Arrow: same context, margin only.
| Context | KV | CPU layers | Decode | GPU use |
|---|---|---|---|---|
| 98304 | q8_0 | 1 | 53.9 to 56.9 t/s | 37 to 48 % |
| 131072 | q8_0 | 7 | 33.2 t/s | |
| 163840 | q8_0 | 12 | 28.0 t/s | |
| 262144 | q4_0 | 14 | 18.4 t/s | 2 to 20 % |
These rows use UD-Q5_K_XL at the default margin. They made 98304 look like a hard ceiling for two days. It was not.
The fix: -fitt 128
At 262144 the default margin pushed 10 layers to the CPU and left about 5 GiB of VRAM unused. Measured with UD-Q4_K_XL and q4_0 KV:
Bars on one 0–16303 MiB scale (GPU0 usable).
Re-tune -fitt every time you change -c.
| Context | -fitt | CPU layers | VRAM GPU0 / GPU1 (MiB) | Decode |
|---|---|---|---|---|
| 262144 | 1024 (default) | 10 | 13160 / 10437 | 25.8 t/s |
| 262144 | 256 | 4 | 14180 / 11189 | 36.9 t/s |
| 262144 | 128 | 2 | 14430 / 11445 | 47.2 t/s |
| 229376 | 1024 (default) | 5 | 13572 / 10489 | 38.2 t/s |
| 229376 | 256 | 1 | 14306 / 10759 | 50.2 t/s |
Look at the two 229376 rows. The margin alone moved decode from 38.2 to 50.2 t/s.
-fitt together with -cIf you sweep context at the default margin, you measure the margin, not the context.
Check it with a long prompt
Short prompts hide out-of-memory errors that only show up when batch buffers grow. The chosen config was re-run on a real 12712-token prompt:
-fitt at 262144 | CPU layers | Prefill | Decode |
|---|---|---|---|
| 128 | 2 | 835 t/s | 52.3 t/s |
| 192 | 2 | 834 t/s | 45.9 t/s |
| 256 | 4 | 632 t/s | 41.2 t/s |
128 and 192 give byte-identical allocations, so the decode gap between them is run-to-run noise. 256 really does tip over.
Short prompts reported about 100 t/s prefill. The real prompt gave 835 t/s. Use short-prompt prefill only to compare rows against each other.
Why you never pin layers or the split
The community recipe for two GPUs is -ngl 99 --tensor-split 16,12. It worked here once, with Q4 quants at 98304, at 64 to 66 t/s. Then it failed at every ratio with UD-Q5_K_XL: 16,12, 18,10, 19,9, 2,1, 20,8 and 3,1.
Two reasons:
-ngl 99turns off auto-fit. The log saysn_gpu_layers already set by user to 99, abort.- A manual
16,12actually put about 55 % of the load on the 12 GB card. Left alone, auto-fit balanced both cards at about 86 % and 85 %.
--main-gpu defaults to 0, the 5070 Ti. That is right: the bigger card should hold the KV cache.
Split mode: keep the default
Measured at 163840 with q5_1 KV:
-sm | CPU layers | Decode |
|---|---|---|
| layer (default) | 1 | 52.3 t/s |
| row | 0 | 41.2 t/s |
| tensor | out of memory on load |
row gets every layer onto the GPUs and is still 21 % slower, because cross-GPU traffic over the x4 link costs more than one CPU layer. An upstream Qwen3.8 MTP recipe reports +68 % from tensor or row on two 5060 Ti cards. That does not carry over to mixed cards on an x4 link.
Layer split is pipeline parallelism. The GPUs take turns. High CPU with idle GPUs means the model is still loading, or real CPU layers. Check which.
GPU1 must stay empty

-fitt 128 the 4070 has less than 1 GiB free. 154 MiB of ComfyUI was enough to matter.With -fitt 128 the 4070 has less than 1 GiB free. A ComfyUI service that started at boot held about 154 MiB on it. That was enough to matter, so it is now manual start only.
cd ~
curl -s http://127.0.0.1:8081/unload # free llama-swap's VRAM before a ComfyUI run
systemctl --user start comfyui.service
If you want both at once, raise -fitt to 256 and accept 2 more CPU layers.
Host RAM counts too
-c 262144 on Unsloth Studio cost about 25 GiB of host RAM on a 32 GB machine. A client 502 with no server error can be the kernel's out-of-memory killer. Check journalctl for an oom-kill.
Next step: Unsloth Studio