Skip to main content

Context, fit-target and CPU layers

Local LLM stackNote

How llama.cpp places the model​

C1 / Auto-fit

A smaller reserve.
Fewer layers on the CPU.

context 262144 · KV q4_0

GPU0 · 16 GBGPU1 · 12 GBCPU
hatched: reservestorm: compute bufferscoral: KV cacheslices: layers

47.2 t/s

2 CPU layers

--fit-target 128 MiB

GPU0 RTX 5070 Ti
GPU1 RTX 4070

On a real 12712-token prompt: 52.3 t/s.

The default safety margin was the real context ceiling. Count CPU layers after every change. Tanks are schematic, not a measured memory breakdown. The control below shows exact VRAM measurements.

Recent llama.cpp builds have auto-fit (--fit on). It looks at free memory on each GPU and places layers so the weights, KV cache and compute buffers fit. Layers that do not fit go to the CPU. The llama.cpp multi-GPU docs say layer mode splits in proportion to available memory when --tensor-split is not set.

Auto-fit keeps a safety margin on each GPU. That margin is --fit-target (short form -fitt), in MiB.

-fitt, --fit-target MiB0,MiB1,... target margin per device for --fit
single value broadcast to all devices
default: 1024

The metric: CPU layers​

Auto-fit never fails loudly. When the context grows, it moves layers to the CPU and keeps going. The server loads, answers correctly, and gets slow. VRAM still looks full.

So after every change, count how many layers landed on the CPU. Run llama-server with -v, save the log, then:

cd ~
grep -E 'load_tensors: layer' srv.log | tail -66 | grep -c 'device CPU'

The last 66 lines are the final placement. Speed tracks this number almost linearly.

Speed tracks CPU layers, not 'it loaded'

UD-Q5_K_XL, default marginUD-Q4_K_XL, q4_0

0204060051015t/sCPU layers

Arrow: same context, margin only.

Tap or hover a point, or choose a measurement above.
ContextKVCPU layersDecodeGPU use
98304q8_0153.9 to 56.9 t/s37 to 48 %
131072q8_0733.2 t/s
163840q8_01228.0 t/s
262144q4_01418.4 t/s2 to 20 %

These rows use UD-Q5_K_XL at the default margin. They made 98304 look like a hard ceiling for two days. It was not.

The fix: -fitt 128​

At 262144 the default margin pushed 10 layers to the CPU and left about 5 GiB of VRAM unused. Measured with UD-Q4_K_XL and q4_0 KV:

context (-c)
--fit-target (MiB)
GPU013160 MiBGPU110437 MiBCPU layers 1025.8t/s

Bars on one 0–16303 MiB scale (GPU0 usable).

Re-tune -fitt every time you change -c.

Context-fittCPU layersVRAM GPU0 / GPU1 (MiB)Decode
2621441024 (default)1013160 / 1043725.8 t/s
262144256414180 / 1118936.9 t/s
262144128214430 / 1144547.2 t/s
2293761024 (default)513572 / 1048938.2 t/s
229376256114306 / 1075950.2 t/s

Look at the two 229376 rows. The margin alone moved decode from 38.2 to 50.2 t/s.

Always re-tune -fitt together with -c

If you sweep context at the default margin, you measure the margin, not the context.

Check it with a long prompt​

Short prompts hide out-of-memory errors that only show up when batch buffers grow. The chosen config was re-run on a real 12712-token prompt:

-fitt at 262144CPU layersPrefillDecode
1282835 t/s52.3 t/s
1922834 t/s45.9 t/s
2564632 t/s41.2 t/s

128 and 192 give byte-identical allocations, so the decode gap between them is run-to-run noise. 256 really does tip over.

Short-prompt prefill is mostly overhead

Short prompts reported about 100 t/s prefill. The real prompt gave 835 t/s. Use short-prompt prefill only to compare rows against each other.

Why you never pin layers or the split​

The community recipe for two GPUs is -ngl 99 --tensor-split 16,12. It worked here once, with Q4 quants at 98304, at 64 to 66 t/s. Then it failed at every ratio with UD-Q5_K_XL: 16,12, 18,10, 19,9, 2,1, 20,8 and 3,1.

Two reasons:

  1. -ngl 99 turns off auto-fit. The log says n_gpu_layers already set by user to 99, abort.
  2. A manual 16,12 actually put about 55 % of the load on the 12 GB card. Left alone, auto-fit balanced both cards at about 86 % and 85 %.

--main-gpu defaults to 0, the 5070 Ti. That is right: the bigger card should hold the KV cache.

Split mode: keep the default​

Measured at 163840 with q5_1 KV:

-smCPU layersDecode
layer (default)152.3 t/s
row041.2 t/s
tensorout of memory on load

row gets every layer onto the GPUs and is still 21 % slower, because cross-GPU traffic over the x4 link costs more than one CPU layer. An upstream Qwen3.8 MTP recipe reports +68 % from tensor or row on two 5060 Ti cards. That does not carry over to mixed cards on an x4 link.

About 45 % GPU use is normal

Layer split is pipeline parallelism. The GPUs take turns. High CPU with idle GPUs means the model is still loading, or real CPU layers. Check which.

GPU1 must stay empty​

With the small fit margin, even a small extra GPU allocation can hurt performance.
With -fitt 128 the 4070 has less than 1 GiB free. 154 MiB of ComfyUI was enough to matter.

With -fitt 128 the 4070 has less than 1 GiB free. A ComfyUI service that started at boot held about 154 MiB on it. That was enough to matter, so it is now manual start only.

cd ~
curl -s http://127.0.0.1:8081/unload # free llama-swap's VRAM before a ComfyUI run
systemctl --user start comfyui.service

If you want both at once, raise -fitt to 256 and accept 2 more CPU layers.

Host RAM counts too​

-c 262144 on Unsloth Studio cost about 25 GiB of host RAM on a 32 GB machine. A client 502 with no server error can be the kernel's out-of-memory killer. Check journalctl for an oom-kill.

Next step: Unsloth Studio