Unsloth Studio and the agent harness
Local LLM stackNote
What it is for
Unsloth Studio is Unsloth's desktop and web app for running and fine-tuning models. It ships its own llama.cpp build. Here it serves the Deepseek Harness (dsh), an agent harness, on the Linux box and on a MacBook over the LAN.
Start it with LAN access:
cd ~
unsloth studio -p 8888 -H 0.0.0.0
What was wrong at the start
U1 / Follow the cause
MTP was supported.
Manual mode used its space.
GPU Memory = Manual, GPU Layers 66
--fit off --gpu-layers 66
nothing reserved for compute buffers + MTP draft context
RTX 4070 out of memory
+ --fit-target 128
never -ngl, never --tensor-split
llama_cpp_supports_mtp: true
spec_fallback_reason: runtime_error
Misleading UI message:MTP could not start for this model on the installed llama.cpp build.
Studio's GPU Memory = Manual with GPU Layers 66 becomes --fit off --gpu-layers 66. With the fitter off, nothing reserves room for compute buffers or the MTP draft context. The 4070 ran out of memory, and Studio then showed:
MTP could not start for this model on the installed llama.cpp build.
That message is wrong. The status API reported llama_cpp_supports_mtp: true and spec_fallback_reason: runtime_error. Manual mode had already used the VRAM that MTP needed.
GPU Memory = Auto, always, plus --fit-target 128 in Extra Arguments. Never -ngl, never --tensor-split.
Presets
A small Python script, uctl, holds all presets in one dictionary. There is no second config file.
| Preset | Context | KV | For |
|---|---|---|---|
q4-headless | 131072 | q8_0 | Daily driver |
q4-agent | 131072 | q8_0 | dsh and Godot MCP work, adds --reasoning-format deepseek and the thinking cap |
q4-max | 262144 | q4_0 | Full window. About 25 % slower alone, much worse with neighbours |
q6-headless | 98304 | q8_0 | Q6_K plain chat |
q6-best | 98304 | q8_0 | Q6_K at 8 to 10 % behind Q4 |
q6-long | 131072 | q4_0 | Q6_K only when you need more than 98304. Costs about 20 % |
cd ~
uctl list # presets
uctl status # context, KV, fit mode, MTP state, template version, VRAM
uctl load q4-agent # load, sync dsh, restart dsh, print status
uctl load q4-max 163840 # same preset, different context
uctl persist q4-agent # make it the Studio UI default too
uctl sync # push Studio's live context into dsh settings
load changes only the running server. persist writes the preset into Studio's settings, so the panel does not bring back Manual mode.
Full context is not free here
U3 / Decode at depth
Full context leaves
less room for neighbours.
MeasuredDecode t/s · brighter cells are faster.
131072 q8_0 · daily
ComfyUI up
131072 q8_0
ComfyUI stopped
131072 q4_0
ComfyUI up
262144 q4_0
ComfyUI up
262144 q4_0
ComfyUI stopped
131072 has slack: a 222 MiB neighbour changes nothing.
At 262144 and 98k depth: 26.7 → 13.2 t/s. The same neighbour halves deep decode.
Decode t/s at different prompt depths, MTP on, --fit-target 128, one slot:
Allocated -c | KV | Neighbour on GPU | 4k | 32k | 98k |
|---|---|---|---|---|---|
| 131072 | q8_0 | ComfyUI up | 52.5 | 44.7 | 32.2 |
| 131072 | q8_0 | ComfyUI stopped | 52.7 | 44.3 | 32.3 |
| 131072 | q4_0 | ComfyUI up | 50.0 | 46.8 | 32.4 |
| 262144 | q4_0 | ComfyUI up | 35.0 | 24.3 | 13.2 |
| 262144 | q4_0 | ComfyUI stopped | 39.2 | 36.1 | 26.7 |
What this shows:
- 131072 has slack. A 222 MiB neighbour changes nothing.
- 262144 is fragile. The same 222 MiB neighbour halves deep-context decode.
- KV type does not change speed. Take q8_0 for quality, and q4_0 only to buy context.
- 65536 buys nothing over 131072.
llama-swap runs 262144 at about 52 t/s. Studio's fork does not. Same GPUs, same GGUF, different binary. Measure on the server you use.
Where a slow 17 t/s came from
Four penalties stacked up: MTP had failed to start, the context was 193280 (past the knee), ComfyUI held 222 MiB, and the prompt was deep. At 131072 the same 98k-deep prompt decodes at 32 t/s instead of 13.
Q6_K against Q4_K_XL
| Model | -c | KV | 4k | 32k | 98k | Verdict |
|---|---|---|---|---|---|---|
| UD-Q4_K_XL | 131072 | q8_0 | 52.5 | 44.7 | 32.2 | Daily driver |
| UD-Q6_K | 98304 | q8_0 | 48.6 | 41.5 | The Q6 knee | |
| UD-Q6_K | 131072 | q8_0 | 27.3 | 19.9 | 11.5 | Loads, then spills to the CPU |
| UD-Q6_K | 131072 | q4_0 | 37.9 | 37.2 | 28.3 | Fits |
The recorded sum mixes GB and GiB. Confirm the units against the original fit log before using it as an overflow calculation.
Q6_K at 131072 with q8_0 loads fine, and status still says auto. Studio's fit log has the sum: 20.5 GB weights + 6.3 GB KV + 0.79 GB MTP reserve is more than the 26.9 GiB free. Gate on decode speed at 4k, or on the est. KV cache log line. Never on whether the load returned.
Every Run settings field
| Field | Set to | Why |
|---|---|---|
| Context Length | per preset | Allocating the native maximum is not free here |
| KV Cache Dtype | q8_0, q4_0 only to buy context | q5 types have no CUDA flash-attention kernel |
| Speculative Decoding | MTP | About 1.3x at shallow depth, never harmful |
| Draft Tokens | auto (2) | 2 against 3 made about 1 % difference |
| Parallel Slots | 1 | One user. More slots cost compute buffers |
| Batch sizes | auto | |
| Tensor Parallelism | off | Two different cards. Layer split measured faster |
| Vision | on | The projector stays on the CPU until an image arrives |
| GPU Memory | Auto | See above |
| GPU Layers | empty | Only read in Manual mode |
| Chat Template | Custom, froggeric v22.4 | See Reasoning and sampling |
| Extra Arguments | --fit-target 128 -t 8 -tb 16 --no-mmproj-offload --image-min-tokens 1024 | Plus the reasoning flags on agent presets |
The q4-agent launch, as a plain command
The same settings as one llama-server command, useful outside Studio:
cd ~
~/.local/bin/llama-server \
-m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf -a Qwen3.8-27B-UD-Q4_K_XL \
--host 0.0.0.0 --port 8888 \
-c 131072 --cache-type-k q8_0 --cache-type-v q8_0 -fa 1 \
--jinja --chat-template-file ~/models/qwen3.8-froggeric-v22.4.jinja \
--parallel 1 --no-context-shift --fit-target 128 -t 8 -tb 16 \
--no-mmproj-offload --image-min-tokens 1024 \
--reasoning-format deepseek --reasoning-budget 12288 \
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer." \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0
A saved copy of this command in my notes uses -c 196608. Which value is live is not recorded. 131072 is the measured choice.
Keeping the harness in step
dsh stores contextWindow and maxTokens per model in ~/.dsh/settings.yaml. If Studio reloads at a different context and dsh keeps the old numbers, prompts overflow or get truncated without an error.
Two layers keep them in step:
uctl loadrewrites both values for whatever Studio has loaded, withmaxTokenscapped atmin(32768, ctx / 4), then restarts dsh.- A systemd user timer runs the same sync every 60 seconds. It catches models loaded from the Studio UI. It only writes when a number changed.
The MacBook has its own copy of settings.yaml. uctl sync updates it over SSH, and caches the last pushed values so a sleeping laptop is not probed 1440 times a day.
The dsh desktop app reads its settings only at launch. The sync prints a reminder instead of quitting the app, because quitting would kill a running agent session.
dsh config suggestions that turned out not to exist
A list of tips suggested YAML keys such as agent.session_policy, scheduler.concurrency_limit, pipeline_mode and tool_execution_timeout. None of them exist in dsh. dsh is a plugin stack: you patch an existing plugin id in ~/.dsh/cordis.patch.yml. Always check with dsh --profile web --dump-config before you add config.
The one real gap was the agent-instructions plugin, which loads AGENTS.md. It ships disabled in the web profile, so project rules never reached the model. It is now enabled with maxBytes: 65536.
Next step: Reasoning and sampling
- Leads to
- Reasoning and sampling