Skip to main content

Unsloth Studio and the agent harness

Local LLM stackNote

What it is for​

Unsloth Studio is Unsloth's desktop and web app for running and fine-tuning models. It ships its own llama.cpp build. Here it serves the Deepseek Harness (dsh), an agent harness, on the Linux box and on a MacBook over the LAN.

Start it with LAN access:

cd ~
unsloth studio -p 8888 -H 0.0.0.0

What was wrong at the start​

U1 / Follow the cause

MTP was supported.
Manual mode used its space.

llama_cpp_supports_mtp: true
spec_fallback_reason: runtime_error

Misleading UI message:
MTP could not start for this model on the installed llama.cpp build.

The status API says support exists. The failure is memory allocation.

Studio's GPU Memory = Manual with GPU Layers 66 becomes --fit off --gpu-layers 66. With the fitter off, nothing reserves room for compute buffers or the MTP draft context. The 4070 ran out of memory, and Studio then showed:

MTP could not start for this model on the installed llama.cpp build.

That message is wrong. The status API reported llama_cpp_supports_mtp: true and spec_fallback_reason: runtime_error. Manual mode had already used the VRAM that MTP needed.

Fix

GPU Memory = Auto, always, plus --fit-target 128 in Extra Arguments. Never -ngl, never --tensor-split.

Presets​

A small Python script, uctl, holds all presets in one dictionary. There is no second config file.

PresetContextKVFor
q4-headless131072q8_0Daily driver
q4-agent131072q8_0dsh and Godot MCP work, adds --reasoning-format deepseek and the thinking cap
q4-max262144q4_0Full window. About 25 % slower alone, much worse with neighbours
q6-headless98304q8_0Q6_K plain chat
q6-best98304q8_0Q6_K at 8 to 10 % behind Q4
q6-long131072q4_0Q6_K only when you need more than 98304. Costs about 20 %
cd ~
uctl list # presets
uctl status # context, KV, fit mode, MTP state, template version, VRAM
uctl load q4-agent # load, sync dsh, restart dsh, print status
uctl load q4-max 163840 # same preset, different context
uctl persist q4-agent # make it the Studio UI default too
uctl sync # push Studio's live context into dsh settings

load changes only the running server. persist writes the preset into Studio's settings, so the panel does not bring back Manual mode.

Full context is not free here​

U3 / Decode at depth

Full context leaves
less room for neighbours.

Measured

Decode t/s · brighter cells are faster.

131072 q8_0 · daily
ComfyUI up

4k52.532k44.798k32.2

131072 q8_0
ComfyUI stopped

4k52.732k44.398k32.3

131072 q4_0
ComfyUI up

4k50.032k46.898k32.4

262144 q4_0
ComfyUI up

4k35.032k24.398k13.2

262144 q4_0
ComfyUI stopped

4k39.232k36.198k26.7

131072 has slack: a 222 MiB neighbour changes nothing.

At 262144 and 98k depth: 26.7 → 13.2 t/s. The same neighbour halves deep decode.

Decode t/s at different prompt depths, MTP on, --fit-target 128, one slot:

Allocated -cKVNeighbour on GPU4k32k98k
131072q8_0ComfyUI up52.544.732.2
131072q8_0ComfyUI stopped52.744.332.3
131072q4_0ComfyUI up50.046.832.4
262144q4_0ComfyUI up35.024.313.2
262144q4_0ComfyUI stopped39.236.126.7

What this shows:

  • 131072 has slack. A 222 MiB neighbour changes nothing.
  • 262144 is fragile. The same 222 MiB neighbour halves deep-context decode.
  • KV type does not change speed. Take q8_0 for quality, and q4_0 only to buy context.
  • 65536 buys nothing over 131072.
This contradicts the llama-swap page, and both are right

llama-swap runs 262144 at about 52 t/s. Studio's fork does not. Same GPUs, same GGUF, different binary. Measure on the server you use.

Where a slow 17 t/s came from​

Four penalties stacked up: MTP had failed to start, the context was 193280 (past the knee), ComfyUI held 222 MiB, and the prompt was deep. At 131072 the same 98k-deep prompt decodes at 32 t/s instead of 13.

Q6_K against Q4_K_XL​

Model-cKV4k32k98kVerdict
UD-Q4_K_XL131072q8_052.544.732.2Daily driver
UD-Q6_K98304q8_048.641.5The Q6 knee
UD-Q6_K131072q8_027.319.911.5Loads, then spills to the CPU
UD-Q6_K131072q4_037.937.228.3Fits
A bad fit does not error
Unverified

The recorded sum mixes GB and GiB. Confirm the units against the original fit log before using it as an overflow calculation.

Q6_K at 131072 with q8_0 loads fine, and status still says auto. Studio's fit log has the sum: 20.5 GB weights + 6.3 GB KV + 0.79 GB MTP reserve is more than the 26.9 GiB free. Gate on decode speed at 4k, or on the est. KV cache log line. Never on whether the load returned.

Every Run settings field​

FieldSet toWhy
Context Lengthper presetAllocating the native maximum is not free here
KV Cache Dtypeq8_0, q4_0 only to buy contextq5 types have no CUDA flash-attention kernel
Speculative DecodingMTPAbout 1.3x at shallow depth, never harmful
Draft Tokensauto (2)2 against 3 made about 1 % difference
Parallel Slots1One user. More slots cost compute buffers
Batch sizesauto
Tensor ParallelismoffTwo different cards. Layer split measured faster
VisiononThe projector stays on the CPU until an image arrives
GPU MemoryAutoSee above
GPU LayersemptyOnly read in Manual mode
Chat TemplateCustom, froggeric v22.4See Reasoning and sampling
Extra Arguments--fit-target 128 -t 8 -tb 16 --no-mmproj-offload --image-min-tokens 1024Plus the reasoning flags on agent presets

The q4-agent launch, as a plain command​

The same settings as one llama-server command, useful outside Studio:

cd ~
~/.local/bin/llama-server \
-m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf -a Qwen3.8-27B-UD-Q4_K_XL \
--host 0.0.0.0 --port 8888 \
-c 131072 --cache-type-k q8_0 --cache-type-v q8_0 -fa 1 \
--jinja --chat-template-file ~/models/qwen3.8-froggeric-v22.4.jinja \
--parallel 1 --no-context-shift --fit-target 128 -t 8 -tb 16 \
--no-mmproj-offload --image-min-tokens 1024 \
--reasoning-format deepseek --reasoning-budget 12288 \
--reasoning-budget-message "I have used my thinking budget. Concluding now and giving the answer." \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0

A saved copy of this command in my notes uses -c 196608. Which value is live is not recorded. 131072 is the measured choice.

Keeping the harness in step​

dsh stores contextWindow and maxTokens per model in ~/.dsh/settings.yaml. If Studio reloads at a different context and dsh keeps the old numbers, prompts overflow or get truncated without an error.

Two layers keep them in step:

  1. uctl load rewrites both values for whatever Studio has loaded, with maxTokens capped at min(32768, ctx / 4), then restarts dsh.
  2. A systemd user timer runs the same sync every 60 seconds. It catches models loaded from the Studio UI. It only writes when a number changed.

The MacBook has its own copy of settings.yaml. uctl sync updates it over SSH, and caches the last pushed values so a sleeping laptop is not probed 1440 times a day.

Restart the desktop app yourself

The dsh desktop app reads its settings only at launch. The sync prints a reminder instead of quitting the app, because quitting would kill a running agent session.

dsh config suggestions that turned out not to exist

A list of tips suggested YAML keys such as agent.session_policy, scheduler.concurrency_limit, pipeline_mode and tool_execution_timeout. None of them exist in dsh. dsh is a plugin stack: you patch an existing plugin id in ~/.dsh/cordis.patch.yml. Always check with dsh --profile web --dump-config before you add config.

The one real gap was the agent-instructions plugin, which loads AGENTS.md. It ships disabled in the web profile, so project rules never reached the model. It is now enabled with maxBytes: 65536.

Next step: Reasoning and sampling