Skip to main content

Clients and harnesses

Local LLM stackNote

Who talks to what​

ClientServerModel
Deepseek Harness (dsh), Linux and macOS desktopUnsloth Studio, port 8888q4-agent preset
Mindshub Cowork (MacBook, over the LAN)llama-swap, port 8081qwen3.8-27b for all roles
opencode, including OMO agent pinsllama-swapqwen3.8-27b
Chat bot agentsllama-swapqwen3.8-27b-uncensored
Speech-to-speech profilesllama-swapqwen3.8-27b
Security agent toolingllama-swapmodel list pulled live from /v1/models

Any OpenAI-compatible client works. Point it at http://<server>:8081/v1 or http://<server>:8888/v1.

Rules for clients​

Match context and max_tokens​

A client that thinks the window is bigger than the server's sends prompts the server cannot take. A client that thinks it is smaller truncates prompts without an error. Both happened:

  • opencode advertised 196608 after the server dropped to 131072.
  • The MacBook's dsh settings stayed at 65536 against a live 98304.

On Unsloth Studio, uctl syncs this for dsh. See Unsloth Studio. For everything else, update the client when you change -c.

Leave max_tokens room for thinking​

Thinking tokens count against max_tokens. A 35B model on this box reasoned for 600 to 1100 tokens before a tool call. At max_tokens: 1024 it returned nothing. At 4096 it worked. Set 4096 or more for any agent loop, and keep the server's --reasoning-budget well below it.

Do not split roles across models​

Mindshub Cowork has planning, coding and routing roles. It is tempting to give each a different model. llama-swap holds one model at a time, and there is no VRAM for a second. Each role switch would be a full reload of 15 seconds or more, 30 to 60 times per task. Keep all roles on one model.

Most clients send no sampler values​

An audit found that opencode, LiteLLM (with drop_params: true) and the chat agents send no sampler settings. They inherit whatever the server sets. So the server flags are what every caller gets. A client that does send temperature still wins, so check new clients.

Prompt caching​

L4 / Cache the prefix

Change the start,
repeat the work.

Stable prefix

systemreusetoolsreusehistoryreusenewprocess

reused prefix

Early change

systemprocesstoolsprocesshistoryprocessnewprocess

everything after a change is processed again

Change at the end

systemreusetoolsreusehistoryreusenewprocess

the earlier prefix stays cached

Prompt order: system prompt → tool list → history… → new message.

Keep the start stable. Put changing data at the end. The plus marker shows where data changed.

llama.cpp reuses the KV cache for a prompt prefix it has already seen. If anything near the start of the prompt changes, the whole prompt is processed again. Keep the system prompt and tool list stable, and put changing data at the end.

In dsh this was already mostly true: the system prompt is static, and the tool catalog stays the same across modes on purpose.

Further reading:

Reaching the server from another machine​

llama-swap listens on the LAN, and the firewall allows port 8081 only from the local subnet. Unsloth Studio is started with -H 0.0.0.0 so the MacBook can reach it. Neither config sets an API key, so never expose these ports to the internet.

cd ~
curl -s http://<server>:8081/v1/models

Next step: Pitfalls and fixes