Clients and harnesses
Local LLM stackNote
Who talks to what
| Client | Server | Model |
|---|---|---|
| Deepseek Harness (dsh), Linux and macOS desktop | Unsloth Studio, port 8888 | q4-agent preset |
| Mindshub Cowork (MacBook, over the LAN) | llama-swap, port 8081 | qwen3.8-27b for all roles |
| opencode, including OMO agent pins | llama-swap | qwen3.8-27b |
| Chat bot agents | llama-swap | qwen3.8-27b-uncensored |
| Speech-to-speech profiles | llama-swap | qwen3.8-27b |
| Security agent tooling | llama-swap | model list pulled live from /v1/models |
Any OpenAI-compatible client works. Point it at http://<server>:8081/v1 or http://<server>:8888/v1.
Rules for clients
Match context and max_tokens
A client that thinks the window is bigger than the server's sends prompts the server cannot take. A client that thinks it is smaller truncates prompts without an error. Both happened:
- opencode advertised 196608 after the server dropped to 131072.
- The MacBook's dsh settings stayed at 65536 against a live 98304.
On Unsloth Studio, uctl syncs this for dsh. See Unsloth Studio. For everything else, update the client when you change -c.
Leave max_tokens room for thinking
Thinking tokens count against max_tokens. A 35B model on this box reasoned for 600 to 1100 tokens before a tool call. At max_tokens: 1024 it returned nothing. At 4096 it worked. Set 4096 or more for any agent loop, and keep the server's --reasoning-budget well below it.
Do not split roles across models
Mindshub Cowork has planning, coding and routing roles. It is tempting to give each a different model. llama-swap holds one model at a time, and there is no VRAM for a second. Each role switch would be a full reload of 15 seconds or more, 30 to 60 times per task. Keep all roles on one model.
Most clients send no sampler values
An audit found that opencode, LiteLLM (with drop_params: true) and the chat agents send no sampler settings. They inherit whatever the server sets. So the server flags are what every caller gets. A client that does send temperature still wins, so check new clients.
Prompt caching
L4 / Cache the prefix
Change the start,
repeat the work.
Stable prefix
reused prefix
Early change
everything after a change is processed again
Change at the end
the earlier prefix stays cached
Prompt order: system prompt → tool list → history… → new message.
llama.cpp reuses the KV cache for a prompt prefix it has already seen. If anything near the start of the prompt changes, the whole prompt is processed again. Keep the system prompt and tool list stable, and put changing data at the end.
In dsh this was already mostly true: the system prompt is static, and the tool catalog stays the same across modes on purpose.
Further reading:
- Claude Code, llama.cpp, and the Hidden Prompt Cache Killer by Mykola Aleksandrov.
- Mastering Host-Memory Prompt Caching in llama-server, a llama.cpp discussion.
- llama.cpp constantly reprocessing huge prompts on r/LocalLLaMA.
Reaching the server from another machine
llama-swap listens on the LAN, and the firewall allows port 8081 only from the local subnet. Unsloth Studio is started with -H 0.0.0.0 so the MacBook can reach it. Neither config sets an API key, so never expose these ports to the internet.
cd ~
curl -s http://<server>:8081/v1/models
Next step: Pitfalls and fixes
- Leads to
- Pitfalls and fixes