Skip to main content

OrcaSAQ-2 Cyber and DFlash2

Local LLM stackNote

What OrcaSAQ-2 Cyber is​

orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF is a mixed-precision imatrix quant of an abliterated Qwen3.8-27B. It is the same architecture as the stack's main model: qwen35, 65 blocks, embedding length 5120, built-in MTP head (nextn_predict_layers = 1), 262144 native context. The GGUF is 15.68 GB, about 2 GB lighter than UD-Q4_K_XL, so auto-fit keeps it resident with more headroom.

Because the base architecture matches, nothing new was needed: the shared mmproj-Qwen3.8-27B-F16.gguf and the froggeric chat template both apply unchanged, and the sampler row, fit-target rule and MTP flags are identical. The only behavioural difference is that the model is uncensored.

No meaningful refusals

This model is abliterated and refuses almost nothing. Keep it off any agent that executes tools or shell on its own output unless that is exactly what you want. The censored main model is the default for agent work; this twin is opt-in.

The third binary​

The overview says two servers, two binaries. Serving DFlash2 forced a correction: there are three llama.cpp binaries on this box, and no single one serves every preset.

BinaryPathCarriesq4_0 V-cacheDFlash2 drafter
Fork (build 601)~/.local/bin/llama-serverMoE-offload + MTP patchesGPU kernel (only binary with it)Cannot load (expected 81, got 58)
Upstream (build 200)~/src/llama-upstream/build/bin/llama-serverDFlash2 support (PR #27342)runs on CPU (8–14x prefill loss)Works, acceptance 0.54–0.60
Studio (build 10798)~/.unsloth/llama.cpp/build/bin/llama-serverUnsloth forkruns on CPULoads at half the acceptance

The two needs pull opposite ways. A q4_0 V-cache has a CUDA kernel only on the fork, so a long-context q4_0 preset must run on the fork. DFlash2 only loads on upstream. You cannot have both in one server.

DFlash2: the drafter that only loads on upstream​

DFlash2 is an external draft model, unlike MTP's built-in head. The drafter for this base is z-lab/Qwen3.8-27B-DFlash2-GGUF; the Q4_K_M copy (~940 MiB on GPU0) lives at ~/.ollama/local-gguf/dflash2/. Because OrcaSAQ shares the base model, the same drafter drafts for it.

Serving it on the fork fails at load:

done_getting_tensors: wrong number of tensors; expected 81, got 58

The fork's build predates z-lab's DFlash2 format. On the upstream binary the same flags work:

--spec-type draft-dflash -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--spec-draft-n-max 5 --spec-draft-p-min 0
One warning that is not an error

During load, upstream prints:

llama_init_from_model: failed to initialize the context: dflash requires
ctx_other to be set (this warning is normal during memory fitting)

It is benign fit-pass noise, as the parenthetical says. The server continues, loads the drafter, and binds. Gate on /health, not on this line. Treating it as fatal cost a wrong "it does not work" conclusion.

The drafter needs ~940 MiB free on GPU0, which is only there below 163840 context. DFlash2 therefore caps context at 131072 on this box unless you free room elsewhere.

Measured numbers​

Warm (two throwaway requests first), median on real-code prompts at temp 1.0, all with cpu_kv = 0.

PresetBinaryContextKVDrafterDecodeAcceptance
orca-speedupstream131072q8_0DFlash2 n599.2 t/s0.60
orca-maxfork262144q4_0MTP n3~59 t/s0.56
(OrcaSAQ on fork + MTP)fork131072q8_0MTP n370.8 t/s0.57
base q4-speed, for scaleupstream131072q8_0DFlash2 n591.6 t/s0.82

DFlash2 on upstream beats the fork's MTP by about +40 % at the same context, and the lighter weights put OrcaSAQ ahead of the base model's DFlash2 number despite lower draft acceptance. The prefill figures on short prompts are overhead-dominated and are not reported; see Benchmarking honestly.

The presets​

Both are in uctl.py and show in the qbitx Presets pane:

  • orca-speed — upstream, 131072, q8_0, DFlash2 n5. The fast daily and agent slot.
  • orca-max — fork, 262144, q4_0, MTP n3. Full native context; q4_0 V-cache needs the fork, so it cannot use DFlash2.

The same GGUF is also a llama-swap preset on port 8081 (orcasaq2-cyber-27b, fork + MTP) for the general clients that do not go through Unsloth Studio.

A broken template path was hiding under this

Adding the preset surfaced that the shared template flag pointed at qwen3.8-froggeric-v22.3.jinja, but only v22.4 exists on disk. That silently broke every preset's load with failed to open file. Fixed once in the shared macro. If a preset refuses to load after a template upgrade, check the version in the path first.

Next step: Pitfalls and fixes