File size: 17,882 Bytes
f6bc13b 7a53d78 f6bc13b 7a53d78 f6bc13b 18520bf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 | ---
license: mit
base_model: zai-org/GLM-5.3-Flash
tags:
- dgx-spark
- gb10
- sm121
- sglang
- speculative-decoding
- dflash2
- nvfp4
- recipe
---
# GLM-5.3-Flash + DFlash2 on 2Γ NVIDIA DGX Spark (GB10) β SGLang TP=2 recipe
Hey β I'm one of the folks running DGX Sparks at home, and this community's recipes are
the only reason my cluster works at all. tonyd2wild's GB10 forensics, MiaAI-Lab's
dual-Spark configs, hasso5703's DFlash2 writeup, LibertAIDAI's quant card β I've leaned
on all of them, so here's mine back.
This is GLM-5.3-Flash with the incoai DFlash2 drafter on the **SGLang** path (the
PR [#36507](https://github.com/sgl-project/sglang/pull/36507) branch everyone will get
by default once it merges). Getting it to boot on GB10 took a night and four fixes
nobody had written down yet β they're all here with patches and probes, so your
bring-up should take an hour instead. If you hit something new, open an issue and
I'll dig in with you.
## Current status (2026-08-29) β measured, in production
Running config: `start-LC4.sh` β fp8 KV + DFlash2 (D=5) + 8 concurrent streams + vision,
`--chunked-prefill-size 4096`, 131k context, ~242k-token KV pool.
| metric | value |
|---|---:|
| code single-stream | 28.6 tok/s |
| prose single-stream | 23.6 tok/s |
| **c8 aggregate** | **77.4 tok/s** (8/8 concurrent) |
| **c12 aggregate** | **83.2 tok/s** (12/12) |
| TTFT @16k warm | ~6.6 s |
| 100k-token prompt | PASS, 104 s |
| correctness under load | 44/44 (c1/c4/c8) |
| vision (image input) | working |
All warmed, temp 0, `stream:false`, n=5 medians, stock clocks. Three things that were
believed impossible on this path when we started are now measured working: fp8 KV cache
(upstream PR [#36904](https://github.com/sgl-project/sglang/pull/36904)), >2 concurrent
DFlash streams (issue [#36889](https://github.com/sgl-project/sglang/issues/36889)),
and 100k-token prompts (issue [#36941](https://github.com/sgl-project/sglang/issues/36941)).
See **LADDER.md** for every experiment including the failures and our own retractions,
**RESULTS.md** for the head-to-head against the EXL3+vLLM lane on this same rig.
## Full recipe below
Provenance chain (see FINDINGS.md for how each was verified):
| Artifact | Pin | Role |
|---|---|---|
| `lmsysorg/sglang:glm-5.3-flash-arm64` | digest `sha256:73f9294b78e38β¦`, pushed 2026-08-27T05:22Z | serving image (glm5_next + DFLASH infra) |
| sglang `refs/pull/36507/head` | `c4d5d45e506dcd978a65661a503eda1a272c39a4` | branch the image tracks; head now includes PR #36708 |
| PR #36708 (merged into that branch 18:23Z) | +31/β5 on `models/glm5_next.py` | DFLASH capture adapter β **newer than the image; we patch it in** (IMPLEMENTATION.md) |
| `LibertAIDAI/GLM-5.3-Flash-NVFP4` | 181 GiB, 48-ish shards β record actual count at download | target weights; card's own test = this image, 2Γ GB10 TP=2 |
| `incoai/GLM-5.3-Flash-DFlash2` | 1B BF16, single shard, **gated + research-only license** | drafter, block size 8 |
## 0. Hardware and current state
- 2Γ DGX Spark (GB10, sm_121, aarch64), 121.7 GB unified memory, 2.7 TB free disk each.
- `spark-1` (rank 0, serves HTTP) / `spark-2` (rank 1) β CX-7 back-to-back DAC, dual-rail
RoCE: `enp1s0f1np1` (10.10.10.1β.2/30) + `enP2p1s0f1np1` (10.10.11.1β.2/30), both MTU
9000. HCA twins `rocep1s0f1,roceP2p1s0f1` (matches `ibv_devices`; note `roceP2p1s0f0`
also exists β it is not ours).
- **Production today:** Qwen3.8-Flash-Next-NVFP4, containers `qwen38-flash-next-head` /
`qwen38-flash-next-worker`, spark-1:8899. It stays up until cutover. Both lanes cannot
run at once (each wants ~100 GB/node) β bring-up is a maintenance window with the Qwen
lane stopped, rollback is restarting it (Β§9).
- GLM lane port: **8901** (deliberately β 8899 so no client or probe can ever confuse
lanes mid-migration). `--served-model-name glm-5.3-flash-dflash2`.
## 1. Human-required steps β blocking, do these first
1. **Request access to `incoai/GLM-5.3-Flash-DFlash2`** (gate is `manual` β a human at
inco approves; wall-clock unknown, so file the request before anything else).
2. **Read the license before accepting.** It is *research-and-evaluation only*; no
commercial/production use without written consent (`contact@inco.ai`). Whether this
cluster's use qualifies is the operator's call, not this document's.
3. Place the HF token where the download step expects it (`~/.env.glm53` on spark-1,
`HF_TOKEN=β¦`). Never paste it into a shell command or this repo.
4. Approve the maintenance window (Qwen lane down for the duration of Β§5βΒ§8).
## 2. Preflight (both nodes, before the window)
```bash
# hotplug flag must be ABSENT; earlyoom must be STOPPED before launch
test ! -e /etc/nvidia/cx7-hotplug-enabled && echo hotplug-flag OK
sudo systemctl stop earlyoom && systemctl is-active earlyoom
# MTU 9000 end-to-end on BOTH rails (from spark-1):
ping -M do -s 8972 -c 3 10.10.10.2 # rail 1 β expect 0% loss
ping -M do -s 8972 -c 3 10.10.11.2 # rail 2 β expect 0% loss
# free page cache before the big load (GB10 unified-memory ritual):
sync && sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'
```
## 3. Image
Follow IMPLEMENTATION.md Β§Steps 0β3: pull `lmsysorg/sglang:glm-5.3-flash-arm64`, check
whether upstream already rebuilt it past the 18:23Z merge, otherwise overlay the one
patched file β local tag `glm53-flash-dflash:c4d5d45e5`, `docker save | ssh spark-2
docker load`, **verify digests match on both nodes**, and run the patch-presence probe
(with its negative control) on both. Never build independently on the worker.
## 4. Weights (tmux on spark-1; poll, don't block)
```bash
set -a; source ~/.env.glm53; set +a # HF_TOKEN for the gated drafter
huggingface-cli download LibertAIDAI/GLM-5.3-Flash-NVFP4 # ~181 GiB
huggingface-cli download incoai/GLM-5.3-Flash-DFlash2 # ~2 GiB, gated
```
Then rsync the HF cache to spark-2 (both nodes need both repos), record
`safetensor shard count and missing=0` for BOTH repos on BOTH nodes in the deploy log,
and from then on run containers with `HF_HUB_OFFLINE=1` (a gated repo re-check at boot
fails without the token; offline mode sidesteps it β but only after the cache is
complete). Check cache ownership on both nodes afterward; container writes as root
through the bind mount.
## 5. Launch
One script, `start-glm53-dflash.sh <node-rank>`; run **rank 1 on spark-2 FIRST**, then
rank 0 on spark-1 (worker-first; matches the proven dual-Spark DFlash2 deploy).
```bash
#!/usr/bin/env bash
# start-glm53-dflash.sh <0|1>
set -euo pipefail
RANK="${1:?usage: start-glm53-dflash.sh <0|1>}"
docker run -d --name "glm53-dflash-rank${RANK}" \
--gpus all --network host --ipc host \
--ulimit memlock=-1:-1 --cap-add IPC_LOCK --device /dev/infiniband \
--memory 115g --memory-swap 115g \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
-e HF_HUB_OFFLINE=1 \
-e NCCL_IB_HCA=rocep1s0f1,roceP2p1s0f1 \
-e NCCL_IB_MERGE_NICS=1 \
-e NCCL_SOCKET_IFNAME=enp1s0f1np1,enP2p1s0f1np1 \
-e GLOO_SOCKET_IFNAME=enp1s0f1np1 \
-e TP_SOCKET_IFNAME=enp1s0f1np1 \
-e NCCL_CUMEM_ENABLE=0 \
-e NCCL_NVLS_ENABLE=0 \
glm53-flash-dflash:c4d5d45e5 \
python3 -m sglang.launch_server \
--model-path LibertAIDAI/GLM-5.3-Flash-NVFP4 \
--served-model-name glm-5.3-flash-dflash2 \
--trust-remote-code \
--tp-size 2 --nnodes 2 --node-rank "$RANK" \
--dist-init-addr 10.10.10.1:50051 \
--attention-backend dsa \
--dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
--moe-runner-backend flashinfer_cutlass \
--kv-cache-dtype bfloat16 \
--disable-shared-experts-fusion \
--reasoning-parser glm45 --tool-call-parser glm47 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path incoai/GLM-5.3-Flash-DFlash2 \
--speculative-num-draft-tokens 8 \
--mem-fraction-static 0.80 \
--context-length 65536 --max-running-requests 2 \
--disable-flashinfer-autotune \
--stream-interval 1 --sleep-on-idle \
--host 0.0.0.0 --port 8901
```
Boot is minutes-long (weight load + CUDA graph capture). Poll
`curl -s http://10.0.x.x:8901/v1/models` from the Mac mini (over the tailnet β never
loopback) and tail `docker logs -f glm53-dflash-rank0`.
### Why each flag (and which are UNMEASURED)
| Flag | Why | Evidence |
|---|---|---|
| `--attention-backend dsa`, `--dsa-*-backend tilelang` | GLM-5.3's 11 deepseek-sparse-attention layers; tilelang is the backend the quant author ran on GB10 | LibertAIDAI card, tested 2Γ GB10 TP=2 |
| `--moe-runner-backend flashinfer_cutlass` | NVFP4 routed-expert path on Blackwell | same |
| `--kv-cache-dtype bfloat16` | the GB10-tested config; fp8 KV on this arch on sm_121 is unproven in SGLang (vLLM needed a CTA-tile cap for GB10's 101 KB smem) | same; fp8 KV = UNMEASURED, try later for KV headroom |
| `--disable-shared-experts-fusion`, `--reasoning-parser glm45`, `--tool-call-parser glm47` | quant author's tested config; parsers are GLM-4.x-lineage compatible | same |
| `--speculative-algorithm DFLASH` + drafter path | the point of this recipe | PR #36708; **combination UNMEASURED on GB10** |
| `--speculative-num-draft-tokens 8` | drafter block size is 8 (7 draft + 1); 8 measured optimal for the Qwen3.8 GB10 DFlash2 deploy | drafter card; forum 380732 |
| **no** `--speculative-draft-attention-backend fa4` | incoai's quickstart flag, written for GB300; fa4 on sm_121 unverified β the proven GB10 DFlash2 deploy used the default | risk #1 in IMPLEMENTATION.md |
| `--mem-fraction-static 0.80` | 0.84 was the card's value *without* a drafter; drafter adds ~2.3 GiB + capture buffers. 0.90 is the GB10 ceiling for a 15 GiB model; **0.95 hard-reboots GB10 in graph capture**. With ~91 GiB weights/node we start low. Raise to 0.84 only past all gates, one step, watching host free mem | UNMEASURED for this model; cliffs from forum 380732 |
| `--context-length 65536 --max-running-requests 2` | quant author's tested envelope; also caps contention while the tilelang-collapse risk (G7) is unretired. Model supports 1M; raising ctx is a later, gated experiment | LibertAIDAI card |
| `--disable-flashinfer-autotune` | autotuner's 25β40 GB transient allocations are invisible to SGLang accounting on unified memory; also non-deterministic boots | hasso5703 recipe |
| `--memory 115g` docker cap | fail as container-OOM, not host wedge (wedged GB10 = unplug/replug recovery) | GB10 unified-memory lesson |
| `--stream-interval 1`, `--sleep-on-idle` | client token-count fidelity; idle CPU-spin fix | forum 380732 |
| `NCCL_IB_HCA` both twins, `MERGE_NICS=1`, three socket vars (NCCL both rails, GLOO/TP first rail only) | dual-rail fabric β184 Gb/s; Gloo/TP steer TCP control plane | cluster baseline; **confirm inside the container, not the shell β G2** |
| `NCCL_CUMEM_ENABLE=0`, `NCCL_NVLS_ENABLE=0` | required in the proven GB10 dual-Spark DFlash2 deploy | forum 380732 |
## 6. Verification gates (in order; a failed gate is a hard stop)
- **G1 β patch presence** (IMPLEMENTATION.md Step 3), both nodes, WITH the negative
control against the unpatched image.
- **G2 β env at point of effect:** `docker exec glm53-dflash-rank0 env | grep -E
'NCCL|GLOO|TP_SOCKET'` on both nodes. A variable you did not confirm arrived is a
variable you did not set.
- **G3 β boot log:** DFLASH worker init lines present; no silent fallback to
non-speculative; no NaN/assert warnings from tilelang/DSA init. `NET/IB` (RoCE) in
NCCL init lines, both HCAs listed.
- **G4 β identity from the artifact:** `docker exec glm53-dflash-rank0 cat /proc/1/cmdline
| tr '\0' ' '` must show `LibertAIDAI/GLM-5.3-Flash-NVFP4` and the DFLASH flags.
SGLang will echo whatever served-name you configured β argv is the evidence,
`/v1/models` is a label.
- **G5 β real generation, cross-tailnet, anti-echo:** from the Mac mini, `stream:false`,
a prompt whose correct answer shares no 10-gram with the prompt. Assert the completion
is not a prompt echo (the known sm_121 vLLM failure shape), is coherent, and
`usage.completion_tokens` > 50. Never probe from loopback.
- **G6 β losslessness spot-check (one-time):** same 5 prompts, temp 0, against a
DFLASH-off launch (drop the three speculative flags) β outputs must match token-for-token.
DFlash2 is lossless by construction; this catches a broken capture/verify path, which
is exactly the part we patched in. Costs one extra boot cycle; worth it once.
- **G7 β acceptance, by NAME, with profile:** scrape `/metrics`, match metric names
containing `spec_accept` (gauges on SGLang β sample DURING active decode, not idle).
Expect the code-vs-prose spread (Qwen3.8 reference: ~5/8 code, ~3/8 prose β GLM values
UNMEASURED). A flat/degenerate accept profile with normal-looking tok/s = broken
drafter, stop. Then the contention probe: 2 concurrent code generations, scan outputs
for the token-collapse signature (runs of `!` / token-0 floods) seen once on our Qwen
lane's tilelang path.
- **G8 β fabric really carrying traffic:**
`/sys/class/infiniband/rocep1s0f1/ports/1/counters/port_xmit_data` (and the P2 twin)
advancing by GB-scale deltas during a long decode, both rails. `/sys/class/net`
statistics stay near zero for RDMA β that is expected, not idle.
Append every gate's numbers + exact commands to the deploy log.
## 7. Benchmark protocol (only after all gates)
Every number is recorded with **prompt name, max_tokens, and clock state** β all three
move results more than most effects being measured.
1. Warm: 2Γ real 800-token generations (cold penalty ~30% on this cluster's experience;
returns after idle).
2. `stream:false`, read `usage.completion_tokens`, wall-clock from response timing.
3. Fixed prompt set: `code-1` (implement a nontrivial function, ~200-token prompt) and
`prose-1` (essay), 800 max_tokens, nβ₯5 each, report median Β± spread.
4. Concurrency: c1 and c2 (the max-running-requests cap). Aggregate and per-stream.
5. Same protocol once with the three DFLASH flags removed β the speedup ratio, measured
not vibed. Reference points, clearly not ours: same model/hardware on vLLM+MTP =
21.8 tok/s decode (tonyd2wild, 2026-08-27); Qwen3.8-27B+DFlash2 dual-Spark = ~87 tok/s
code / ~41 prose (forum 380732). **GLM+DFlash2: UNMEASURED until this step.**
## 8. Cutover
Only after Β§6+Β§7: repoint clients (Hermes first) from :8899 to :8901, watch a matching
request appear in spark-1's GLM `/metrics` (provenance by metrics, not by client UI),
then decommission the Qwen containers in a later, separate decision. Keep the Qwen
image + cache untouched for rollback regardless.
## 9. Rollback (to the Qwen lane)
```bash
# both nodes:
docker stop glm53-dflash-rank0 glm53-dflash-rank1 2>/dev/null || docker ps # stop GLM lane
# then relaunch the Qwen lane with its EXISTING start script (worker first),
# containers qwen38-flash-next-worker / -head, image qwen38-flashnext-dspark:local
```
Verify rollback the same way as deploy: argv (G4), cross-tailnet generation (G5).
Weights and caches for both lanes coexist on disk (2.7 TB free) β rollback is a restart,
never a re-download.
## 10. Known issues β the honest list
| Issue | Status |
|---|---|
| **Nobody has run GLM-5.3+DFlash2 anywhere public** β drafter downloads: 0; enabling PR merged 2026-08-27 18:23Z | every combined number UNMEASURED; expect at least one gate to fail on first bring-up |
| Published image predates the DFLASH adapter by 13 h | one-file overlay, IMPLEMENTATION.md; check for upstream rebuild first |
| `fa4` draft attention backend (incoai quickstart) on sm_121 | unverified; omitted β default used, per the proven GB10 DFlash2 deploy |
| Drafter license | research/eval only β production use needs inco's written consent; operator decision |
| Drafter gate is manual | blocking human step; 0 downloads implies no queue history to estimate approval latency |
| mem-fraction-static with drafter on 91 GiB-weight model | UNMEASURED; start 0.80, ceiling 0.84; GB10 hard-reboots near 0.95 in graph capture (recovery: unplug/replug, power button is dead when wedged) |
| DFlash2 Γ YaRN | incompatible (forum 380732 build). We run native 65536 ctx β do not bolt YaRN on later to extend; the model is natively 1M, extend via `--context-length` + KV budget instead, as a gated experiment |
| tilelang collapse-under-contention (`!`-floods) seen on our Qwen lane | mitigation: max-running-requests 2 + G7 contention probe; not yet observed on glm5_next path |
| Prefix/radix cache on SM_121 hit `CUBLAS_STATUS_INTERNAL_ERROR` on one other model | not disabling preemptively; if hit, add `--disable-radix-cache` and log it |
| fp8 draft KV / fp8 target KV / ctx > 65536 / max-running > 2 / TP4 | all later, gated experiments; each currently UNMEASURED on this stack |
## Prior art, pinning, and license notes (2026-08-28 sweep)
- **Prior art:** [Tutanka01/glm5.3-flash-2x-dgx-spark-nvfp4](https://github.com/Tutanka01/glm5.3-flash-2x-dgx-spark-nvfp4)
(created 2026-08-26) published an SGLang TP=2 GB10 deployment of this model ~30h before ours,
with DFlash2 as a secondary profile. Our claim is accordingly the narrower one: the first
*dedicated, tuned and fully-measured* DFlash2-on-SGLang recipe (D-sweep, concurrency ladder,
fp8-KV port, multimodal). Credit where due.
- **Baseline pin:** every number in RESULTS/LADDER was measured on SGLang branch
`xinyuan/glm-5.3-flash-support` @ `aa8c950a3` (+ our patches). sm_121 fixes are landing
upstream at pace (e.g. #36755 now upstream β drop that patch on rebase; #36649 trtllm-gen
sparse decode); numbers move with the base.
- **Drafter license:** `incoai/GLM-5.3-Flash-DFlash2` is **cc-by-nc-nd-4.0** β non-commercial
AND no-derivatives: do not ship modified or requantized drafters.
|