File size: 17,882 Bytes
f6bc13b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7a53d78
 
 
 
f6bc13b
7a53d78
 
 
 
 
 
f6bc13b
18520bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
---
license: mit
base_model: zai-org/GLM-5.3-Flash
tags:
- dgx-spark
- gb10
- sm121
- sglang
- speculative-decoding
- dflash2
- nvfp4
- recipe
---

# GLM-5.3-Flash + DFlash2 on 2Γ— NVIDIA DGX Spark (GB10) β€” SGLang TP=2 recipe

Hey β€” I'm one of the folks running DGX Sparks at home, and this community's recipes are
the only reason my cluster works at all. tonyd2wild's GB10 forensics, MiaAI-Lab's
dual-Spark configs, hasso5703's DFlash2 writeup, LibertAIDAI's quant card β€” I've leaned
on all of them, so here's mine back.

This is GLM-5.3-Flash with the incoai DFlash2 drafter on the **SGLang** path (the
PR [#36507](https://github.com/sgl-project/sglang/pull/36507) branch everyone will get
by default once it merges). Getting it to boot on GB10 took a night and four fixes
nobody had written down yet β€” they're all here with patches and probes, so your
bring-up should take an hour instead. If you hit something new, open an issue and
I'll dig in with you.

## Current status (2026-08-29) β€” measured, in production

Running config: `start-LC4.sh` β€” fp8 KV + DFlash2 (D=5) + 8 concurrent streams + vision,
`--chunked-prefill-size 4096`, 131k context, ~242k-token KV pool.

| metric | value |
|---|---:|
| code single-stream | 28.6 tok/s |
| prose single-stream | 23.6 tok/s |
| **c8 aggregate** | **77.4 tok/s** (8/8 concurrent) |
| **c12 aggregate** | **83.2 tok/s** (12/12) |
| TTFT @16k warm | ~6.6 s |
| 100k-token prompt | PASS, 104 s |
| correctness under load | 44/44 (c1/c4/c8) |
| vision (image input) | working |

All warmed, temp 0, `stream:false`, n=5 medians, stock clocks. Three things that were
believed impossible on this path when we started are now measured working: fp8 KV cache
(upstream PR [#36904](https://github.com/sgl-project/sglang/pull/36904)), >2 concurrent
DFlash streams (issue [#36889](https://github.com/sgl-project/sglang/issues/36889)),
and 100k-token prompts (issue [#36941](https://github.com/sgl-project/sglang/issues/36941)).
See **LADDER.md** for every experiment including the failures and our own retractions,
**RESULTS.md** for the head-to-head against the EXL3+vLLM lane on this same rig.

## Full recipe below

Provenance chain (see FINDINGS.md for how each was verified):

| Artifact | Pin | Role |
|---|---|---|
| `lmsysorg/sglang:glm-5.3-flash-arm64` | digest `sha256:73f9294b78e38…`, pushed 2026-08-27T05:22Z | serving image (glm5_next + DFLASH infra) |
| sglang `refs/pull/36507/head` | `c4d5d45e506dcd978a65661a503eda1a272c39a4` | branch the image tracks; head now includes PR #36708 |
| PR #36708 (merged into that branch 18:23Z) | +31/βˆ’5 on `models/glm5_next.py` | DFLASH capture adapter β€” **newer than the image; we patch it in** (IMPLEMENTATION.md) |
| `LibertAIDAI/GLM-5.3-Flash-NVFP4` | 181 GiB, 48-ish shards β€” record actual count at download | target weights; card's own test = this image, 2Γ— GB10 TP=2 |
| `incoai/GLM-5.3-Flash-DFlash2` | 1B BF16, single shard, **gated + research-only license** | drafter, block size 8 |

## 0. Hardware and current state

- 2Γ— DGX Spark (GB10, sm_121, aarch64), 121.7 GB unified memory, 2.7 TB free disk each.
- `spark-1` (rank 0, serves HTTP) / `spark-2` (rank 1) β€” CX-7 back-to-back DAC, dual-rail
  RoCE: `enp1s0f1np1` (10.10.10.1↔.2/30) + `enP2p1s0f1np1` (10.10.11.1↔.2/30), both MTU
  9000. HCA twins `rocep1s0f1,roceP2p1s0f1` (matches `ibv_devices`; note `roceP2p1s0f0`
  also exists β€” it is not ours).
- **Production today:** Qwen3.8-Flash-Next-NVFP4, containers `qwen38-flash-next-head` /
  `qwen38-flash-next-worker`, spark-1:8899. It stays up until cutover. Both lanes cannot
  run at once (each wants ~100 GB/node) β€” bring-up is a maintenance window with the Qwen
  lane stopped, rollback is restarting it (Β§9).
- GLM lane port: **8901** (deliberately β‰  8899 so no client or probe can ever confuse
  lanes mid-migration). `--served-model-name glm-5.3-flash-dflash2`.

## 1. Human-required steps β€” blocking, do these first

1. **Request access to `incoai/GLM-5.3-Flash-DFlash2`** (gate is `manual` β€” a human at
   inco approves; wall-clock unknown, so file the request before anything else).
2. **Read the license before accepting.** It is *research-and-evaluation only*; no
   commercial/production use without written consent (`contact@inco.ai`). Whether this
   cluster's use qualifies is the operator's call, not this document's.
3. Place the HF token where the download step expects it (`~/.env.glm53` on spark-1,
   `HF_TOKEN=…`). Never paste it into a shell command or this repo.
4. Approve the maintenance window (Qwen lane down for the duration of Β§5–§8).

## 2. Preflight (both nodes, before the window)

```bash
# hotplug flag must be ABSENT; earlyoom must be STOPPED before launch
test ! -e /etc/nvidia/cx7-hotplug-enabled && echo hotplug-flag OK
sudo systemctl stop earlyoom && systemctl is-active earlyoom

# MTU 9000 end-to-end on BOTH rails (from spark-1):
ping -M do -s 8972 -c 3 10.10.10.2   # rail 1 β€” expect 0% loss
ping -M do -s 8972 -c 3 10.10.11.2   # rail 2 β€” expect 0% loss

# free page cache before the big load (GB10 unified-memory ritual):
sync && sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'
```

## 3. Image

Follow IMPLEMENTATION.md Β§Steps 0–3: pull `lmsysorg/sglang:glm-5.3-flash-arm64`, check
whether upstream already rebuilt it past the 18:23Z merge, otherwise overlay the one
patched file β†’ local tag `glm53-flash-dflash:c4d5d45e5`, `docker save | ssh spark-2
docker load`, **verify digests match on both nodes**, and run the patch-presence probe
(with its negative control) on both. Never build independently on the worker.

## 4. Weights (tmux on spark-1; poll, don't block)

```bash
set -a; source ~/.env.glm53; set +a       # HF_TOKEN for the gated drafter
huggingface-cli download LibertAIDAI/GLM-5.3-Flash-NVFP4        # ~181 GiB
huggingface-cli download incoai/GLM-5.3-Flash-DFlash2           # ~2 GiB, gated
```

Then rsync the HF cache to spark-2 (both nodes need both repos), record
`safetensor shard count and missing=0` for BOTH repos on BOTH nodes in the deploy log,
and from then on run containers with `HF_HUB_OFFLINE=1` (a gated repo re-check at boot
fails without the token; offline mode sidesteps it β€” but only after the cache is
complete). Check cache ownership on both nodes afterward; container writes as root
through the bind mount.

## 5. Launch

One script, `start-glm53-dflash.sh <node-rank>`; run **rank 1 on spark-2 FIRST**, then
rank 0 on spark-1 (worker-first; matches the proven dual-Spark DFlash2 deploy).

```bash
#!/usr/bin/env bash
# start-glm53-dflash.sh <0|1>
set -euo pipefail
RANK="${1:?usage: start-glm53-dflash.sh <0|1>}"

docker run -d --name "glm53-dflash-rank${RANK}" \
  --gpus all --network host --ipc host \
  --ulimit memlock=-1:-1 --cap-add IPC_LOCK --device /dev/infiniband \
  --memory 115g --memory-swap 115g \
  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
  -e HF_HUB_OFFLINE=1 \
  -e NCCL_IB_HCA=rocep1s0f1,roceP2p1s0f1 \
  -e NCCL_IB_MERGE_NICS=1 \
  -e NCCL_SOCKET_IFNAME=enp1s0f1np1,enP2p1s0f1np1 \
  -e GLOO_SOCKET_IFNAME=enp1s0f1np1 \
  -e TP_SOCKET_IFNAME=enp1s0f1np1 \
  -e NCCL_CUMEM_ENABLE=0 \
  -e NCCL_NVLS_ENABLE=0 \
  glm53-flash-dflash:c4d5d45e5 \
  python3 -m sglang.launch_server \
    --model-path LibertAIDAI/GLM-5.3-Flash-NVFP4 \
    --served-model-name glm-5.3-flash-dflash2 \
    --trust-remote-code \
    --tp-size 2 --nnodes 2 --node-rank "$RANK" \
    --dist-init-addr 10.10.10.1:50051 \
    --attention-backend dsa \
    --dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
    --moe-runner-backend flashinfer_cutlass \
    --kv-cache-dtype bfloat16 \
    --disable-shared-experts-fusion \
    --reasoning-parser glm45 --tool-call-parser glm47 \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path incoai/GLM-5.3-Flash-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --mem-fraction-static 0.80 \
    --context-length 65536 --max-running-requests 2 \
    --disable-flashinfer-autotune \
    --stream-interval 1 --sleep-on-idle \
    --host 0.0.0.0 --port 8901
```

Boot is minutes-long (weight load + CUDA graph capture). Poll
`curl -s http://10.0.x.x:8901/v1/models` from the Mac mini (over the tailnet β€” never
loopback) and tail `docker logs -f glm53-dflash-rank0`.

### Why each flag (and which are UNMEASURED)

| Flag | Why | Evidence |
|---|---|---|
| `--attention-backend dsa`, `--dsa-*-backend tilelang` | GLM-5.3's 11 deepseek-sparse-attention layers; tilelang is the backend the quant author ran on GB10 | LibertAIDAI card, tested 2Γ— GB10 TP=2 |
| `--moe-runner-backend flashinfer_cutlass` | NVFP4 routed-expert path on Blackwell | same |
| `--kv-cache-dtype bfloat16` | the GB10-tested config; fp8 KV on this arch on sm_121 is unproven in SGLang (vLLM needed a CTA-tile cap for GB10's 101 KB smem) | same; fp8 KV = UNMEASURED, try later for KV headroom |
| `--disable-shared-experts-fusion`, `--reasoning-parser glm45`, `--tool-call-parser glm47` | quant author's tested config; parsers are GLM-4.x-lineage compatible | same |
| `--speculative-algorithm DFLASH` + drafter path | the point of this recipe | PR #36708; **combination UNMEASURED on GB10** |
| `--speculative-num-draft-tokens 8` | drafter block size is 8 (7 draft + 1); 8 measured optimal for the Qwen3.8 GB10 DFlash2 deploy | drafter card; forum 380732 |
| **no** `--speculative-draft-attention-backend fa4` | incoai's quickstart flag, written for GB300; fa4 on sm_121 unverified β€” the proven GB10 DFlash2 deploy used the default | risk #1 in IMPLEMENTATION.md |
| `--mem-fraction-static 0.80` | 0.84 was the card's value *without* a drafter; drafter adds ~2.3 GiB + capture buffers. 0.90 is the GB10 ceiling for a 15 GiB model; **0.95 hard-reboots GB10 in graph capture**. With ~91 GiB weights/node we start low. Raise to 0.84 only past all gates, one step, watching host free mem | UNMEASURED for this model; cliffs from forum 380732 |
| `--context-length 65536 --max-running-requests 2` | quant author's tested envelope; also caps contention while the tilelang-collapse risk (G7) is unretired. Model supports 1M; raising ctx is a later, gated experiment | LibertAIDAI card |
| `--disable-flashinfer-autotune` | autotuner's 25–40 GB transient allocations are invisible to SGLang accounting on unified memory; also non-deterministic boots | hasso5703 recipe |
| `--memory 115g` docker cap | fail as container-OOM, not host wedge (wedged GB10 = unplug/replug recovery) | GB10 unified-memory lesson |
| `--stream-interval 1`, `--sleep-on-idle` | client token-count fidelity; idle CPU-spin fix | forum 380732 |
| `NCCL_IB_HCA` both twins, `MERGE_NICS=1`, three socket vars (NCCL both rails, GLOO/TP first rail only) | dual-rail fabric β‰ˆ184 Gb/s; Gloo/TP steer TCP control plane | cluster baseline; **confirm inside the container, not the shell β€” G2** |
| `NCCL_CUMEM_ENABLE=0`, `NCCL_NVLS_ENABLE=0` | required in the proven GB10 dual-Spark DFlash2 deploy | forum 380732 |

## 6. Verification gates (in order; a failed gate is a hard stop)

- **G1 β€” patch presence** (IMPLEMENTATION.md Step 3), both nodes, WITH the negative
  control against the unpatched image.
- **G2 β€” env at point of effect:** `docker exec glm53-dflash-rank0 env | grep -E
  'NCCL|GLOO|TP_SOCKET'` on both nodes. A variable you did not confirm arrived is a
  variable you did not set.
- **G3 β€” boot log:** DFLASH worker init lines present; no silent fallback to
  non-speculative; no NaN/assert warnings from tilelang/DSA init. `NET/IB` (RoCE) in
  NCCL init lines, both HCAs listed.
- **G4 β€” identity from the artifact:** `docker exec glm53-dflash-rank0 cat /proc/1/cmdline
  | tr '\0' ' '` must show `LibertAIDAI/GLM-5.3-Flash-NVFP4` and the DFLASH flags.
  SGLang will echo whatever served-name you configured β€” argv is the evidence,
  `/v1/models` is a label.
- **G5 β€” real generation, cross-tailnet, anti-echo:** from the Mac mini, `stream:false`,
  a prompt whose correct answer shares no 10-gram with the prompt. Assert the completion
  is not a prompt echo (the known sm_121 vLLM failure shape), is coherent, and
  `usage.completion_tokens` > 50. Never probe from loopback.
- **G6 β€” losslessness spot-check (one-time):** same 5 prompts, temp 0, against a
  DFLASH-off launch (drop the three speculative flags) β€” outputs must match token-for-token.
  DFlash2 is lossless by construction; this catches a broken capture/verify path, which
  is exactly the part we patched in. Costs one extra boot cycle; worth it once.
- **G7 β€” acceptance, by NAME, with profile:** scrape `/metrics`, match metric names
  containing `spec_accept` (gauges on SGLang β€” sample DURING active decode, not idle).
  Expect the code-vs-prose spread (Qwen3.8 reference: ~5/8 code, ~3/8 prose β€” GLM values
  UNMEASURED). A flat/degenerate accept profile with normal-looking tok/s = broken
  drafter, stop. Then the contention probe: 2 concurrent code generations, scan outputs
  for the token-collapse signature (runs of `!` / token-0 floods) seen once on our Qwen
  lane's tilelang path.
- **G8 β€” fabric really carrying traffic:**
  `/sys/class/infiniband/rocep1s0f1/ports/1/counters/port_xmit_data` (and the P2 twin)
  advancing by GB-scale deltas during a long decode, both rails. `/sys/class/net`
  statistics stay near zero for RDMA β€” that is expected, not idle.

Append every gate's numbers + exact commands to the deploy log.

## 7. Benchmark protocol (only after all gates)

Every number is recorded with **prompt name, max_tokens, and clock state** β€” all three
move results more than most effects being measured.

1. Warm: 2Γ— real 800-token generations (cold penalty ~30% on this cluster's experience;
   returns after idle).
2. `stream:false`, read `usage.completion_tokens`, wall-clock from response timing.
3. Fixed prompt set: `code-1` (implement a nontrivial function, ~200-token prompt) and
   `prose-1` (essay), 800 max_tokens, nβ‰₯5 each, report median Β± spread.
4. Concurrency: c1 and c2 (the max-running-requests cap). Aggregate and per-stream.
5. Same protocol once with the three DFLASH flags removed β†’ the speedup ratio, measured
   not vibed. Reference points, clearly not ours: same model/hardware on vLLM+MTP =
   21.8 tok/s decode (tonyd2wild, 2026-08-27); Qwen3.8-27B+DFlash2 dual-Spark = ~87 tok/s
   code / ~41 prose (forum 380732). **GLM+DFlash2: UNMEASURED until this step.**

## 8. Cutover

Only after Β§6+Β§7: repoint clients (Hermes first) from :8899 to :8901, watch a matching
request appear in spark-1's GLM `/metrics` (provenance by metrics, not by client UI),
then decommission the Qwen containers in a later, separate decision. Keep the Qwen
image + cache untouched for rollback regardless.

## 9. Rollback (to the Qwen lane)

```bash
# both nodes:
docker stop glm53-dflash-rank0 glm53-dflash-rank1 2>/dev/null || docker ps  # stop GLM lane
# then relaunch the Qwen lane with its EXISTING start script (worker first),
# containers qwen38-flash-next-worker / -head, image qwen38-flashnext-dspark:local
```

Verify rollback the same way as deploy: argv (G4), cross-tailnet generation (G5).
Weights and caches for both lanes coexist on disk (2.7 TB free) β€” rollback is a restart,
never a re-download.

## 10. Known issues β€” the honest list

| Issue | Status |
|---|---|
| **Nobody has run GLM-5.3+DFlash2 anywhere public** β€” drafter downloads: 0; enabling PR merged 2026-08-27 18:23Z | every combined number UNMEASURED; expect at least one gate to fail on first bring-up |
| Published image predates the DFLASH adapter by 13 h | one-file overlay, IMPLEMENTATION.md; check for upstream rebuild first |
| `fa4` draft attention backend (incoai quickstart) on sm_121 | unverified; omitted β€” default used, per the proven GB10 DFlash2 deploy |
| Drafter license | research/eval only β€” production use needs inco's written consent; operator decision |
| Drafter gate is manual | blocking human step; 0 downloads implies no queue history to estimate approval latency |
| mem-fraction-static with drafter on 91 GiB-weight model | UNMEASURED; start 0.80, ceiling 0.84; GB10 hard-reboots near 0.95 in graph capture (recovery: unplug/replug, power button is dead when wedged) |
| DFlash2 Γ— YaRN | incompatible (forum 380732 build). We run native 65536 ctx β€” do not bolt YaRN on later to extend; the model is natively 1M, extend via `--context-length` + KV budget instead, as a gated experiment |
| tilelang collapse-under-contention (`!`-floods) seen on our Qwen lane | mitigation: max-running-requests 2 + G7 contention probe; not yet observed on glm5_next path |
| Prefix/radix cache on SM_121 hit `CUBLAS_STATUS_INTERNAL_ERROR` on one other model | not disabling preemptively; if hit, add `--disable-radix-cache` and log it |
| fp8 draft KV / fp8 target KV / ctx > 65536 / max-running > 2 / TP4 | all later, gated experiments; each currently UNMEASURED on this stack |

## Prior art, pinning, and license notes (2026-08-28 sweep)
- **Prior art:** [Tutanka01/glm5.3-flash-2x-dgx-spark-nvfp4](https://github.com/Tutanka01/glm5.3-flash-2x-dgx-spark-nvfp4)
  (created 2026-08-26) published an SGLang TP=2 GB10 deployment of this model ~30h before ours,
  with DFlash2 as a secondary profile. Our claim is accordingly the narrower one: the first
  *dedicated, tuned and fully-measured* DFlash2-on-SGLang recipe (D-sweep, concurrency ladder,
  fp8-KV port, multimodal). Credit where due.
- **Baseline pin:** every number in RESULTS/LADDER was measured on SGLang branch
  `xinyuan/glm-5.3-flash-support` @ `aa8c950a3` (+ our patches). sm_121 fixes are landing
  upstream at pace (e.g. #36755 now upstream β€” drop that patch on rebase; #36649 trtllm-gen
  sparse decode); numbers move with the base.
- **Drafter license:** `incoai/GLM-5.3-Flash-DFlash2` is **cc-by-nc-nd-4.0** β€” non-commercial
  AND no-derivatives: do not ship modified or requantized drafters.