File size: 17,667 Bytes
00d52e8 269cd26 344991c c84a0db 9eae477 a486ba9 16cff2a 50d8e1e 2db1b02 a158c6c b8a968e e91cfbc 0cee402 e91cfbc d72ca7c 498eadf d8b622a 1605678 2f5385e da788b0 18a14c3 da788b0 18a14c3 d2c1c38 1cc17a5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 | # Optimization ladder β GLM-5.3-Flash + DFlash2, 2Γ DGX Spark (GB10), SGLang TP=2
All numbers: warmed, temp 0, stream:false, 800 max_tokens, n=5 medians, stock clocks,
code-1/prose-1 prompts as in RESULTS.md. Keep rule: beats incumbent code median with the
19Γ21 gate + token-0 collapse scan passing.
## Round 1 (autonomous, night of 2026-08-27β28)
| rung | change | code | prose | verdict |
|---|---|---|---|---|
| first light | β | 27.6 | 20.7 | baseline |
| L2 | flashinfer autotune ON (+`--enable-metrics`) | **28.4** (26.9β29.8) | 20.6 | **KEPT** |
| L3 | + fp8 draft KV | 27.0 (23.9β28.4) | 20.7 | reverted β convert overhead exceeds bandwidth saved on a ~2 GB drafter cache |
| L4 | + fp8 target KV | β | β | **BOOT FAILED in 64 s**; failure logs lost to teardown (ladder now saves them); autopsy queued round 2 |
| L5 | ctx 131072, max-total 262144 | 28.0 (27.6β30.8) | **21.3** | reverted by keep-rule β **but note: 2Γ context for β1.4% code speed; worth keeping for serving. Operator's call. Confound: carried L3's fp8-draft-KV flag** |
Round-1 net: +3% code. The big lever (fp8 target KV) is blocked, not disproven.
## Round 2
| rung | change | code | prose | verdict |
|---|---|---|---|---|
| L1v2 | decode+v1 kernels stock tiles | β | β | BOOT FAILED, same 169,984 B smem error β `sparse_attention_fwd_kernel_v1` is in the verify path and must stay small |
| L1v3 | only `sparse_mla_fwd_decode_partial` stock | 28.5 (25.7β33.2) | 20.2 | kept (wash) β **finding: the GB10 small tile is ~free; tile geometry is not the bottleneck** |
Kernel tile map for sm_121 (measured): `qo_len` multi-token kernel β small tile REQUIRED;
`sparse_attention_fwd_kernel_v1` β small tile REQUIRED (verify path); `sparse_mla_fwd_decode_partial`
β either (no measurable difference).
| D6 | speculative-num-draft-tokens 8β6 | **29.7** (27.2β31.0) | **21.5** (20.4β23.4) | **KEPT β new incumbent.** Shorter draft block wastes less verify on doomed tokens at our acceptance profile |
Cumulative: 27.6 β 29.7 code (+7.6%) since first light.
| L4v2 | fp8 target KV (tilelang DSA) | β | β | **CLOSED: architecturally unsupported** β SGLang raises `tilelang DSA ... on CUDA requires a bfloat16 KV cache` at arg resolution. Round-1's 64s death explained |
| L4v3 | trtllm DSA backends + fp8 KV | β | β | **DIED: `TllmGenFmhaRunner: Unsupported arch`** β trtllm FMHA does not support sm_121 |
**fp8 target KV verdict on GB10 + SGLang + GLM-5.3: impossible on this stack today.**
tilelang forbids it on CUDA by upstream policy; trtllm kernels don't build for the chip.
(vLLM reached fp8 KV on GB10 only via hand-patched CTA tile caps β see tonyd2wild's recipe.)
| D4 | draft tokens 4 | 28.3 (28.2β28.7) | **23.3** (23.3β23.3) | reverted on code β but best prose of the campaign (+8% vs D6): code peaks at D=6, prose at D=4; D=5 queued |
| V4/V5 (first attempt) | upstream aa8c950a3 refresh + novel fat tile | β | β | both died pre-kernel on a partial-overlay skew (`hc_attn_to_mlp` β model file and communicator_mhc.py are a coupled pair); rebuilt as V4b/V5b with both files |
| V4b | upstream aa8c950a3 (official mHC capture fix + kpool changes) + our tiles | 29.5 (29.5β29.6) | **22.5** | **KEPT as incumbent** β code statistically tied with D6 (29.7, whose spread contains V4b), prose +4.7%, and it replaces our hand-guard with the official fix. Tie broken on provenance |
| V5b | novel fat tile 64/1/256 | β | β | smem 104,448 B > 101,376 β **3,072 bytes over.** Single-stage halved the request exactly as modeled |
| D5 | draft tokens 5 | 29.6 (29.6β29.6) | **23.3** (23.0β23.3) | **best combined config** β D6's code with D4's prose; takes incumbency on tie-break |
| V5c | fat tile 64/1/128 | β | β | smem 103,424 B β still 2 KB over. **Fat-tile chapter closed with a complete map**: 64-wide needs β₯103.4 KB in any shape; GB10 ceiling 101.4; 32-wide fits and costs nothing |
**Decoupled drafter** (killing the TP=2 all-reduce tax on the 1B draft model): mapped, viable, NOT attempted β
requires a separate drafter-server topology (`--decoupled-spec-role verifier/drafter-rank` + bind/connect
endpoints + rank). This is the #1 next-frontier item; expected the largest structural gain.
| **FINAL** | **v4b image + D5 flags** | **29.4** (28.2β30.4) | **23.4** (20.7β25.4) | **SHIP CONFIG** β best combined; official upstream code |
| FA3 | draft attention fa3 | β | β | CLOSED: `FA3 requires SM>=80 and SM<=90` β sm_121 outside the window. Draft-attention map complete: flashinfer only |
| SERVE | FINAL + ctx 131072 / max-total 262144 | 29.3 (25.2β31.8) | 22.7 (21.9β23.5) | **PRODUCTION CONFIG** β 2Γ context window at statistical parity; Hermes cut over to this endpoint 2026-08-28 |
Campaign complete. Closeout data in RESULTS.md; morning package in MORNING.md.
## Concurrency curves (2026-08-28, agent-style prompts, 400 tok, aggregate tok/s)
| c | DFlash (max_req 8) | no-spec (max_req 12) |
|---|---|---|
| 1 | **37.3** | 14.5 |
| 2 | **49.5** | 27.1 |
| 4 | **51.7** (12.9/stream) | 36.8 |
| 8 | 47.3 | **55.1** |
| 12 | 48.8 | **55.0** |
**Correction:** our earlier "DFlash verify saturates at c1" was an artifact of
`max_running_requests=2` β with headroom, DFlash scales and dominates to ~c8.
~55 aggregate is the machine's bandwidth plateau either way. **Production config
changed to DFlash + max_running 8** (best latency at every c<8, ~94% of the fleet
ceiling at c12).
## fp8-KV tilelang CUDA port (2026-08-28) β WORKS ON GB10
The gate was plumbing, not kernels: a complete raw-fp8 sparse kernel (`sparse_mla_fwd_decode_partial_fp8`)
ships in-tree, HIP-gated in 3 plumbing sites (dispatch, pool layout, fused-quant write). Port = relax the
gate (CUDA + both-backends-tilelang + SM89+), route the raw 512 B/token layout on CUDA, key the raw write
on layout not platform. Kernel verified by execution on sm_121 first (one-hot exact 0.000000; negative
control fails at 91%; fp8 GEMMs lower and run).
| check | result |
|---|---|
| boot-gate negative control (mixed backends + fp8) | **died as required** (ValueError) |
| point-of-effect log line | present (stock code cannot print it) |
| 19Γ21 gate / token-0 collapse | PASS / none |
| decode (n=5 medians) | 29.2 code / 22.9 prose β **parity** with same-image bf16 (29.8) |
| temp-0 vs bf16, 5 prompts | **4/5 exact** |
| 32K-depth codeword probe | PASS |
| TTFT@16k warm | **6.6 s vs 7.9 s bf16 β 17% faster prefill** |
| open item | raw-layout accounting resolves max_running_requests to 1; retune attempt (max-total 393216 + mamba-ratio 3) held at 1 and shrank the pool β the constraint is the mamba slice, not the token budget. Documented as PR known-limitation; fp8 = interactive config, bf16-T8 = fleet config |
Kernel analysis + port: this rig (Fable agent), validated per the house protocol.
## Multi-stream unlock (2026-08-28 pm) β the "1-stream fp8" open item is CLOSED, and it was never about fp8
The engine logs the whole story at boot:
`max_running_requests is capped to 2 by the mamba state cache (max_mamba_cache_size=10, 5 state slots per request)`.
Every DFlash config on this rig β bf16 T8 included, since it shipped identical memory args β was
silently capped at **2 concurrent spec streams**; the flat c2βc12 aggregate we called a "bandwidth
plateau" was queueing. The ~55 tok/s no-spec ceiling was the cap's shadow, not the machine's.
Fix (config FP8T8b, one coherent package): `--max-mamba-cache-size 40` (8 streams x 5 slots) +
`--mamba-ssm-dtype bfloat16` (halves state: 35 MB/slot vs 75) + `--mem-fraction-static 0.90`.
A first attempt (48 slots, fp32 ssm, 0.88) died at pool allocation β weights leave only ~2.3 GB of
fraction slack, so the ssm dtype lever is what makes 40 slots fit. KV pool grew to 84,288 tokens as
a side effect; 10.8 GB runtime headroom kept deliberately (GB10 OOM can wedge the node).
| c | FP8T8b (fp8-KV, DFlash, 40 slots) | old T8 (bf16, capped@2) | no-spec T12 |
|---|---|---|---|
| 1 | 36.9 | 37.3 | 14.5 |
| 2 | 48.2 | 49.5 | 27.1 |
| 4 | 46.9 | 51.7 | 36.8 |
| 8 | **80.3** (10.0/stream, 8/8) | 47.3 | 55.1 |
| 12 | **85.4** (7.1/stream, 12/12) | 48.8 | 55.0 |
+70% at c8, +75% at c12 over the capped curve; +55% over the no-spec "plateau". Decode batches
observed at 5-8 running requests β first true multi-stream DFlash on this rig. Trade-off, measured:
single-stream code median 27.4 vs 29.2 on the 10-slot fp8 config (~6%, consistent with bf16 ssm
states shaving accept length; deep-batch accept len ~3.1 vs ~4.2 single). Correctness: 19x21 gate
PASS, no token-0 collapse. **Production is now FP8T8b** β fp8-KV prefill wins AND the concurrency
curve, one config. (c-sweep prompts: 12 distinct short code/infra prompts, 400 max_tokens, temp 0,
stream:false, warmed; clocks stock.)
## Vision unlock (2026-08-28 pm) β GLM-5.3-Flash is MULTIMODAL, and it works on this stack
Correction of our own record: GLM-5.3-Flash has a full 24-layer vision tower with image AND
video tokens (upstream config: Glm5NextForConditionalGeneration). The LibertAIDAI NVFP4 quant
ships all 347 model.visual.* tensors, and the SGLang #36507 branch implements the vision path.
Our deployments simply never passed --enable-multimodal.
Config FP8T8V = FP8T8b + `--enable-multimodal`. Results:
- vision gate: PASS (64x64 solid-red data-URL probe answered "Red", temp 0)
- single-stream: 29.3 code / 23.8 prose β no measurable vision-tower tax (matches best fp8)
- c-sweep: c1 34.5 / c2 51.1 / c4 44.5 / c8 78.7 (8/8) / c12 79.5 β concurrency intact
(c12 delta vs FP8T8b's 85.4 is single-run noise territory; c4 dip reproduces in both)
**Production is now FP8T8V: fp8-KV + 8 concurrent streams + image input, one config.**
Not yet measured: vision quality beyond the smoke probe, video input, vision+DFlash accept
interaction, vision under concurrency. Treat image support as verified-working, not benchmarked.
## Pool expansion (2026-08-28 pm) β FP8T8X, production
FP8T8V + mem-fraction 0.90->0.92, single variable. KV pool 84,288 -> 244,032 tokens (2.9x;
the 262,144 ask minus page rounding), both target-fp8 and draft-bf16 pools. 8.9GB runtime
headroom kept. Gate PASS; c8 78.1 / c12 83.5 (curve unchanged); single-stream 27.0 code β
within the day's 27.0-29.3 noise band for this config family. Production = start-FP8T8X.sh:
fp8-KV + 8 streams + vision + 244k-token pool. Context-length raise beyond 131k = next rung
(needs rope/prefill verification, not just pool).
## Analyst pass (2026-08-28 pm) β c4 dip explained, D=7 rejected, one retraction
- The "c4 dip" is a HARNESS ARTIFACT, not an engine effect: engine decode is monotone in batch
(bs1 ~37-41 -> bs8 ~97-100 tok/s from decode-batch logs; cuda graph active at bs=4, no batch
split, mamba usage 0.30). The sweep harness takes wall=max over c different prompts and one
low-acceptance straggler (p2/p3 in the prompt list) sets the wall; c8 amortizes the same tail
over 2x the tokens (c4 1600 tok in 37.0s vs c8 3200 in 41.0s β impossible unless a shared
straggler). Harness now prints per-stream walls + straggler id. Low-c aggregates in earlier
sweeps understate the engine; c8/c12 figures stand (~80-96% homogeneous-model fit).
- RETRACTION: "D=5, max 6, realize ~5.9" was wrong β the arg counts the bonus token, ceiling
is 5, measured accept 3.4-4.4. Their k=7 = our D=8. Corrected in RESULTS.md.
- D=7 experiment REJECTED by arithmetic before spending a boot: predicted +1.4% structured
(ceiling +8% at perfect acceptance), -9% prose. The real gap is verify cost per token (~56%
cheaper on the EXL3 stack at identical step rates) -> next frontier is the decoupled drafter
attacking the ~40 ms fixed step floor.
## Straggler probe result (2026-08-28 pm) β mechanism REVISED by its own negative control
Per-prompt c1 probe: all 8 harness prompts land 10.3-14.0s (within +-20%) β the "slow prompt"
hypothesis is refuted (the predicted regex straggler was fastest). Revised mechanism, consistent
with all walls including the EXL3 lane's per-stream data (c4 min 7.9s ~= c1, max 18.7s):
**admission serialization** β chunked prefills enter one at a time, so wall = last-admitted
stream's start delay + decode; at c4 the fixed ramp amortizes over half the tokens of c8, which
reads as a "dip". Unchanged conclusions: the engine's decode is monotone in batch size, and
low-concurrency wall-clock aggregates understate steady-state engine throughput on BOTH stacks.
Structured re-measure same session: 48.8-49.7 tok/s (x3) vs 43.3 earlier β run-to-run band.
## Night queues 2+3 (2026-08-28 evening) β four rungs, one discovery, one blocker
| rung | change | verdict |
|---|---|---|
| N1 | num-continuous-decode-steps 3 | NULL β exact baseline; scheduler loop is not the 40ms floor |
| N2 | torch.compile | BOOT-FAIL β cuda-graph capture OOM at fraction 0.92; retry would need capped graph set + lower fraction, plus a compile tax on every boot. Not production-shaped |
| N3 | context 262k + fraction 0.93 | Speed-neutral (28.7 code, 80.2 c12), pool fully funded at 262,144 β BUT the 64k depth probe KILLED the worker (silent rank1 death mid-prefill, no traceback, docker state incoherent). **SCOPE CORRECTED 2026-08-29: this death is specific to THIS memory-tight config (bigger pool = less runtime headroom), NOT a general long-prefill failure β production (131k ctx / fraction 0.92) completes 86k-token prompts in 112 s. See the long-context section below.** NOT promoted |
| N4 | enable-mixed-chunk | NULL on this protocol β short-prompt sweep cannot exercise it; needs a prefill-interference test before final verdict |
Conclusion of the verify-cost hunt: both cheap levers dead; the code-speed gap vs EXL3 is
quant-level. Next real levers: decoupled drafter (architecture), and the 64k crash bisect
(reliability before capacity). Production remains FP8T8X.
## Long-context: what is actually true (2026-08-29, corrected)
Two earlier claims of ours were WRONG and are retracted here; the corrections cost one night and
are the reason this section exists.
- **RETRACTED "prefill >40k kills the worker."** That 600 s / timeout data point was the FIRST
request after a boot β TileLang JIT compilation for new shapes. The same config later did 86k
tokens in 112 s. Cold-start contamination, in a project whose own methodology notes warn about
exactly this.
- **RETRACTED "the worker dies at ~62k."** Scope error: it dies on the memory-tight N3 config
(ctx 262k, fraction 0.93). Production presents the same pressure as a timeout, not a kill.
- **RETRACTED our own first probe ladder.** All probes reused one filler string, so later ones
were served from the radix cache (prefill rate climbed 494 β 2,914 tok/s across the ladder β
the tell). Re-run with unique content per probe and `POST /flush_cache` before each.
Controlled A/B, unique uncached prompts, single variable, same night, same rig:
| prompt | chunked-prefill 8192 (production) | chunked-prefill 2048 |
|---|---|---|
| 86,444 tok | PASS 111.6 s | PASS 100.3 s |
| 102,644 tok | **client timeout at 300 s** (worker survived; node ssh unresponsive during) | **PASS 65.5 s** |
**Guidance: use `--chunked-prefill-size 2048` for long-context serving on GB10.** It is at least as
fast at every size measured and is the difference between completing and not completing above
~100k tokens. Mechanism (dense `[Ξ£q Γ Ξ£k/4]` fp32 indexer logits scaling with TOTAL prompt length,
with `_should_chunk_mqa_logits` defined but never called on the kpool path) and our public
correction: sglang issue #36941.
## Chunked-prefill size sweep (2026-08-29) β new production config
Single variable, everything else held at production settings. 102k probe = unique uncached
prompt, JIT pre-warmed, `POST /flush_cache` first.
| chunk | code | prose | c8 | c12 | 102,644-token prompt |
|---|---:|---:|---:|---:|---|
| 8192 (old prod) | 27.0-29.3 | 23.8 | 78.1 | 83.5 | **timeout at 300 s (reproduced twice)** |
| **4096 (NEW PROD)** | **28.6** | 23.6 | **77.4** | **83.2** | **PASS 104.0 s** |
| 2048 | 29.3 | 23.3 | 73.5 (~6% below band) | 82.4 | PASS 65.5 s |
**Production = `start-LC4.sh`** (chunk 4096): keeps the concurrency band AND completes 100k+
prompts. Promotion gated on 19x21 + a 44-request correctness matrix under load: 44/44 clean
(c1 8/8, c4 12/12, c8 24/24). 2048 is the right choice only if long-context latency matters
more than 8-way throughput.
### Mechanism: CONFIRMED (after one failed attempt)
Three builds, same 102,644-token unique prompt, chunk 8192, JIT pre-warmed, within one hour:
| build | result |
|---|---|
| stock | timeout at 300 s (x2) |
| stock + guard wired into kpool path (live-verified via container class MRO) | timeout at 300 s |
| same + logits chunked unconditionally at 512 rows (diagnostic build) | **PASS 135.0 s** |
So the dense `[Ξ£q Γ Ξ£k/4]` fp32 logits block IS the cause β chunking it, and nothing else, fixes
the case. The middle row is explained by the guard's own log line, which the diagnostic build
emitted 143 times: `num_q=4340 num_k=25664 guard_said_chunk=True budget=217873612`. The guard
correctly says "chunk" but its ~218 MB budget yields ~2,122-row chunks (25,664 x 4 bytes/row),
which is still too coarse here; 512-row chunks work. **Two defects, not one: the guard is never
called on the kpool path, AND its default budget is ~4x too generous for GB10 unified memory.**
Reported upstream on #36941. Note the first version of this section said the mechanism was NOT
established β that was written between the second and third builds and is kept in git history.
|