File size: 4,402 Bytes
00d52e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
269cd26
 
 
 
 
 
 
 
 
 
344991c
 
 
c84a0db
 
9eae477
 
 
 
 
 
a486ba9
 
 
16cff2a
 
 
50d8e1e
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
# Optimization ladder β€” GLM-5.3-Flash + DFlash2, 2Γ— DGX Spark (GB10), SGLang TP=2

All numbers: warmed, temp 0, stream:false, 800 max_tokens, n=5 medians, stock clocks,
code-1/prose-1 prompts as in RESULTS.md. Keep rule: beats incumbent code median with the
19Γ—21 gate + token-0 collapse scan passing.

## Round 1 (autonomous, night of 2026-08-27β†’28)
| rung | change | code | prose | verdict |
|---|---|---|---|---|
| first light | β€” | 27.6 | 20.7 | baseline |
| L2 | flashinfer autotune ON (+`--enable-metrics`) | **28.4** (26.9–29.8) | 20.6 | **KEPT** |
| L3 | + fp8 draft KV | 27.0 (23.9–28.4) | 20.7 | reverted β€” convert overhead exceeds bandwidth saved on a ~2 GB drafter cache |
| L4 | + fp8 target KV | β€” | β€” | **BOOT FAILED in 64 s**; failure logs lost to teardown (ladder now saves them); autopsy queued round 2 |
| L5 | ctx 131072, max-total 262144 | 28.0 (27.6–30.8) | **21.3** | reverted by keep-rule β€” **but note: 2Γ— context for βˆ’1.4% code speed; worth keeping for serving. Operator's call. Confound: carried L3's fp8-draft-KV flag** |

Round-1 net: +3% code. The big lever (fp8 target KV) is blocked, not disproven.

## Round 2
| rung | change | code | prose | verdict |
|---|---|---|---|---|
| L1v2 | decode+v1 kernels stock tiles | β€” | β€” | BOOT FAILED, same 169,984 B smem error β€” `sparse_attention_fwd_kernel_v1` is in the verify path and must stay small |
| L1v3 | only `sparse_mla_fwd_decode_partial` stock | 28.5 (25.7–33.2) | 20.2 | kept (wash) β€” **finding: the GB10 small tile is ~free; tile geometry is not the bottleneck** |

Kernel tile map for sm_121 (measured): `qo_len` multi-token kernel β†’ small tile REQUIRED;
`sparse_attention_fwd_kernel_v1` β†’ small tile REQUIRED (verify path); `sparse_mla_fwd_decode_partial`
β†’ either (no measurable difference).

| D6 | speculative-num-draft-tokens 8β†’6 | **29.7** (27.2–31.0) | **21.5** (20.4–23.4) | **KEPT β€” new incumbent.** Shorter draft block wastes less verify on doomed tokens at our acceptance profile |

Cumulative: 27.6 β†’ 29.7 code (+7.6%) since first light.
| L4v2 | fp8 target KV (tilelang DSA) | β€” | β€” | **CLOSED: architecturally unsupported** β€” SGLang raises `tilelang DSA ... on CUDA requires a bfloat16 KV cache` at arg resolution. Round-1's 64s death explained |

| L4v3 | trtllm DSA backends + fp8 KV | β€” | β€” | **DIED: `TllmGenFmhaRunner: Unsupported arch`** β€” trtllm FMHA does not support sm_121 |

**fp8 target KV verdict on GB10 + SGLang + GLM-5.3: impossible on this stack today.**
tilelang forbids it on CUDA by upstream policy; trtllm kernels don't build for the chip.
(vLLM reached fp8 KV on GB10 only via hand-patched CTA tile caps β€” see tonyd2wild's recipe.)

| D4 | draft tokens 4 | 28.3 (28.2–28.7) | **23.3** (23.3–23.3) | reverted on code β€” but best prose of the campaign (+8% vs D6): code peaks at D=6, prose at D=4; D=5 queued |
| V4/V5 (first attempt) | upstream aa8c950a3 refresh + novel fat tile | β€” | β€” | both died pre-kernel on a partial-overlay skew (`hc_attn_to_mlp` β€” model file and communicator_mhc.py are a coupled pair); rebuilt as V4b/V5b with both files |

| V4b | upstream aa8c950a3 (official mHC capture fix + kpool changes) + our tiles | 29.5 (29.5–29.6) | **22.5** | **KEPT as incumbent** β€” code statistically tied with D6 (29.7, whose spread contains V4b), prose +4.7%, and it replaces our hand-guard with the official fix. Tie broken on provenance |
| V5b | novel fat tile 64/1/256 | β€” | β€” | smem 104,448 B > 101,376 β€” **3,072 bytes over.** Single-stage halved the request exactly as modeled |

| D5 | draft tokens 5 | 29.6 (29.6–29.6) | **23.3** (23.0–23.3) | **best combined config** β€” D6's code with D4's prose; takes incumbency on tie-break |
| V5c | fat tile 64/1/128 | β€” | β€” | smem 103,424 B β€” still 2 KB over. **Fat-tile chapter closed with a complete map**: 64-wide needs β‰₯103.4 KB in any shape; GB10 ceiling 101.4; 32-wide fits and costs nothing |

**Decoupled drafter** (killing the TP=2 all-reduce tax on the 1B draft model): mapped, viable, NOT attempted β€”
requires a separate drafter-server topology (`--decoupled-spec-role verifier/drafter-rank` + bind/connect
endpoints + rank). This is the #1 next-frontier item; expected the largest structural gain.

Finishing: FINAL (v4b image + D5 flags) β†’ fa3 draft attention β†’ TTFT/concurrency β†’ G6 matrix β†’ closeout.