sync: latest measured numbers
Browse files
LADDER.md
CHANGED
|
@@ -39,4 +39,8 @@ tilelang forbids it on CUDA by upstream policy; trtllm kernels don't build for t
|
|
| 39 |
| D4 | draft tokens 4 | 28.3 (28.2β28.7) | **23.3** (23.3β23.3) | reverted on code β but best prose of the campaign (+8% vs D6): code peaks at D=6, prose at D=4; D=5 queued |
|
| 40 |
| V4/V5 (first attempt) | upstream aa8c950a3 refresh + novel fat tile | β | β | both died pre-kernel on a partial-overlay skew (`hc_attn_to_mlp` β model file and communicator_mhc.py are a coupled pair); rebuilt as V4b/V5b with both files |
|
| 41 |
|
| 42 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
| D4 | draft tokens 4 | 28.3 (28.2β28.7) | **23.3** (23.3β23.3) | reverted on code β but best prose of the campaign (+8% vs D6): code peaks at D=6, prose at D=4; D=5 queued |
|
| 40 |
| V4/V5 (first attempt) | upstream aa8c950a3 refresh + novel fat tile | β | β | both died pre-kernel on a partial-overlay skew (`hc_attn_to_mlp` β model file and communicator_mhc.py are a coupled pair); rebuilt as V4b/V5b with both files |
|
| 41 |
|
| 42 |
+
| V4b | upstream aa8c950a3 (official mHC capture fix + kpool changes) + our tiles | 29.5 (29.5β29.6) | **22.5** | **KEPT as incumbent** β code statistically tied with D6 (29.7, whose spread contains V4b), prose +4.7%, and it replaces our hand-guard with the official fix. Tie broken on provenance |
|
| 43 |
+
| V5b | novel fat tile 64/1/256 | β | β | smem 104,448 B > 101,376 β **3,072 bytes over.** Single-stage halved the request exactly as modeled |
|
| 44 |
+
|
| 45 |
+
block_I must divide topk (2048): only 32 or 64 are legal. Final tile candidate V5c = 64/1/128
|
| 46 |
+
(reduction workspace shrinks with threads) queued. In progress: D5 β V5c β decoupled drafter, fa3, prefill/TTFT, closeout.
|