sync: latest measured numbers
Browse files
LADDER.md
CHANGED
|
@@ -42,5 +42,11 @@ tilelang forbids it on CUDA by upstream policy; trtllm kernels don't build for t
|
|
| 42 |
| V4b | upstream aa8c950a3 (official mHC capture fix + kpool changes) + our tiles | 29.5 (29.5β29.6) | **22.5** | **KEPT as incumbent** β code statistically tied with D6 (29.7, whose spread contains V4b), prose +4.7%, and it replaces our hand-guard with the official fix. Tie broken on provenance |
|
| 43 |
| V5b | novel fat tile 64/1/256 | β | β | smem 104,448 B > 101,376 β **3,072 bytes over.** Single-stage halved the request exactly as modeled |
|
| 44 |
|
| 45 |
-
|
| 46 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
| V4b | upstream aa8c950a3 (official mHC capture fix + kpool changes) + our tiles | 29.5 (29.5β29.6) | **22.5** | **KEPT as incumbent** β code statistically tied with D6 (29.7, whose spread contains V4b), prose +4.7%, and it replaces our hand-guard with the official fix. Tie broken on provenance |
|
| 43 |
| V5b | novel fat tile 64/1/256 | β | β | smem 104,448 B > 101,376 β **3,072 bytes over.** Single-stage halved the request exactly as modeled |
|
| 44 |
|
| 45 |
+
| D5 | draft tokens 5 | 29.6 (29.6β29.6) | **23.3** (23.0β23.3) | **best combined config** β D6's code with D4's prose; takes incumbency on tie-break |
|
| 46 |
+
| V5c | fat tile 64/1/128 | β | β | smem 103,424 B β still 2 KB over. **Fat-tile chapter closed with a complete map**: 64-wide needs β₯103.4 KB in any shape; GB10 ceiling 101.4; 32-wide fits and costs nothing |
|
| 47 |
+
|
| 48 |
+
**Decoupled drafter** (killing the TP=2 all-reduce tax on the 1B draft model): mapped, viable, NOT attempted β
|
| 49 |
+
requires a separate drafter-server topology (`--decoupled-spec-role verifier/drafter-rank` + bind/connect
|
| 50 |
+
endpoints + rank). This is the #1 next-frontier item; expected the largest structural gain.
|
| 51 |
+
|
| 52 |
+
Finishing: FINAL (v4b image + D5 flags) β fa3 draft attention β TTFT/concurrency β G6 matrix β closeout.
|