randomllama commited on
Commit
2f5385e
·
verified ·
1 Parent(s): 79122b8

sync: latest measured numbers

Browse files
Files changed (1) hide show
  1. LADDER.md +10 -0
LADDER.md CHANGED
@@ -156,3 +156,13 @@ fp8-KV + 8 streams + vision + 244k-token pool. Context-length raise beyond 131k
156
  (ceiling +8% at perfect acceptance), -9% prose. The real gap is verify cost per token (~56%
157
  cheaper on the EXL3 stack at identical step rates) -> next frontier is the decoupled drafter
158
  attacking the ~40 ms fixed step floor.
 
 
 
 
 
 
 
 
 
 
 
156
  (ceiling +8% at perfect acceptance), -9% prose. The real gap is verify cost per token (~56%
157
  cheaper on the EXL3 stack at identical step rates) -> next frontier is the decoupled drafter
158
  attacking the ~40 ms fixed step floor.
159
+
160
+ ## Straggler probe result (2026-08-28 pm) — mechanism REVISED by its own negative control
161
+ Per-prompt c1 probe: all 8 harness prompts land 10.3-14.0s (within +-20%) — the "slow prompt"
162
+ hypothesis is refuted (the predicted regex straggler was fastest). Revised mechanism, consistent
163
+ with all walls including the EXL3 lane's per-stream data (c4 min 7.9s ~= c1, max 18.7s):
164
+ **admission serialization** — chunked prefills enter one at a time, so wall = last-admitted
165
+ stream's start delay + decode; at c4 the fixed ramp amortizes over half the tokens of c8, which
166
+ reads as a "dip". Unchanged conclusions: the engine's decode is monotone in batch size, and
167
+ low-concurrency wall-clock aggregates understate steady-state engine throughput on BOTH stacks.
168
+ Structured re-measure same session: 48.8-49.7 tok/s (x3) vs 43.3 earlier — run-to-run band.