GLM-5.2-EXL3-TR3-3.0bpw / RELEASE_TEST_SUITE.md
brandonmusic's picture
RELEASE_TEST_SUITE.md: pin v31 (base refresh sic3828fd; wheel superset-verified)
9297b9f verified
|
Raw History Blame Contribute Delete
144 kB

Release Test Suite โ€” GLM-5.2-EXL3-TR3-3.0bpw

Measured 2026-07-25 on 4x RTX PRO 6000 Blackwell 96 GB (TP4/DCP4, MTP-3 greedy) with the published sha-pinned image and the shipped server.sh / docker-compose.yml preset, using llm_decode_bench.py with the same flags as run_release_benchmarks.sh. Every raw benchmark log is embedded verbatim beneath its table.

Section 8 reports an independent third-party evaluation run by malaiwah/glm52-exl3-vast.


2026-07-25 update (3) โ€” per-model runtime scoping + per-ordinal arch cache key (v26)

Follow-up to the v20 rebase, from CodeRabbit review on local-inference-lab/vllm#139.

Bug: the EXL3 rank-sliced runtime cache was keyed only on device, dtype, shape, topk and planner settings. The cached entry owns mutable Trellis/prefill scratch and parity staging buffers. A target MoE layer and the rank-sliced MTP-78 draft layer match on every one of those components -- same hidden/intermediate size, same local expert count, same topk, same planner env, and both resolve max_num_batched_tokens from the same scheduler config -- so the draft was reusing the target's scratch. That defeats the target/draft isolation their independently captured CUDA graphs rely on.

Fix: the cache key is now scoped to the owning quant config, so each model gets exactly one runtime. This is deliberately coarser than per-layer: the prefill arena is ~1054 MiB, so per-layer runtimes would need tens of GiB per rank across 75+ layers.

Runtime: verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a (sha256:8753406fโ€ฆ).

Measured effect (the fix is observable in memory, not in the log -- the planner line uses info_once and is deduplicated):

Rank-sliced runtimes GPU KV cache Concurrency @524K
pre-fix (shared scratch) 1 (target+draft collide) 1,115,904 tokens 2.13x
v26 (scoped) 2 (isolated) 998,656 tokens 1.90x

The 119,552-token KV reduction corresponds to ~1056 MiB, matching the 1054.2 MiB arena -- direct evidence that a second runtime is allocated. Lower KV capacity is the intended cost of the isolation.

Second fix in this runtime (v26). CodeRabbit correctly rejected a first attempt at the compile-cache device key: torch.cuda.get_device_capability() / get_device_name() already resolve against torch.cuda.current_device(), so passing that ordinal explicitly was a no-op and left the process-wide key free to freeze whichever GPU was current on the first call. The real fix memoizes the architecture key per device ordinal and threads the ordinal through _static_compile_cache_context, which is lru_cached on the compile callable and would otherwise have re-frozen the identity at that layer. The returned key still omits the ordinal, so GPUs of the same architecture keep sharing compiled artifacts. This matters on this rig specifically: its four boards report two different device names (Max-Q and non-Max-Q) while sharing compute capability 12.0.

Quality on v26:

Suite Config Result
Estonia c2, 5 runs 5/5 pass, 0 fail, correct rate 1.00
LAVD c5, 5 runs 3 EXACT / 2 NEAR / 0 FAIL, correct rate 1.00

For continuity: the pre-fix runtime measured LAVD 2E/3N/0F. An intermediate scoped build measured 2E/2N/1F on one 5-run sample and 2E/3N/0F on a second; that single failure was a wrong ledger total at 7,114 completion tokens against a 24,576 cap, i.e. an answer-quality miss rather than truncation, and within this profile's run-to-run spread. This runtime shows no failures on either suite.


2026-07-25 update (2) โ€” rebased onto the FINAL Gilded Gnosis v20 base

The runtime is rebased onto the finalized v20 common base voipmonitor/vllm:gilded-gnosis-v20-vllm0c79e41-sic3828fd-fi801d57a-cu132-20260727 (vLLM 5517197, Sparkinfer be0edca, FlashInfer 801d57a, CUTLASS 4.6.0).

New runtime: verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a (sha256:da185fe8โ€ฆ).

Why: the finalized base consolidates the DCP prefill auto-policy and the corrected workspace accounting, resolving the >8k DCP prefill collapse present in earlier v20 candidates, plus long-context MTP alignment and deterministic dynamic-MoE output. EXL3 is enabled by rebasing the EXL3 source layer onto that pinned stack (the base ships no EXL3/Trellis loader), keeping EXL3 quantization separate while sharing the corrected runtime. Of the 12 runtime overlay files only models/deepseek_v2.py and v1/attention/backends/mla/indexer.py differ in the new base; both were re-derived from the v20-final versions with the EXL3 edits replayed on top, so v20's DCP/indexer work is preserved rather than overwritten by the overlay.

Config change: server.sh / docker-compose.yml now set the DCP policy that the base launcher resolves from DCP_*=auto for TP4/DCP4 โ€” query split 1, full-CKV gather 1, top-k owner merge 1, indexer shards 0, CKV prefetch depth 1, prefetch workspace 1024 MiB, and DCP_PREFILL_WORKSPACE=1 (VLLM_DCP_PROJECT_BEFORE_MERGE=1 + VLLM_B12X_MLA_DCP_GATHER_IN_WORKSPACE=1). These are set explicitly because the preset calls vllm serve directly and bypasses /usr/local/bin/serve-gilded-gnosis.sh. Note VLLM_DCP_QUERY_SPLIT moves from 0 to 1 versus the previous preset.

Boot assertions observed: engine v0.11.2.dev280+gilded.gnosis.v20.vllm5517197.sibe0edca.fi801d57a.cu132.20260725, vLLM is using nccl==2.30.4, EXL3 rank-sliced runtime planned: Trellis m=1..32 block_m=8, prefill trellis block_m=64 arena=1054.2MiB capacity=3072 chunk=128 topk=8, Preallocated 30.8 MiB for 2 persistent CKV execution lane(s), Using native CKV layer prefetch with depth=1 and 2 workspace slots, GPU KV cache 1,115,904 tokens (2.13x at 524,288). The base's InstantTensor loader line does not appear on this path because the EXL3 checkpoint loads through the EXL3 rank-sliced loader.

Regression (no degradation):

Suite Config Result
Estonia c2, 5 runs 5/5 pass, 0 fail, correct rate 1.00, 70.0 tok/s aggregate
LAVD c5, 5 runs 2 EXACT / 3 NEAR / 0 FAIL, correct rate 1.00, 55.6 tok/s aggregate

All sections below were measured on the previous (v21-mtp78tr3) image and remain valid for the checkpoint itself; only the runtime base changed.


2026-07-25 update โ€” MTP layer-78 is now EXL3 tr3

This checkpoint now ships the MTP (layer 78) routed experts in EXL3 Trellis tr3 (3.0 bpw), matching layers 3-77; the previous BF16 MTP head is retired. The layer-78 file drops from 19.9 GB to 4.24 GB (โˆ’15.66 GB), freeing ~3.9 GiB/rank.

Requires the updated runtime: verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a (sha256:9b1befc1โ€ฆ) plus VLLM_EXL3_TRELLIS_MIN_M=1 (the compose / server.sh default in this repo). The prior v20 image cannot load a tr3 MTP layer. The two loader fixes are in vLLM PR #139 (local-inference-lab/vllm#139); no Sparkinfer change was needed (validated against #49 / si1a88b38).

Re-measured on the same 4ร— RTX PRO 6000 (TP4/DCP4, MTP-3, util 0.96, VLLM_EXL3_TRELLIS_MIN_M=1, auto-profiled KV):

Metric BF16-MTP build (Sections below) tr3-MTP build (this)
GPU KV cache @ 0.96 util 680K tok (1.3ร— @ 524K) 1,132,544 tok (2.16ร— @ 524K)
Prefill 8k / 64k / 128k (tok/s) 2,551 / โ€” / 1,833 2,521 / 1,916 / 1,765
Decode C1 / C4 / C8 (tok/s) 87.5 / 219.3 / 308.1 89.7 / 225.3 / 293.5
Estonia (long-ctx retrieval) PASS 30/30 PASS 10/10
LAVD (ledger consistency) 18E / 11N / 1F EXACT 5 / NEAR 5 / FAIL 0

Decode is neutral (MTP is lossless); the real gain is **+66% KV-cache / concurrency headroom** from the freed VRAM. Accuracy is unchanged โ€” the detailed Sections 1-8 below were measured on the prior BF16-MTP build and remain representative for quality; only the image, the layer-78 format, and the KV/serving preset changed.


What this quantization costs

BF16 full precision 1,506 GB
EXL3 TR3 3.0bpw 316.5 GB (tr3-MTP build; was 332.2 GB with the BF16 MTP head)
Size vs BF16 21.0% (a 79.0% reduction)
Effective whole-model rate ~3.45 bpw (routed experts, now including MTP layer-78, are a flat 3.0; attention, shared experts, embeddings, LM head, and the non-expert parts of layer 78 stay BF16)

BF16 weights alone would need 16x 96 GB cards. This fits on 4 with room for the KV cache.

Image

verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff

This is the current runtime. The per-benchmark sections further down were measured on the earlier v21-mtp78tr3 image and are retained as the checkpoint-level record; see the dated update sections above for what changed in the runtime since, and for the re-verification run on this image.

Serving preset (as shipped)

Setting Value
GPU_MEMORY_UTILIZATION 0.96
MAX_MODEL_LEN 524288 (tr3-MTP build; was 262144)
NUM_GPU_BLOCKS_OVERRIDE empty โ†’ auto-profile (~1,132,544 KV tokens @ 0.96; was pinned 1024)
VLLM_EXL3_TRELLIS_MIN_M 1 (required for the tr3 MTP draft's small-m GEMMs; was 4)
MAX_NUM_BATCHED_TOKENS 3072
MAX_NUM_SEQS 8
MTP enabled, 3 tokens, greedy draft
ENABLE_ASYNC_SCHEDULING 0 (correctness guard)
VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE 0 (lossless setting)
Attention / MoE B12X_MLA_SPARSE / b12x
Quantization exl3
KV cache nvfp4_ds_mla (shipped default) and fp8 (comparison arm)

Only --kv-cache-dtype and (for section 3) --default-chat-template-kwargs were parameterized; every other flag is byte-identical to the published compose file.


Summary

Test Runs nvfp4_ds_mla fp8
Estonia (needle retrieval, 133K ctx) 30 @ c2 PASS 30 / FAIL 0 PASS 30 / FAIL 0
LAVD (ledger consistency) 30 @ c5 18 EXACT / 11 NEAR / 1 FAIL (97%) 15 EXACT / 13 NEAR / 2 FAIL (93%)
Hotel-lights, low tier 30 @ c5 15 EXACT / 15 FAIL (50%) 18 EXACT / 12 FAIL (60%)
Hotel-lights, Max tier 30 @ c5 20 EXACT / 10 FAIL (67%) 20 EXACT / 10 FAIL (67%)
KLD vs BF16 5 0.116138 0.101198
Decode C1 / C8 (tok/s) โ€” 87.5 / 308.1 86.6 / 312.2
Prefill 8k / 128k (tok/s) โ€” 2,551 / 1,833 2,660 / 1,771

1. Estonia โ€” long-context needle retrieval

133,186-token prompt. 30 runs, concurrency 2, temperature 0, repetition penalty 1.25, max_tokens 40000, regex-scored on the final answer line.

KV cache score completed hit max_tokens tok p50 avg latency TTFT gen tok/s
nvfp4_ds_mla PASS 30 / FAIL 0 30/30 0 2,250 38.0 s 0.61 s 67.9
fp8 PASS 30 / FAIL 0 30/30 0 2,370 51.8 s 0.60 s 71.2

100% on both KV formats, no run near the token cap. The repetition penalty matters here: without it the model loops on retrieval phrasing and exhausts the output budget without answering.

Estonia โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Completion Token Statistics Benchmark                                        โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:estonia                                                      โ”‚
โ”‚ Concurrency: 2                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 40000                                        โ”‚
โ”‚ Scoring: \bestonia\b                                                         โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Completion Token Statistics                                                  โ”‚
โ”‚ One optional prefix-cache scout request is used to populate prefill first.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=2.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                       
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                                     โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ estonia                                   โ”‚
โ”‚ prompt         โ”‚ profile:estonia                           โ”‚
โ”‚ prompt chars   โ”‚ 707,372                                   โ”‚
โ”‚ requested runs โ”‚ 30                                        โ”‚
โ”‚ concurrency    โ”‚ 2                                         โ”‚
โ”‚ max tokens     โ”‚ 40000                                     โ”‚
โ”‚ scoring        โ”‚ regex                                     โ”‚
โ”‚ prefill scout  โ”‚ 133,186 prompt tok / 79.72s = 1,671 tok/s โ”‚
โ”‚ correct regex  โ”‚ \bestonia\b                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,105 W | max 1,124 W | limit 1,200 W | over 10m 56s | 273 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚ pโ€ฆ โ”‚ done/โ€ฆ โ”‚     score โ”‚ staโ€ฆ โ”‚ outputโ€ฆ โ”‚ output โ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg reโ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  2 โ”‚  30/30 โ”‚ PASS 30 โ€ฆ โ”‚ โ˜…โ˜…โ˜…โ€ฆ โ”‚   2,250 โ”‚    3,687 โ”‚     67.9 โ”‚    38.0 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=2                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚            value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚            30/30 โ”‚
โ”‚ score                      โ”‚ PASS 30 / FAIL 0 โ”‚
โ”‚ stars                      โ”‚    โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜… ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                0 โ”‚
โ”‚ completion tokens avg      โ”‚            2,536 โ”‚
โ”‚ completion tokens p50      โ”‚            2,250 โ”‚
โ”‚ completion tokens p90      โ”‚            3,687 โ”‚
โ”‚ completion tokens p99      โ”‚            4,725 โ”‚
โ”‚ elapsed avg                โ”‚            38.0s โ”‚
โ”‚ TTFT avg                   โ”‚            0.61s โ”‚
โ”‚ aggregate gen tok/s        โ”‚             67.9 โ”‚
โ”‚ mean per-request gen tok/s โ”‚             67.9 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: completion-token p50/p90/p99 tells how many decode tokens the 
model needed to reach its final answer under this engine/config. Correctness is 
scored from the final non-empty answer line by default, matching the GLM 
dense-MLA vs NSA benchmark methodology. The prefill scout is not a scored 
answer; it is the max_tokens=1 prefix-cache warmup and its prompt_tokens/TTFT 
value is reported as scout prefill speed. Concurrency Results groups completed 
requests by parallelism; Completed Requests shows the latest individual finished
answers.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/estonia-nvfp4_ds_mla-c2-r30-rp125.json
Estonia โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Completion Token Statistics Benchmark                                        โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:estonia                                                      โ”‚
โ”‚ Concurrency: 2                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 40000                                        โ”‚
โ”‚ Scoring: \bestonia\b                                                         โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Completion Token Statistics                                                  โ”‚
โ”‚ One optional prefix-cache scout request is used to populate prefill first.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=2.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                       
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                                     โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ estonia                                   โ”‚
โ”‚ prompt         โ”‚ profile:estonia                           โ”‚
โ”‚ prompt chars   โ”‚ 707,372                                   โ”‚
โ”‚ requested runs โ”‚ 30                                        โ”‚
โ”‚ concurrency    โ”‚ 2                                         โ”‚
โ”‚ max tokens     โ”‚ 40000                                     โ”‚
โ”‚ scoring        โ”‚ regex                                     โ”‚
โ”‚ prefill scout  โ”‚ 133,186 prompt tok / 78.91s = 1,688 tok/s โ”‚
โ”‚ correct regex  โ”‚ \bestonia\b                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,101 W | max 1,127 W | limit 1,200 W | over 14m 28s | 361 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚ pโ€ฆ โ”‚ done/โ€ฆ โ”‚     score โ”‚ staโ€ฆ โ”‚ outputโ€ฆ โ”‚ output โ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg reโ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  2 โ”‚  30/30 โ”‚ PASS 30 โ€ฆ โ”‚ โ˜…โ˜…โ˜…โ€ฆ โ”‚   2,370 โ”‚    4,513 โ”‚     71.2 โ”‚    51.8 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=2                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚            value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚            30/30 โ”‚
โ”‚ score                      โ”‚ PASS 30 / FAIL 0 โ”‚
โ”‚ stars                      โ”‚    โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜… ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                0 โ”‚
โ”‚ completion tokens avg      โ”‚            3,644 โ”‚
โ”‚ completion tokens p50      โ”‚            2,370 โ”‚
โ”‚ completion tokens p90      โ”‚            4,513 โ”‚
โ”‚ completion tokens p99      โ”‚           26,158 โ”‚
โ”‚ elapsed avg                โ”‚            51.8s โ”‚
โ”‚ TTFT avg                   โ”‚            0.60s โ”‚
โ”‚ aggregate gen tok/s        โ”‚             71.2 โ”‚
โ”‚ mean per-request gen tok/s โ”‚             67.3 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: completion-token p50/p90/p99 tells how many decode tokens the 
model needed to reach its final answer under this engine/config. Correctness is 
scored from the final non-empty answer line by default, matching the GLM 
dense-MLA vs NSA benchmark methodology. The prefill scout is not a scored 
answer; it is the max_tokens=1 prefix-cache warmup and its prompt_tokens/TTFT 
value is reported as scout prefill speed. Concurrency Results groups completed 
requests by parallelism; Completed Requests shows the latest individual finished
answers.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/estonia-fp8-c2-r30-rp125.json

2. LAVD โ€” long-context ledger consistency

48,302-char structured ledger; the model must find human data-entry errors, apply the repair rule, and return ticket count and hours. Ground truth 72, 46.0. 30 runs, concurrency 5, temperature 0, repetition penalty 1.15, max_tokens 24576.

KV cache EXACT NEAR FAIL pass (E+N) hit max_tokens tok p50 avg latency gen tok/s
nvfp4_ds_mla 18 11 1 29/30 (97%) 0 9,206 180.6 s 53.3
fp8 15 13 2 28/30 (93%) 0 8,908 190.4 s 52.2

All three failures are single-axis near-misses, not parse errors or truncation: 65, 42.5 (count -7) ยท 72, 41.75 (count exact, hours -4.25) ยท 66, 45.25 (count -6). Each used fewer tokens than the ~9K median, i.e. they stopped searching early rather than looping.

LAVD โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LAVD Context Consistency Test                                                โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:lavd-test                                                    โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 24576                                        โ”‚
โ”‚ Scoring: EXACT / NEAR / FAIL numeric pair                                    โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LAVD Context Consistency Test                                                โ”‚
โ”‚ Arithmetic is intentionally simple; the test checks whether the model keeps  โ”‚
โ”‚ a long structured context consistent, finds human data-entry errors, applies โ”‚
โ”‚ the repair rule, and returns the final ticket count and hours. Built-in      โ”‚
โ”‚ profile run at fixed concurrency C=5.                                        โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                               
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value             โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ lavd-test         โ”‚
โ”‚ prompt         โ”‚ profile:lavd-test โ”‚
โ”‚ prompt chars   โ”‚ 48,302            โ”‚
โ”‚ requested runs โ”‚ 30                โ”‚
โ”‚ concurrency    โ”‚ 5                 โ”‚
โ”‚ max tokens     โ”‚ 24576             โ”‚
โ”‚ scoring        โ”‚ ledger_lavd       โ”‚
โ”‚ expected       โ”‚ 72, 46            โ”‚
โ”‚ prompt sha256  โ”‚ 5c83674d5f0fd2a7  โ”‚
โ”‚ dataset sha256 โ”‚ 612f8041bbca048c  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,130 W | max 1,143 W | limit 1,200 W | over 19m 20s | 483 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 18 / NEAR 11โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚   9,206 โ”‚ 11,799 โ”‚     53.3 โ”‚ 180.6 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 18 / NEAR 11 / FAIL 1 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜†โ˜†โ˜†โ˜† ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           0 โ”‚
โ”‚ completion tokens avg      โ”‚                       9,517 โ”‚
โ”‚ completion tokens p50      โ”‚                       9,206 โ”‚
โ”‚ completion tokens p90      โ”‚                      11,799 โ”‚
โ”‚ completion tokens p99      โ”‚                      14,634 โ”‚
โ”‚ elapsed avg                โ”‚                      180.6s โ”‚
โ”‚ TTFT avg                   โ”‚                       1.99s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        53.3 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        53.3 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                     
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 14 โ”‚ 5 โ”‚  5,721 โ”‚ FAIL  โ”‚ 65, 42.5 (count -7, hours -3.50) | 65, 42.5 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the parsed final numeric pair is exactly 72, 46.0. 
NEAR means both count and hours are within the configured tolerance; FAIL means 
the answer was unparseable or outside tolerance. The 10-slot quality bar is a 
rounded distribution: โ˜…=EXACT, โ˜†=NEAR, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/lavd-nvfp4_ds_mla-c5-r30-rp115.json
LAVD โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LAVD Context Consistency Test                                                โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:lavd-test                                                    โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 24576                                        โ”‚
โ”‚ Scoring: EXACT / NEAR / FAIL numeric pair                                    โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LAVD Context Consistency Test                                                โ”‚
โ”‚ Arithmetic is intentionally simple; the test checks whether the model keeps  โ”‚
โ”‚ a long structured context consistent, finds human data-entry errors, applies โ”‚
โ”‚ the repair rule, and returns the final ticket count and hours. Built-in      โ”‚
โ”‚ profile run at fixed concurrency C=5.                                        โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                               
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value             โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ lavd-test         โ”‚
โ”‚ prompt         โ”‚ profile:lavd-test โ”‚
โ”‚ prompt chars   โ”‚ 48,302            โ”‚
โ”‚ requested runs โ”‚ 30                โ”‚
โ”‚ concurrency    โ”‚ 5                 โ”‚
โ”‚ max tokens     โ”‚ 24576             โ”‚
โ”‚ scoring        โ”‚ ledger_lavd       โ”‚
โ”‚ expected       โ”‚ 72, 46            โ”‚
โ”‚ prompt sha256  โ”‚ 5c83674d5f0fd2a7  โ”‚
โ”‚ dataset sha256 โ”‚ 612f8041bbca048c  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,127 W | max 1,138 W | limit 1,200 W | over 19m 45s | 493 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 15 / NEAR 13โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚   8,908 โ”‚ 13,173 โ”‚     52.2 โ”‚ 190.4 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 15 / NEAR 13 / FAIL 2 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โ˜†โ˜†โ˜†โ˜†โœ• ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           0 โ”‚
โ”‚ completion tokens avg      โ”‚                       9,831 โ”‚
โ”‚ completion tokens p50      โ”‚                       8,908 โ”‚
โ”‚ completion tokens p90      โ”‚                      13,173 โ”‚
โ”‚ completion tokens p99      โ”‚                      14,658 โ”‚
โ”‚ elapsed avg                โ”‚                      190.4s โ”‚
โ”‚ TTFT avg                   โ”‚                       1.95s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        52.2 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        52.3 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                       
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                  โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 13 โ”‚ 5 โ”‚  9,982 โ”‚ FAIL  โ”‚ 72, 41.75 (count +0, hours -4.25) | 72, 41.75 โ”‚
โ”‚ 23 โ”‚ 5 โ”‚  7,131 โ”‚ FAIL  โ”‚ 66, 45.25 (count -6, hours -0.75) | 66, 45.25 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the parsed final numeric pair is exactly 72, 46.0. 
NEAR means both count and hours are within the configured tolerance; FAIL means 
the answer was unparseable or outside tolerance. The 10-slot quality bar is a 
rounded distribution: โ˜…=EXACT, โ˜†=NEAR, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/lavd-fp8-c5-r30-rp115.json

3. Hotel-lights โ€” reasoning, low tier vs Max tier

100-room light-cycling puzzle, expected answer 48. 30 runs each, concurrency 5, temperature 0, exact numeric scoring. Run at both reasoning tiers.

Tier KV cache EXACT FAIL pass hit max_tokens avg completion tok wall
low nvfp4_ds_mla 15 15 50% 2 18,903 39 min
low fp8 18 12 60% 3 20,806 44 min
Max nvfp4_ds_mla 20 10 67% 4 37,389 77 min
Max fp8 20 10 67% 6 38,346 81 min

Max tier buys +7 to +17 points for 2x the reasoning tokens and 2x the wall time. Both KV formats converge to the same 67% at Max. Some runs still exhaust even a 60,000-token budget, so this task can absorb unbounded reasoning.

Reasoning-tier gotcha, worth knowing before reproducing. This model's chat_template.jinja line 2 reads:

{%- set effective_reasoning_effort = 'high' if reasoning_effort is defined and reasoning_effort == 'high' else 'max' -%}

Only the literal string high selects the lower tier. Anything else โ€” including max, or omitting the field โ€” selects Max. The shipped compose sends --default-chat-template-kwargs '{"reasoning_effort":"high"}', so the default preset runs the lower tier. The Max rows above were produced by setting that server default to max; each run logged the resolved value from the live container to confirm the tier applied.

Hotel-lights low tier โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:hotel-lights                                                 โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 40000                                        โ”‚
โ”‚ Scoring: EXACT / FAIL final number                                           โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ A compact reasoning profile with a known numeric answer. It checks whether   โ”‚
โ”‚ the model handles repeated toggles plus the cat reset rule and returns 48.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=5.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ hotel-lights                       โ”‚
โ”‚ prompt         โ”‚ profile:hotel-lights               โ”‚
โ”‚ prompt chars   โ”‚ 385                                โ”‚
โ”‚ requested runs โ”‚ 30                                 โ”‚
โ”‚ concurrency    โ”‚ 5                                  โ”‚
โ”‚ max tokens     โ”‚ 40000                              โ”‚
โ”‚ scoring        โ”‚ numeric_exact                      โ”‚
โ”‚ expected       โ”‚ 48                                 โ”‚
โ”‚ prefill scout  โ”‚ 102 prompt tok / 0.19s = 526 tok/s โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,129 W | max 1,145 W | limit 1,200 W | over 38m 36s | 964 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 15 / NEAR 0 โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚  18,048 โ”‚ 29,997 โ”‚     52.4 โ”‚ 361.2 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 15 / NEAR 0 / FAIL 15 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โœ•โœ•โœ•โœ•โœ• ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           2 โ”‚
โ”‚ completion tokens avg      โ”‚                      18,903 โ”‚
โ”‚ completion tokens p50      โ”‚                      18,048 โ”‚
โ”‚ completion tokens p90      โ”‚                      29,997 โ”‚
โ”‚ completion tokens p99      โ”‚                      40,000 โ”‚
โ”‚ elapsed avg                โ”‚                      361.2s โ”‚
โ”‚ TTFT avg                   โ”‚                       0.31s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        52.4 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        51.8 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                            
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  4 โ”‚ 5 โ”‚ 11,559 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚  5 โ”‚ 5 โ”‚ 40,000 โ”‚ FAIL  โ”‚ 31 (expected 48, got 31) | Let's check $k=2 \times โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ 3 \times 5 \times 7 \times 11 \times 13 \times 17  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ \times 19 \times 23 \times 29 \times 31 \          โ”‚
โ”‚  7 โ”‚ 5 โ”‚ 18,848 โ”‚ FAIL  โ”‚ 52 (expected 48, got 52) | 52                      โ”‚
โ”‚  8 โ”‚ 5 โ”‚  9,706 โ”‚ FAIL  โ”‚ 45 (expected 48, got 45) | 45                      โ”‚
โ”‚ 13 โ”‚ 5 โ”‚ 15,266 โ”‚ FAIL  โ”‚ 47 (expected 48, got 47) | 47                      โ”‚
โ”‚ 14 โ”‚ 5 โ”‚ 17,935 โ”‚ FAIL  โ”‚ 50 (expected 48, got 50) | 50                      โ”‚
โ”‚ 17 โ”‚ 5 โ”‚ 25,984 โ”‚ FAIL  โ”‚ 47 (expected 48, got 47) | 47                      โ”‚
โ”‚ 18 โ”‚ 5 โ”‚  6,827 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 20 โ”‚ 5 โ”‚ 20,794 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 22 โ”‚ 5 โ”‚ 18,403 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 24 โ”‚ 5 โ”‚ 11,375 โ”‚ FAIL  โ”‚ 42 (expected 48, got 42) | 42                      โ”‚
โ”‚ 26 โ”‚ 5 โ”‚ 40,000 โ”‚ FAIL  โ”‚ 46 (expected 48, got 46) | Let me review k=46      โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ˜…=EXACT, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotel-nvfp4_ds_mla-c5-r30.json
Hotel-lights low tier โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:hotel-lights                                                 โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 40000                                        โ”‚
โ”‚ Scoring: EXACT / FAIL final number                                           โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ A compact reasoning profile with a known numeric answer. It checks whether   โ”‚
โ”‚ the model handles repeated toggles plus the cat reset rule and returns 48.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=5.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ hotel-lights                       โ”‚
โ”‚ prompt         โ”‚ profile:hotel-lights               โ”‚
โ”‚ prompt chars   โ”‚ 385                                โ”‚
โ”‚ requested runs โ”‚ 30                                 โ”‚
โ”‚ concurrency    โ”‚ 5                                  โ”‚
โ”‚ max tokens     โ”‚ 40000                              โ”‚
โ”‚ scoring        โ”‚ numeric_exact                      โ”‚
โ”‚ expected       โ”‚ 48                                 โ”‚
โ”‚ prefill scout  โ”‚ 102 prompt tok / 0.19s = 524 tok/s โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,121 W | max 1,136 W | limit 1,200 W | over 43m 45s | 1,092 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 18 / NEAR 0 โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚  20,108 โ”‚ 30,229 โ”‚     52.0 โ”‚ 400.1 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 18 / NEAR 0 / FAIL 12 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โœ•โœ•โœ•โœ• ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           3 โ”‚
โ”‚ completion tokens avg      โ”‚                      20,806 โ”‚
โ”‚ completion tokens p50      โ”‚                      20,108 โ”‚
โ”‚ completion tokens p90      โ”‚                      30,229 โ”‚
โ”‚ completion tokens p99      โ”‚                      40,000 โ”‚
โ”‚ elapsed avg                โ”‚                      400.1s โ”‚
โ”‚ TTFT avg                   โ”‚                       0.31s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        52.0 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        51.6 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                            
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  5 โ”‚ 5 โ”‚  9,772 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚  6 โ”‚ 5 โ”‚  8,645 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚  9 โ”‚ 5 โ”‚  9,616 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 11 โ”‚ 5 โ”‚ 40,000 โ”‚ FAIL  โ”‚ 17801 (expected 48, got 17801) | What about $m=37  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ \times 37 \times 13 = 17801 >                      โ”‚
โ”‚ 12 โ”‚ 5 โ”‚ 23,780 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 14 โ”‚ 5 โ”‚ 17,984 โ”‚ FAIL  โ”‚ 46 (expected 48, got 46) | 46                      โ”‚
โ”‚ 15 โ”‚ 5 โ”‚ 40,000 โ”‚ FAIL  โ”‚ 0 (expected 48, got 0) | So $M(70) = 70$. Count =  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ 0.                                                 โ”‚
โ”‚ 16 โ”‚ 5 โ”‚ 14,794 โ”‚ FAIL  โ”‚ 47 (expected 48, got 47) | 47                      โ”‚
โ”‚ 19 โ”‚ 5 โ”‚ 16,012 โ”‚ FAIL  โ”‚ 47 (expected 48, got 47) | 47                      โ”‚
โ”‚ 23 โ”‚ 5 โ”‚ 11,105 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 26 โ”‚ 5 โ”‚  8,702 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 30 โ”‚ 5 โ”‚ 40,000 โ”‚ FAIL  โ”‚ 2 (expected 48, got 2) | So the elements in $D_2$  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ that are $\ge                                      โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ˜…=EXACT, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotel-fp8-c5-r30.json
Hotel-lights Max tier โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:hotel-lights                                                 โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 60000                                        โ”‚
โ”‚ Scoring: EXACT / FAIL final number                                           โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ A compact reasoning profile with a known numeric answer. It checks whether   โ”‚
โ”‚ the model handles repeated toggles plus the cat reset rule and returns 48.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=5.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ hotel-lights                       โ”‚
โ”‚ prompt         โ”‚ profile:hotel-lights               โ”‚
โ”‚ prompt chars   โ”‚ 385                                โ”‚
โ”‚ requested runs โ”‚ 30                                 โ”‚
โ”‚ concurrency    โ”‚ 5                                  โ”‚
โ”‚ max tokens     โ”‚ 60000                              โ”‚
โ”‚ scoring        โ”‚ numeric_exact                      โ”‚
โ”‚ expected       โ”‚ 48                                 โ”‚
โ”‚ prefill scout  โ”‚ 102 prompt tok / 0.20s = 507 tok/s โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,130 W | max 1,144 W | limit 1,200 W | over 77m 13s | 1,927 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 20 / NEAR 0 โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚  34,282 โ”‚ 60,000 โ”‚     51.6 โ”‚ 725.4 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 20 / NEAR 0 / FAIL 10 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โœ•โœ•โœ• ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           4 โ”‚
โ”‚ completion tokens avg      โ”‚                      37,389 โ”‚
โ”‚ completion tokens p50      โ”‚                      34,282 โ”‚
โ”‚ completion tokens p90      โ”‚                      60,000 โ”‚
โ”‚ completion tokens p99      โ”‚                      60,000 โ”‚
โ”‚ elapsed avg                โ”‚                      725.4s โ”‚
โ”‚ TTFT avg                   โ”‚                       0.31s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        51.6 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        51.2 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                            
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  1 โ”‚ 5 โ”‚ 24,913 โ”‚ FAIL  โ”‚ 50 (expected 48, got 50) | 50                      โ”‚
โ”‚  7 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 283 (expected 48, got 283) | Let's test $k' = 2    โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ \cdot 191 \cdot 283                                โ”‚
โ”‚ 11 โ”‚ 5 โ”‚ 32,608 โ”‚ FAIL  โ”‚ 40 (expected 48, got 40) | 40                      โ”‚
โ”‚ 14 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 8 (expected 48, got 8) | Divisors of 32: 1, 2, 4,  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ 8,                                                 โ”‚
โ”‚ 16 โ”‚ 5 โ”‚ 30,373 โ”‚ FAIL  โ”‚ 52 (expected 48, got 52) | 52                      โ”‚
โ”‚ 21 โ”‚ 5 โ”‚ 31,931 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 23 โ”‚ 5 โ”‚ 36,285 โ”‚ FAIL  โ”‚ 47 (expected 48, got 47) | 47                      โ”‚
โ”‚ 25 โ”‚ 5 โ”‚ 25,730 โ”‚ FAIL  โ”‚ 32 (expected 48, got 32) | 32                      โ”‚
โ”‚ 28 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 62 (expected 48, got 62) | m=62: 62,               โ”‚
โ”‚ 29 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 83 (expected 48, got 83) | Let's re-verify Room 83 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ˜…=EXACT, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotelmax-nvfp4_ds_mla-c5-r30.json
Hotel-lights Max tier โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Prompt: profile:hotel-lights                                                 โ”‚
โ”‚ Concurrency: 5                                                               โ”‚
โ”‚ Measured runs: 30 | Max tokens: 60000                                        โ”‚
โ”‚ Scoring: EXACT / FAIL final number                                           โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Completion Stats โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Hotel Lights Reasoning Test                                                  โ”‚
โ”‚ A compact reasoning profile with a known numeric answer. It checks whether   โ”‚
โ”‚ the model handles repeated toggles plus the cat reset rule and returns 48.   โ”‚
โ”‚ Built-in profile run at fixed concurrency C=5.                               โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Profile                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ field          โ”‚ value                              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ profile        โ”‚ hotel-lights                       โ”‚
โ”‚ prompt         โ”‚ profile:hotel-lights               โ”‚
โ”‚ prompt chars   โ”‚ 385                                โ”‚
โ”‚ requested runs โ”‚ 30                                 โ”‚
โ”‚ concurrency    โ”‚ 5                                  โ”‚
โ”‚ max tokens     โ”‚ 60000                              โ”‚
โ”‚ scoring        โ”‚ numeric_exact                      โ”‚
โ”‚ expected       โ”‚ 48                                 โ”‚
โ”‚ prefill scout  โ”‚ 102 prompt tok / 0.19s = 525 tok/s โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,124 W | max 1,138 W | limit 1,200 W | over 80m 42s | 2,012 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Concurrency Results                                                             
โ•ญโ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ•ฎ
โ”‚  โ”‚ donโ€ฆ โ”‚               score โ”‚ sโ€ฆ โ”‚ outputโ€ฆ โ”‚ outpuโ€ฆ โ”‚ aggregaโ€ฆ โ”‚ avg โ€ฆ โ”‚ โ€ฆ โ”‚
โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ค
โ”‚  โ”‚ 30/โ€ฆ โ”‚ EXACT 20 / NEAR 0 โ€ฆ โ”‚ โ˜…โ€ฆ โ”‚  34,446 โ”‚ 60,000 โ”‚     51.7 โ”‚ 742.4 โ”‚ โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ•ฏ
Selected C=5                                                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ metric                     โ”‚                       value โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ completed                  โ”‚                       30/30 โ”‚
โ”‚ score                      โ”‚ EXACT 20 / NEAR 0 / FAIL 10 โ”‚
โ”‚ stars                      โ”‚               โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โ˜…โœ•โœ•โœ• ๐Ÿ‘ โ”‚
โ”‚ hit max_tokens             โ”‚                           6 โ”‚
โ”‚ completion tokens avg      โ”‚                      38,346 โ”‚
โ”‚ completion tokens p50      โ”‚                      34,446 โ”‚
โ”‚ completion tokens p90      โ”‚                      60,000 โ”‚
โ”‚ completion tokens p99      โ”‚                      60,000 โ”‚
โ”‚ elapsed avg                โ”‚                      742.4s โ”‚
โ”‚ TTFT avg                   โ”‚                       0.31s โ”‚
โ”‚ aggregate gen tok/s        โ”‚                        51.7 โ”‚
โ”‚ mean per-request gen tok/s โ”‚                        51.5 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Failed Final Answers                                                            
โ•ญโ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚  # โ”‚ C โ”‚ tokens โ”‚ score โ”‚ final answer                                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  3 โ”‚ 5 โ”‚ 31,254 โ”‚ FAIL  โ”‚ 45 (expected 48, got 45) | 45                      โ”‚
โ”‚  5 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 31 (expected 48, got 31) | What about $n=31 \      โ”‚
โ”‚ 13 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 33 (expected 48, got 33) | D=11: 11, 33,           โ”‚
โ”‚ 16 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 4 (expected 48, got 4) | What about $m=25 \cdot 4  โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ =                                                  โ”‚
โ”‚ 17 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 38903 (expected 48, got 38903) | $38903 /          โ”‚
โ”‚ 18 โ”‚ 5 โ”‚ 17,288 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 19 โ”‚ 5 โ”‚ 19,525 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 24 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 3 (expected 48, got 3) | Guest 7: 7 times. $7      โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ \equiv 1 \pmod 3$. R -> G -> R.                    โ”‚
โ”‚ 27 โ”‚ 5 โ”‚ 26,455 โ”‚ FAIL  โ”‚ 49 (expected 48, got 49) | 49                      โ”‚
โ”‚ 30 โ”‚ 5 โ”‚ 60,000 โ”‚ FAIL  โ”‚ 5 (expected 48, got 5) | m=35: divs=[1,5,7,35].    โ”‚
โ”‚    โ”‚   โ”‚        โ”‚       โ”‚ n=1: 1->0. n=5:                                    โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ˜…=EXACT, โœ•=FAIL.

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotelmax-fp8-c5-r30.json

4. KLD vs BF16 reference

5 cold-reload runs per dtype (fresh container each run), 2,047 scored positions, WikiText-2 @ ctx 2048, speculative_config=None. Lower is better.

run nvfp4_ds_mla fp8
1 0.115928 0.100616
2 0.115569 0.102350
3 0.117059 0.100503
4 0.115955 0.100163
5 0.116181 0.102360
mean 0.116138 0.101198
sd 0.000559 0.001069

fp8 shows 0.0149 (-12.9%) lower divergence from the BF16 reference. nvfp4_ds_mla remains the shipped default because it costs roughly half the KV bytes per token โ€” that is what provides the context headroom at this preset โ€” and it matched or beat fp8 on the retrieval and ledger tasks.

This is the honest measure of what quantization costs: divergence is small but not zero.

5. Prefill throughput

--standalone-prefill --prefill-only --prefill-contexts 8k,64k,128k.

Context nvfp4 TTFT nvfp4 tok/s fp8 TTFT fp8 tok/s
8k 3.21 s 2,551 3.08 s 2,660
64k 32.96 s 1,957 33.80 s 1,909
128k 70.31 s 1,833 72.78 s 1,771
Prefill โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1, 2, 4, 8, 16, 32, 64, 128]                            โ”‚
โ”‚ Decode contexts: ['0', '16k', '32k', '64k', '128k']                          โ”‚
โ”‚ Decode: skipped (--prefill-only) | Max tokens: 8192                          โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill-only: standalone cold profile (auto) | Sustained decode: 0 cells     โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: standalone cold profile ['8k', '64k', '128k']
Calibrating padding text (run=zamzztpcfkqj, up to 128k)...
  Token targeting: single-point estimate from 8k (use --token-targeting exact 
for /tokenize binary search)
  Calibrated: 6.18 chars/token (cached, source=8k)
  8k: 50,601 chars (~8,191 tokens)
  16k: 101,202 chars (~16,383 tokens)
  32k: 202,404 chars (~32,767 tokens)
  64k: 404,809 chars (~65,535 tokens)
  128k: 809,618 chars (~131,071 tokens)
Done.



llm-decode-bench v0.4.29
Prefill Speed (C=1, client ISL / TTFT)                                          
                                                                                
                                                                PCIe rx/tx      
  Context    Tokens   TTFT (s)   Client tok/s   Server tok/s           avg   N  
 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ 
  8k          8,202       3.21          2,551      2,565 (2)   21578/21298   2  
  64k        64,516      32.96          1,957      1,963 (1)   77950/79445   1  
  128k      128,890      70.31          1,833      1,838 (1)   88481/90763   1  
                                                                                
Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the 
prefix-cache scout request that decode needs anyway. Server tok/s is optional 
Prometheus validation when the engine exports prefill counters and the exact 
counter delta is uncontaminated; for vLLM this uses newly computed KV tokens, 
not request prompt tokens.


Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/prefill-nvfp4_ds_mla.json
Prefill โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1, 2, 4, 8, 16, 32, 64, 128]                            โ”‚
โ”‚ Decode contexts: ['0', '16k', '32k', '64k', '128k']                          โ”‚
โ”‚ Decode: skipped (--prefill-only) | Max tokens: 8192                          โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill-only: standalone cold profile (auto) | Sustained decode: 0 cells     โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: standalone cold profile ['8k', '64k', '128k']
Calibrating padding text (run=vzteflzjnjfk, up to 128k)...
  Token targeting: single-point estimate from 8k (use --token-targeting exact 
for /tokenize binary search)
  Calibrated: 6.18 chars/token (cached, source=8k)
  8k: 50,601 chars (~8,191 tokens)
  16k: 101,202 chars (~16,383 tokens)
  32k: 202,404 chars (~32,767 tokens)
  64k: 404,809 chars (~65,535 tokens)
  128k: 809,618 chars (~131,071 tokens)
Done.



llm-decode-bench v0.4.29
Prefill Speed (C=1, client ISL / TTFT)                                          
                                                                                
                                                                PCIe rx/tx      
  Context    Tokens   TTFT (s)   Client tok/s   Server tok/s           avg   N  
 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ 
  8k          8,201       3.08          2,660      2,674 (2)   25261/22621   2  
  64k        64,515      33.80          1,909      1,915 (1)   80688/78187   1  
  128k      128,889      72.78          1,771      1,776 (1)   85527/85895   1  
                                                                                
Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the 
prefix-cache scout request that decode needs anyway. Server tok/s is optional 
Prometheus validation when the engine exports prefill counters and the exact 
counter delta is uncontaminated; for vLLM this uses newly computed KV tokens, 
not request prompt tokens.


Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/prefill-fp8.json

6. Decode throughput, C1-C8

--concurrency 1,2,3,4,5,6,7,8 --contexts 0 --duration 20 --max-tokens 8192 --skip-prefill --temperature 0. Aggregate tokens/sec:

concurrency 1 2 3 4 5 6 7 8
nvfp4_ds_mla 87.5 147.4 183.7 219.3 242.4 274.9 291.1 308.1
fp8 86.6 143.3 184.7 217.2 241.2 270.3 289.6 312.2

Per-user decode (nvfp4): 87.5 / 73.7 / 61.2 / 54.8 / 48.5 / 45.8 / 41.6 / 38.5 tok/s โ€” 3.5x aggregate scaling C1 to C8. Single-stream figures are for the shipped lossless configuration (VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE=0); the lossy alternative trades correctness for about +10 tok/s at C1 and is deliberately disabled.

Decode C1-C8 โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1, 2, 3, 4, 5, 6, 7, 8]                                 โ”‚
โ”‚ Decode contexts: ['0']                                                       โ”‚
โ”‚ Duration: 20.0s per decode test | Max tokens: 8192                           โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill: skipped | Sustained decode: 8 cells                                 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.



llm-decode-bench v0.4.29
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 2 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Sustained Decode                                                             โ”‚
โ”‚ Steady-state decode throughput after the engine has admitted the requested   โ”‚
โ”‚ concurrency and passed warmup. Use this as the main tuning/regression signal โ”‚
โ”‚ for kernels, NCCL, DCP, MTP, and scheduler changes.                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate tok/s  + TTFT/ITL                                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \  โ”‚        โ”‚        โ”‚        โ”‚        โ”‚        โ”‚       โ”‚        โ”‚       โ”‚
โ”‚ conc   โ”‚      1 โ”‚      2 โ”‚      3 โ”‚      4 โ”‚      5 โ”‚     6 โ”‚      7 โ”‚     8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0      โ”‚   87.5 โ”‚  147.4 โ”‚  183.7 โ”‚  219.3 โ”‚  242.4 โ”‚ 274.9 โ”‚  291.1 โ”‚ 308.1 โ”‚
โ”‚        โ”‚ 161/11 โ”‚ 253/13 โ”‚ 363/16 โ”‚ 398/18 โ”‚ 442/20 โ”‚ 472/โ€ฆ โ”‚ 499/23 โ”‚ 542/โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default 
(continuous completion_tokens when the server supports it). Prometheus is kept 
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s                                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚    2 โ”‚    3 โ”‚    4 โ”‚    5 โ”‚    6 โ”‚    7 โ”‚    8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 87.5 โ”‚ 73.7 โ”‚ 61.2 โ”‚ 54.8 โ”‚ 48.5 โ”‚ 45.8 โ”‚ 41.6 โ”‚ 38.5 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Client request latency: p50 / p90 ms                          
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚   1 โ”‚   2 โ”‚   3 โ”‚   4 โ”‚   5 โ”‚   6 โ”‚   7 โ”‚   8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc 
coordinate. ITL is computed from observed generated tokens, including streams 
stopped at the measurement boundary; a missing ITL means no stream produced at 
least two measured output tokens. Per-request tok/s and request latency are 
shown in separate per-cell matrices. Completion/sample counts and full 
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate 
tok/s remains the primary throughput signal. 
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary                                                                
โ•ญโ”€โ”€โ”€โ”ฌโ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ โ€ฆ โ”‚ โ”‚ mode  โ”‚ GPU avg/โ€ฆ โ”‚ Mem โ€ฆ โ”‚ W avg/โ€ฆ โ”‚ T โ€ฆ โ”‚ CPUโ€ฆ โ”‚ VRโ€ฆ โ”‚ PCIe rx/tx aโ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”ผโ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   48% โ”‚ 1067/1โ€ฆ โ”‚ 86C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   10945/10913 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   49% โ”‚ 1104/1โ€ฆ โ”‚ 89C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   17508/17924 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   48% โ”‚ 1126/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   22034/21826 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   49% โ”‚ 1136/1โ€ฆ โ”‚ 89C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   24354/24152 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚    99/99% โ”‚   47% โ”‚ 1137/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   26736/26290 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚    99/99% โ”‚   47% โ”‚ 1141/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   29873/29618 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚    99/99% โ”‚   47% โ”‚ 1144/1โ€ฆ โ”‚ 90C โ”‚  79C โ”‚ 94โ€ฆ โ”‚   31554/31721 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚   99/100% โ”‚   46% โ”‚ 1144/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 94โ€ฆ โ”‚   33035/33073 โ”‚
โ•ฐโ”€โ”€โ”€โ”ดโ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,062 W | max 1,147 W | limit 1,200 W | over 3m 51s | 97 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Hardware summary is sampled from nvidia-smi during the measured part of each 
cell. Whole-run GPU power is the sampled sum of GPU power draw across the 
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is 
a coarse live diagnostic, not a per-kernel NCCL profiler.

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 3 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Burst / E2E Decode                                                           โ”‚
โ”‚ Not run. Re-run with --run-burst to append a finite client-facing request    โ”‚
โ”‚ burst after Sustained Decode. This is intentionally disabled by default      โ”‚
โ”‚ because it adds another full decode matrix.                                  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Primary Summary โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Primary matrices repeated last so the important numbers are visible without  โ”‚
โ”‚ scrolling back through diagnostics.                                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate decode tok/s                                                       
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚     2 โ”‚     3 โ”‚     4 โ”‚     5 โ”‚     6 โ”‚     7 โ”‚     8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 87.5 โ”‚ 147.4 โ”‚ 183.7 โ”‚ 219.3 โ”‚ 242.4 โ”‚ 274.9 โ”‚ 291.1 โ”‚ 308.1 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1c8-nvfp4_ds_mla.json
Decode C1-C8 โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1, 2, 3, 4, 5, 6, 7, 8]                                 โ”‚
โ”‚ Decode contexts: ['0']                                                       โ”‚
โ”‚ Duration: 20.0s per decode test | Max tokens: 8192                           โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill: skipped | Sustained decode: 8 cells                                 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.



llm-decode-bench v0.4.29
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 2 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Sustained Decode                                                             โ”‚
โ”‚ Steady-state decode throughput after the engine has admitted the requested   โ”‚
โ”‚ concurrency and passed warmup. Use this as the main tuning/regression signal โ”‚
โ”‚ for kernels, NCCL, DCP, MTP, and scheduler changes.                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate tok/s  + TTFT/ITL                                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \  โ”‚        โ”‚        โ”‚        โ”‚        โ”‚        โ”‚       โ”‚        โ”‚       โ”‚
โ”‚ conc   โ”‚      1 โ”‚      2 โ”‚      3 โ”‚      4 โ”‚      5 โ”‚     6 โ”‚      7 โ”‚     8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0      โ”‚   86.6 โ”‚  143.3 โ”‚  184.7 โ”‚  217.2 โ”‚  241.2 โ”‚ 270.3 โ”‚  289.6 โ”‚ 312.2 โ”‚
โ”‚        โ”‚ 153/11 โ”‚ 244/14 โ”‚ 363/16 โ”‚ 393/18 โ”‚ 443/20 โ”‚ 474/โ€ฆ โ”‚ 510/23 โ”‚ 538/โ€ฆ โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default 
(continuous completion_tokens when the server supports it). Prometheus is kept 
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s                                                     
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚    2 โ”‚    3 โ”‚    4 โ”‚    5 โ”‚    6 โ”‚    7 โ”‚    8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 86.6 โ”‚ 71.7 โ”‚ 61.6 โ”‚ 54.3 โ”‚ 48.2 โ”‚ 45.0 โ”‚ 41.4 โ”‚ 39.0 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Client request latency: p50 / p90 ms                          
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚   1 โ”‚   2 โ”‚   3 โ”‚   4 โ”‚   5 โ”‚   6 โ”‚   7 โ”‚   8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚ โ€”/โ€” โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc 
coordinate. ITL is computed from observed generated tokens, including streams 
stopped at the measurement boundary; a missing ITL means no stream produced at 
least two measured output tokens. Per-request tok/s and request latency are 
shown in separate per-cell matrices. Completion/sample counts and full 
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate 
tok/s remains the primary throughput signal. 
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary                                                                
โ•ญโ”€โ”€โ”€โ”ฌโ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ โ€ฆ โ”‚ โ”‚ mode  โ”‚ GPU avg/โ€ฆ โ”‚ Mem โ€ฆ โ”‚ W avg/โ€ฆ โ”‚ T โ€ฆ โ”‚ CPUโ€ฆ โ”‚ VRโ€ฆ โ”‚ PCIe rx/tx aโ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”ผโ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   47% โ”‚ 1067/1โ€ฆ โ”‚ 86C โ”‚  74C โ”‚ 95โ€ฆ โ”‚   10754/10642 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   48% โ”‚ 1100/1โ€ฆ โ”‚ 88C โ”‚  73C โ”‚ 95โ€ฆ โ”‚   17246/17439 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   47% โ”‚ 1117/1โ€ฆ โ”‚ 89C โ”‚  73C โ”‚ 95โ€ฆ โ”‚   21808/21260 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   47% โ”‚ 1126/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 95โ€ฆ โ”‚   24263/24104 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚    99/99% โ”‚   46% โ”‚ 1131/1โ€ฆ โ”‚ 90C โ”‚  74C โ”‚ 95โ€ฆ โ”‚   25999/25642 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚    99/99% โ”‚   46% โ”‚ 1134/1โ€ฆ โ”‚ 90C โ”‚  74C โ”‚ 95โ€ฆ โ”‚   28964/28378 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚   99/100% โ”‚   46% โ”‚ 1138/1โ€ฆ โ”‚ 90C โ”‚  74C โ”‚ 95โ€ฆ โ”‚   30996/31210 โ”‚
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚   99/100% โ”‚   45% โ”‚ 1137/1โ€ฆ โ”‚ 90C โ”‚  73C โ”‚ 95โ€ฆ โ”‚   33155/32814 โ”‚
โ•ฐโ”€โ”€โ”€โ”ดโ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 1,060 W | max 1,139 W | limit 1,200 W | over 3m 50s | 96 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Hardware summary is sampled from nvidia-smi during the measured part of each 
cell. Whole-run GPU power is the sampled sum of GPU power draw across the 
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is 
a coarse live diagnostic, not a per-kernel NCCL profiler.

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 3 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Burst / E2E Decode                                                           โ”‚
โ”‚ Not run. Re-run with --run-burst to append a finite client-facing request    โ”‚
โ”‚ burst after Sustained Decode. This is intentionally disabled by default      โ”‚
โ”‚ because it adds another full decode matrix.                                  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Primary Summary โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Primary matrices repeated last so the important numbers are visible without  โ”‚
โ”‚ scrolling back through diagnostics.                                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate decode tok/s                                                       
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚     2 โ”‚     3 โ”‚     4 โ”‚     5 โ”‚     6 โ”‚     7 โ”‚     8 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 86.6 โ”‚ 143.3 โ”‚ 184.7 โ”‚ 217.2 โ”‚ 241.2 โ”‚ 270.3 โ”‚ 289.6 โ”‚ 312.2 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1c8-fp8.json
Decode C1 dedicated 30 s โ€” nvfp4_ds_mla log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1]                                                      โ”‚
โ”‚ Decode contexts: ['0']                                                       โ”‚
โ”‚ Duration: 30.0s per decode test | Max tokens: 8192                           โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill: skipped | Sustained decode: 1 cells                                 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.



llm-decode-bench v0.4.29
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 2 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Sustained Decode                                                             โ”‚
โ”‚ Steady-state decode throughput after the engine has admitted the requested   โ”‚
โ”‚ concurrency and passed warmup. Use this as the main tuning/regression signal โ”‚
โ”‚ for kernels, NCCL, DCP, MTP, and scheduler changes.                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate tok/s  + TTFT/ITL 
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚           1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 83.5 160/12 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default 
(continuous completion_tokens when the server supports it). Prometheus is kept 
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s    
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 83.5 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Client request      
latency: p50 / p90  
ms                  
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚   1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ โ€”/โ€” โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc 
coordinate. ITL is computed from observed generated tokens, including streams 
stopped at the measurement boundary; a missing ITL means no stream produced at 
least two measured output tokens. Per-request tok/s and request latency are 
shown in separate per-cell matrices. Completion/sample counts and full 
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate 
tok/s remains the primary throughput signal. 
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary                                                                
โ•ญโ”€โ”€โ”€โ”ฌโ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ โ€ฆ โ”‚ โ”‚ mode  โ”‚ GPU avg/โ€ฆ โ”‚ Mem โ€ฆ โ”‚ W avg/โ€ฆ โ”‚ T โ€ฆ โ”‚ CPUโ€ฆ โ”‚ VRโ€ฆ โ”‚ PCIe rx/tx aโ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”ผโ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   47% โ”‚ 1071/1โ€ฆ โ”‚ 86C โ”‚  74C โ”‚ 94โ€ฆ โ”‚   10791/10730 โ”‚
โ•ฐโ”€โ”€โ”€โ”ดโ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 996 W | max 1,099 W | limit 1,200 W | over 48s | 21 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Hardware summary is sampled from nvidia-smi during the measured part of each 
cell. Whole-run GPU power is the sampled sum of GPU power draw across the 
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is 
a coarse live diagnostic, not a per-kernel NCCL profiler.

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 3 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Burst / E2E Decode                                                           โ”‚
โ”‚ Not run. Re-run with --run-burst to append a finite client-facing request    โ”‚
โ”‚ burst after Sustained Decode. This is intentionally disabled by default      โ”‚
โ”‚ because it adds another full decode matrix.                                  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Primary Summary โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Primary matrices repeated last so the important numbers are visible without  โ”‚
โ”‚ scrolling back through diagnostics.                                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate decode     
tok/s                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 83.5 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1-nvfp4_ds_mla.json
Decode C1 dedicated 30 s โ€” fp8 log
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Configuration โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ LLM Inference Benchmark                                                      โ”‚
โ”‚ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000                              โ”‚
โ”‚ Decode concurrency: [1]                                                      โ”‚
โ”‚ Decode contexts: ['0']                                                       โ”‚
โ”‚ Duration: 30.0s per decode test | Max tokens: 8192                           โ”‚
โ”‚ Pre-decode warmup: C=1 max-runnable context for 3s                           โ”‚
โ”‚ Prefill: skipped | Sustained decode: 1 cells                                 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Engine: vLLM 
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722  
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร— 64; local 65,536 ร—
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.



llm-decode-bench v0.4.29
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 2 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Sustained Decode                                                             โ”‚
โ”‚ Steady-state decode throughput after the engine has admitted the requested   โ”‚
โ”‚ concurrency and passed warmup. Use this as the main tuning/regression signal โ”‚
โ”‚ for kernels, NCCL, DCP, MTP, and scheduler changes.                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate tok/s  + TTFT/ITL 
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚           1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 82.4 153/12 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default 
(continuous completion_tokens when the server supports it). Prometheus is kept 
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s    
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 82.4 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Client request      
latency: p50 / p90  
ms                  
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚   1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ โ€”/โ€” โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc 
coordinate. ITL is computed from observed generated tokens, including streams 
stopped at the measurement boundary; a missing ITL means no stream produced at 
least two measured output tokens. Per-request tok/s and request latency are 
shown in separate per-cell matrices. Completion/sample counts and full 
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate 
tok/s remains the primary throughput signal. 
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary                                                                
โ•ญโ”€โ”€โ”€โ”ฌโ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ โ€ฆ โ”‚ โ”‚ mode  โ”‚ GPU avg/โ€ฆ โ”‚ Mem โ€ฆ โ”‚ W avg/โ€ฆ โ”‚ T โ€ฆ โ”‚ CPUโ€ฆ โ”‚ VRโ€ฆ โ”‚ PCIe rx/tx aโ€ฆ โ”‚
โ”œโ”€โ”€โ”€โ”ผโ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0 โ”‚ โ”‚ sustโ€ฆ โ”‚  100/100% โ”‚   46% โ”‚ 1066/1โ€ฆ โ”‚ 86C โ”‚  75C โ”‚ 95โ€ฆ โ”‚   10586/10558 โ”‚
โ•ฐโ”€โ”€โ”€โ”ดโ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Whole-run GPU Power โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ avg 993 W | max 1,099 W | limit 1,200 W | over 48s | 21 samples โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Hardware summary is sampled from nvidia-smi during the measured part of each 
cell. Whole-run GPU power is the sampled sum of GPU power draw across the 
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is 
a coarse live diagnostic, not a per-kernel NCCL profiler.

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Phase 3 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Burst / E2E Decode                                                           โ”‚
โ”‚ Not run. Re-run with --run-burst to append a finite client-facing request    โ”‚
โ”‚ burst after Sustained Decode. This is intentionally disabled by default      โ”‚
โ”‚ because it adds another full decode matrix.                                  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Primary Summary โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ Primary matrices repeated last so the important numbers are visible without  โ”‚
โ”‚ scrolling back through diagnostics.                                          โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ
Aggregate decode     
tok/s                
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ ctx \ conc โ”‚    1 โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ 0          โ”‚ 82.4 โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

Results saved to 
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1-fp8.json

7. Power

Whole-run GPU power across the 4-GPU set (1,200 W aggregate limit):

Run avg max duration
LAVD nvfp4 1,130 W 1,143 W 19m 20s
LAVD fp8 1,127 W 1,138 W 19m 45s
Estonia nvfp4 1,105 W 1,124 W 10m 56s
Estonia fp8 1,101 W 1,127 W 14m 28s

Sustained draw is 92-94% of the aggregate limit. Two of the four cards are 600 W-capable parts software-limited to 300 W, so single-stream decode is partly bounded by the slowest rank's clock.


8. Independent third-party evaluation

Run and published by malaiwah/glm52-exl3-vast on 2026-07-23/24, independently of this repository's author. The original write-up, the client harness, and the raw per-question summary are mirrored under independent-eval/.

Results (pass@1, pooled across repeats)

Benchmark Questions x repeats = n EXL3 3.0bpw GLM-5.2 BF16 (Z.ai published) ~95% CI
AIME 2026 30 x 4 = 120 99.2 99.2 ยฑ1.6
HMMT Feb 2026 33 x 4 = 132 95.5 92.5 ยฑ3.6
GPQA Diamond 198 x 2 = 396 91.4 91.2 ยฑ2.8

All three land within sampling noise of the BF16 reference โ€” no measurable reasoning degradation was detected at 3.0 bits per weight.

GPQA per-question stability across its 2 repeats: 173 of 198 questions correct both times, 16 split 1-of-2, 9 wrong both times. Across all 396 generations: 0 truncations, 0 errors, 1 sample with no extractable answer, 10,561 avg completion tokens, 15.9 h wall time.

How it was run

  • Sampling matched to Z.ai's published eval settings: temperature=1.0, top_p=0.95
  • Max generation: 163,840 tokens (math), 131,072 (GPQA). Zero truncations occurred.
  • No thinking-effort override โ€” server default reasoning mode
  • Math prompt: Z.ai's Explanation: / Exact Answer: / Confidence: system prompt
  • Math datasets: MathArena/aime_2026, MathArena/hmmt_feb_2026
  • Math grading: math-verify symbolic equivalence; fallback chain Exact Answer: line -> last \boxed{} -> none
  • GPQA: Idavidrein/gpqa (gpqa_diamond), simple-evals / Artificial-Analysis MCQ template, options deterministically shuffled per (question, repeat), regex letter extraction
  • Client: async Python harness, 32 concurrent requests (16 GPQA + 8 AIME + 8 HMMT) with all three benchmarks running simultaneously; ~65 tok/s aggregate under that mixed long-reasoning load; 8.63M completion tokens over ~16 h
  • pass@1 computed over all repeats pooled

How this run differed from the shipped preset

Reported by the runner as fp8 KV cache, max_model_len 524288, server version 0.17.0rc1.dev4499+g60c82d972 โ€” corresponding to the earlier image tag v1-gg-60c82d972-spi1937274-cu132-sm120a rather than the published v20-gg6722c1d-si1a88b38. Same model weights, same compose and server script, MTP-3 enabled (speculative decoding affects throughput, not the output distribution).

Reading these numbers fairly

  • The BF16 column is Z.ai's published figures, not a re-measurement on this harness. The math grading used math-verify symbolic equivalence rather than Z.ai's GPT-5.5 judge. This is therefore measured-versus-published, not a controlled head-to-head.
  • HMMT +3.0 over BF16 should not be read as the quant beating full precision โ€” a quantization cannot exceed its source in expectation. With 33 questions, a ยฑ3.6 interval and a different grader, that gap is noise plus methodology.
  • Confidence intervals are simple binomial approximations. Repeats of the same question are correlated, so true intervals are somewhat wider.
  • AIME and HMMT are small sets (30 and 33 questions). GPQA Diamond at 198 x 2 is the most statistically solid of the three.

Known reproducible quirks (seen on multiple quants, likely model-level)

  • HMMT Q20: a common reasoning path converges on 1100 where the gold answer is 20460. Reproduced across different quantizations of this base model; this quant scored 2 of 4 repeats.
  • GPQA idx 79 (dataset order): triggers unusually long reasoning chains.

Incident note from the runner

Three requests stalled mid-run on dropped server connections โ€” client sockets stayed ESTABLISHED while the server no longer tracked the request. All three hit the client read timeout, auto-retried and completed. Zero lost or errored samples in the final data. Suggested hardening: TCP keepalives plus a tighter per-request timeout.

Reproduce with the mirrored harness (point --base-url at any OpenAI-compatible endpoint):

python independent-eval/mathbench.py  --dataset MathArena/aime_2026     --repeats 4 --concurrency 8
python independent-eval/mathbench.py  --dataset MathArena/hmmt_feb_2026 --repeats 4 --concurrency 8
python independent-eval/gpqa_bench.py --repeats 2 --concurrency 16

Reproducing the serving benchmarks

Note (2026-07-25): the docker-compose.yml and server.sh embedded further below reproduce the BF16-MTP Sections 1-8. The repo's live server.sh / docker-compose.yml are the tr3-MTP build (v21 image, NUM_GPU_BLOCKS_OVERRIDE empty โ†’ auto-profile, VLLM_EXL3_TRELLIS_MIN_M=1, MAX_MODEL_LEN=524288); ./server.sh start below pulls and runs that current preset.

hf download brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw --local-dir "$HOME/models/GLM-5.2-EXL3-TR3-3.0bpw"
cd "$HOME/models/GLM-5.2-EXL3-TR3-3.0bpw"
chmod +x server.sh && ./server.sh start
Full docker-compose.yml (all serve flags)
services:
  glm52:
    image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff}
    container_name: glm52-exl3-sparkinfer
    ports:
      - "${BIND_ADDRESS:-127.0.0.1}:${PORT:-8000}:8000"
    gpus: all
    shm_size: "32g"
    ipc: host
    ulimits:
      memlock: -1
      nofile: 1048576
    environment:
      CUDA_VISIBLE_DEVICES: "${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
      CUDA_DEVICE_ORDER: PCI_BUS_ID
      CUDA_DEVICE_MAX_CONNECTIONS: "32"
      CUTE_DSL_ARCH: sm_120a
      TORCH_CUDA_ARCH_LIST: 12.0a
      FLASHINFER_CUDA_ARCH_LIST: 12.0f
      FLASHINFER_DISABLE_VERSION_CHECK: "1"
      OMP_NUM_THREADS: "16"
      PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
      SAFETENSORS_FAST_GPU: "1"
      NCCL_IB_DISABLE: "1"
      NCCL_P2P_LEVEL: SYS
      NCCL_PROTO: LL,LL128,Simple
      VLLM_USE_FLASHINFER_SAMPLER: "1"
      VLLM_USE_B12X_FP8_GEMM: "1"
      VLLM_USE_B12X_SPARSE_INDEXER: "1"
      VLLM_USE_B12X_MOE: "1"
      VLLM_USE_V2_MODEL_RUNNER: "1"
      VLLM_ENABLE_PCIE_ALLREDUCE: "1"
      VLLM_PCIE_ALLREDUCE_BACKEND: b12x
      VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE: "${VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE:-64KB}"
      VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE: "${VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE:-84KB}"
      VLLM_PCIE_DMA_FP8: ag
      B12X_PCIE_DMA_FP8: ag
      VLLM_CPP_AR_1STAGE_NCCL_CUTOFF: 56KB
      VLLM_CPP_AR_IGNORE_CUTOFF_MAX_ROWS: "0"
      VLLM_RTX6K_FUSED_ALLREDUCE_ADD: "0"
      VLLM_RTX6K_FUSED_ALLREDUCE_ADD_END_BARRIER: "0"
      VLLM_USE_AOT_COMPILE: "1"
      VLLM_USE_BREAKABLE_CUDAGRAPH: "0"
      VLLM_USE_FUSED_MOE_GROUPED_TOPK: "1"
      VLLM_USE_B12X_MHC: "1"
      B12X_MHC_MAX_TOKENS: "16384"
      VLLM_USE_B12X_WO_PROJECTION: "1"
      B12X_MLA_SM120_UNIFIED: "1"
      B12X_DENSE_SPLITK_TURBO: "1"
      B12X_W4A16_TC_DECODE: "1"
      B12X_MOE_FORCE_A16: "1"
      VLLM_DISABLE_SHARED_EXPERTS_STREAM: "${VLLM_DISABLE_SHARED_EXPERTS_STREAM:-1}"
      VLLM_DISABLED_KERNELS: MarlinFP8ScaledMMLinearKernel
      VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE: "${VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE:-0}"
      VLLM_B12X_MLA_SPEC_DECODE_MAX_Q: "8"
      VLLM_USE_B12X_DCP_A2A: "1"
      VLLM_DCP_A2A_MAX_TOKENS: "16"
      VLLM_DCP_A2A_LARGE_BACKEND: ag_rs
      VLLM_DCP_GLOBAL_TOPK: "${VLLM_DCP_GLOBAL_TOPK:-1}"
      VLLM_DCP_SHARD_DRAFT: "${VLLM_DCP_SHARD_DRAFT:-1}"
      VLLM_DCP_QUERY_SPLIT: "0"
      VLLM_B12X_MLA_CKV_GATHER: "1"
      VLLM_B12X_MLA_CKV_GATHER_MIN_TOKENS: "512"
      VLLM_B12X_MLA_CKV_GATHER_MAX_TOKENS: "16384"
      ENABLE_MTP: "${ENABLE_MTP:-1}"
      MTP_TOKENS: "${MTP_TOKENS:-3}"
      MTP_DRAFT_SAMPLE_METHOD: "${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
      ENABLE_ASYNC_SCHEDULING: "${ENABLE_ASYNC_SCHEDULING:-0}"
      GLM52_INDEX_TOPK_PATTERN: "${GLM52_INDEX_TOPK_PATTERN:-FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS}"
      NUM_GPU_BLOCKS_OVERRIDE: "${NUM_GPU_BLOCKS_OVERRIDE:-1024}"
      MAX_NUM_BATCHED_TOKENS: "${MAX_NUM_BATCHED_TOKENS:-3072}"
      VLLM_EXL3_TRELLIS_MIN_M: "4"
      VLLM_EXL3_TRELLIS_MAX_M: "32"
      VLLM_EXL3_TRELLIS_BLOCK_M: "8"
      VLLM_EXL3_PREFILL_CHUNK: "128"
      VLLM_CACHE_DIR: /cache/jit/vllm
      TRITON_CACHE_DIR: /cache/jit/triton
      TORCH_EXTENSIONS_DIR: /cache/jit/torch_extensions
      TORCHINDUCTOR_CACHE_DIR: /cache/jit/torchinductor
      FLASHINFER_WORKSPACE_BASE: /cache/jit/flashinfer
      XDG_CACHE_HOME: /cache/jit
      TVM_FFI_CACHE_DIR: /cache/jit/tvm-ffi
      VLLM_MEMORY_PROFILE_INCLUDE_ATTN: "1"
      VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS: "1"
      VLLM_DEBUG_WORKSPACE: "${VLLM_DEBUG_WORKSPACE:-0}"
    volumes:
      - ${MODEL_DIR:-/home/brandonmusic/models/GLM-5.2-EXL3-TR3-3.0bpw}:/model:ro
      - ${CACHE_DIR:-/home/brandonmusic/.cache/glm52-tr3-release}:/cache:rw
    entrypoint:
      - /bin/bash
      - -lc
    command:
      - |
        unset NCCL_GRAPH_FILE NCCL_GRAPH_DUMP_FILE VLLM_B12X_MLA_EXTEND_MAX_CHUNKS
        spec_args=()
        if [[ "$${ENABLE_MTP:-1}" == "1" ]]; then
          printf -v spec_config '{"method":"mtp","num_speculative_tokens":%s,"moe_backend":"triton","draft_sample_method":"%s"}' \
            "$${MTP_TOKENS:-3}" "$${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
          spec_args=(--speculative-config "$$spec_config")
        fi
        if [[ "$${ENABLE_ASYNC_SCHEDULING:-0}" == "1" ]]; then
          async_args=(--async-scheduling)
        else
          async_args=(--no-async-scheduling)
        fi
        index_pattern="$${GLM52_INDEX_TOPK_PATTERN}"
        if [[ "$${#index_pattern}" -ne 78 ]]; then
          printf 'GLM-5.2 index_topk_pattern must cover all 78 layers (got %s)\n' "$${#index_pattern}" >&2
          exit 2
        fi
        printf -v hf_overrides '{"use_index_cache":true,"index_topk_pattern":"%s"}' "$${index_pattern}"
        block_args=()
        if [[ -n "$${NUM_GPU_BLOCKS_OVERRIDE:-}" ]]; then
          block_args=(--num-gpu-blocks-override "$${NUM_GPU_BLOCKS_OVERRIDE}")
        fi
        exec vllm serve /model \
          --served-model-name GLM-5.2-EXL3-TR3-3.0bpw \
          --host 0.0.0.0 --port 8000 --trust-remote-code \
          --tensor-parallel-size 4 \
          --decode-context-parallel-size 4 \
          --dcp-comm-backend a2a \
          --dcp-kv-cache-interleave-size ${DCP_KV_CACHE_INTERLEAVE_SIZE:-64} \
          --seed 0 \
          --quantization exl3 \
          --kv-cache-dtype ${KV_CACHE_DTYPE:-nvfp4_ds_mla} \
          --attention-backend B12X_MLA_SPARSE \
          --moe-backend b12x \
          --load-format safetensors \
          --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[4,8,12,16,20,24,28,32],"custom_ops":["all"],"pass_config":{"fuse_allreduce_rms":true}}' \
          --gpu-memory-utilization ${GPU_MEMORY_UTILIZATION:-0.96} \
          --max-model-len ${MAX_MODEL_LEN:-262144} \
          --max-num-seqs 8 \
          --max-num-batched-tokens $${MAX_NUM_BATCHED_TOKENS:-3072} \
          --max-cudagraph-capture-size 32 \
          --enable-chunked-prefill \
          --enable-prefix-caching \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --default-chat-template-kwargs '{"reasoning_effort":"high"}' \
          --hf-overrides "$${hf_overrides}" \
          "$${block_args[@]}" \
          "$${async_args[@]}" \
          "$${spec_args[@]}"

Full server.sh
#!/usr/bin/env bash
set -Eeuo pipefail

SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"

export IMAGE="${IMAGE:-verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff}"
export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
export PORT="${PORT:-8000}"
export BIND_ADDRESS="${BIND_ADDRESS:-127.0.0.1}"
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.96}"
export MAX_MODEL_LEN="${MAX_MODEL_LEN:-262144}"
export DCP_KV_CACHE_INTERLEAVE_SIZE="${DCP_KV_CACHE_INTERLEAVE_SIZE:-64}"
export VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE="${VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE:-0}"
export VLLM_DCP_SHARD_DRAFT="${VLLM_DCP_SHARD_DRAFT:-1}"
export VLLM_DCP_GLOBAL_TOPK="${VLLM_DCP_GLOBAL_TOPK:-1}"
export VLLM_DISABLE_SHARED_EXPERTS_STREAM="${VLLM_DISABLE_SHARED_EXPERTS_STREAM:-1}"
export VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE="${VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE:-64KB}"
export VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE="${VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE:-84KB}"
export ENABLE_MTP="${ENABLE_MTP:-1}"
export MTP_TOKENS="${MTP_TOKENS:-3}"
export MTP_DRAFT_SAMPLE_METHOD="${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
export ENABLE_ASYNC_SCHEDULING="${ENABLE_ASYNC_SCHEDULING:-0}"
export GLM52_INDEX_TOPK_PATTERN="${GLM52_INDEX_TOPK_PATTERN:-FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS}"
export NUM_GPU_BLOCKS_OVERRIDE="${NUM_GPU_BLOCKS_OVERRIDE:-1024}"
export MAX_NUM_BATCHED_TOKENS="${MAX_NUM_BATCHED_TOKENS:-3072}"
export COMPOSE_PROJECT_NAME="${COMPOSE_PROJECT_NAME:-glm52-exl3-sparkinfer}"

COMPOSE_FILE="${COMPOSE_FILE:-/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d17377be8eb/scratchpad/testsuite/config/docker-compose.yml}"
COMPOSE=(docker compose -f "$COMPOSE_FILE")

usage() {
  cat <<'EOF'
Usage: ./server.sh [start|stop|restart|logs|status|pull]

Environment overrides:
  IMAGE, MODEL_DIR, CACHE_DIR, PORT, BIND_ADDRESS, CUDA_VISIBLE_DEVICES,
  GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, DCP_KV_CACHE_INTERLEAVE_SIZE,
  VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE, VLLM_DCP_SHARD_DRAFT,
  VLLM_DCP_GLOBAL_TOPK,
  VLLM_DISABLE_SHARED_EXPERTS_STREAM,
  VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE,
  VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE,
  ENABLE_MTP, MTP_TOKENS, MTP_DRAFT_SAMPLE_METHOD, ENABLE_ASYNC_SCHEDULING,
  GLM52_INDEX_TOPK_PATTERN,
  NUM_GPU_BLOCKS_OVERRIDE,
  MAX_NUM_BATCHED_TOKENS,
  COMPOSE_PROJECT_NAME, COMPOSE_FILE
EOF
}

require_runtime() {
  command -v docker >/dev/null 2>&1 || {
    echo "docker is required" >&2
    exit 1
  }
  docker compose version >/dev/null
  [[ -f "$COMPOSE_FILE" ]] || {
    echo "Compose file not found: $COMPOSE_FILE" >&2
    exit 1
  }
}

require_model() {
  [[ -f "$MODEL_DIR/config.json" ]] || {
    echo "Model config not found: $MODEL_DIR/config.json" >&2
    exit 1
  }
  [[ -f "$MODEL_DIR/model.safetensors.index.json" ]] || {
    echo "Model index not found: $MODEL_DIR/model.safetensors.index.json" >&2
    exit 1
  }
  mkdir -p "$CACHE_DIR"
}

action="${1:-start}"
require_runtime

case "$action" in
  start)
    require_model
    docker pull "$IMAGE"
    "${COMPOSE[@]}" up -d --force-recreate
    echo "Starting on http://localhost:$PORT"
    echo "Follow startup with: $0 logs"
    ;;
  stop)
    "${COMPOSE[@]}" down
    ;;
  restart)
    require_model
    docker pull "$IMAGE"
    "${COMPOSE[@]}" up -d --force-recreate
    echo "Restarting on http://localhost:$PORT"
    ;;
  logs)
    "${COMPOSE[@]}" logs --tail 100 -f glm52
    ;;
  status)
    "${COMPOSE[@]}" ps
    curl -fsS "http://localhost:$PORT/v1/models" || true
    printf '\n'
    ;;
  pull)
    docker pull "$IMAGE"
    ;;
  -h|--help|help)
    usage
    ;;
  *)
    usage >&2
    exit 2
    ;;
esac

Long-context needle-in-a-haystack (v28, 2026-07-26)

Run on 4x RTX PRO 6000 Blackwell (SM120), TP4 / DCP4, nvfp4_ds_mla KV, MTP-3 draft at layer 78, VLLM_EXL3_TRELLIS_MIN_M=1, max-model-len 524288. A unique authorization code is planted at three depths (0.1 / 0.5 / 0.9) in a filler document and requested back with greedy decoding.

nvfp4_ds_mla KV

target ctx real prompt tokens depth 0.1 depth 0.5 depth 0.9
8k 4,844 HIT HIT HIT
32k 19,316 HIT HIT HIT
65k 39,188 HIT HIT HIT
128k 77,168 HIT HIT HIT
200k 199,783 HIT HIT HIT
300k 299,648 HIT HIT HIT
400k 399,512 HIT HIT HIT
480k 479,396 HIT HIT HIT

24/24 needles recovered. GPU KV cache 959,744-998,400 tokens. Deepest 480k probe 315 s.

fp8 KV

Identical checkpoint, weights, draft and flags; only --kv-cache-dtype changed.

target ctx real prompt tokens depth 0.1 depth 0.5 depth 0.9
8k 8,010 HIT HIT HIT
65k 64,965 HIT HIT HIT
128k 127,854 HIT HIT HIT
200k 199,784 HIT HIT HIT
300k 299,648 HIT HIT HIT
480k 479,396 HIT HIT HIT

18/18 needles recovered. GPU KV cache 648,192 tokens (8-bit vs 4-bit, so a smaller pool than nvfp4_ds_mla at the same utilization). Deepest 480k probe 311 s.

Combined: 42/42 across both KV dtypes, three depths each, to ~480k real prompt tokens -- within ~45k of the 524,288 max-model-len ceiling. No garbled output and zero engine restarts in either lane.

Context: vLLM issue #183 reported that VLLM_EXL3_TRELLIS_MIN_M=1 silently corrupts long-context output. That did not reproduce here in either KV dtype. Independently, the fused Trellis MoE was verified bitwise-correct at m=1,2,3 at this exact geometry (tile 64x256x64x256, block_size_m=8, capacity 32, topk=8) with the scratch arena NaN-poisoned. Note that MIN_M=1 widens the Trellis window so the draft's m=1..3 GEMMs stay on the fused, graph-capturable path; it is not a per-token capture and carries no throughput penalty.

Boot without the MIN_M workaround (v29, 2026-07-27)

Historically, serving an EXL3 rank-sliced tr3 MTP draft required setting VLLM_EXL3_TRELLIS_MIN_M=1 by hand; without it the engine could not start (vLLM issue #183):

RuntimeError: EXL3 eager parity path entered during CUDA graph capture (m=3);
              capture sizes must lie inside the Trellis window [4, 32]

v29 removes the requirement. The backend stamps each layer's draft/target role at construction (runner_type == "draft"), where the vllm-config context is live, and defaults draft layers' Trellis window to MIN_CAPTURABLE_TRELLIS_M=1 automatically. Target layers keep the historical default of 4; an explicit VLLM_EXL3_TRELLIS_MIN_M still overrides both.

Validation on this rig (4x RTX PRO 6000 SM120, TP4/DCP4, tr3 MTP-78, MTP-3), with VLLM_EXL3_TRELLIS_MIN_M entirely unset:

gate result
engine boot + serve PASS (previously guaranteed startup failure)
capture-time window error in logs 0 occurrences
blank-env int('') startup crash 0 occurrences (blank now means unset)
greedy inference PASS
needle 8k/65k/128k x depths 0.1/0.5/0.9 9/9 recovered

The compose files in this repo now leave VLLM_EXL3_TRELLIS_MIN_M unset by default. Fix commits: vLLM PR #139 239ba678b5 + 796ea923f1.

v30 (2026-07-27): env-knob registration + sparkinfer PR#79 module

Delta vs v29 (which carries all correctness fixes): (1) the nine EXL3 env knobs are registered in vllm envs.py -- startup 'Unknown vLLM environment variable' warnings drop 15 -> 8 (the remaining 8 are base-runtime-owned), and the knobs join the torch.compile cache-key factors (one-time ~70 s recompile after changing one); (2) the SparkInfer wheel is rebuilt from the PR#49 branch rebased onto master AFTER PR#79 ("perf(pcie): add exact DCP top-k owner exchange"), so the CUDA-IPC owner-exchange module ships in the image. It is DORMANT here: PR#79 has zero overlap with the EXL3/MoE lane (only +1 line outside its new files), and the vLLM-side owner algorithm is not yet in this image's pinned base. Gates on the pinned v30: boot with VLLM_EXL3_TRELLIS_MIN_M unset PASS, 0 capture-window errors, warnings 15->8 confirmed, greedy inference PASS, tool-calls 4/4.

v31 (2026-07-27): unified v20 base refresh (SparkInfer c3828fd)

Base bump only on the SparkInfer axis (vLLM pin unchanged at 0c79e41, so the 13-file vLLM overlay is byte-identical to v30). Wheel rebuilt from PR#49 (11 commits) on the new integration pin c3828fd and verified a strict superset of the base's canonical SparkInfer (168/168 source files present). Gates on the pinned digest: boot with VLLM_EXL3_TRELLIS_MIN_M unset PASS, warnings 8, 0 capture errors, inference PASS, tool-calls 4/4, KV 963,840.

Correction note for v30: its wheel was built from SparkInfer master rather than the base's integration pin, so the pip install replaced the base's canonical SparkInfer with a tree missing the integration-only PCIe calibration commits. No effect on the published serving configs (they pin DCP controls explicitly, and calibration only engages on 'auto'), but helper/auto-calibration users should prefer v31.