Instructions to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") model = AutoModelForCausalLM.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Trellis
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
- SGLang
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
Download RELEASE_TEST_SUITE.md from brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw: direct link, hf CLI and curl.
- Browser
- Download file 144 kB
-
https://huggingface.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw/resolve/main/RELEASE_TEST_SUITE.md
- Command line
-
hf download hf://brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw/RELEASE_TEST_SUITE.md
-
curl -L -o RELEASE_TEST_SUITE.md https://huggingface.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw/resolve/main/RELEASE_TEST_SUITE.md
Release Test Suite โ GLM-5.2-EXL3-TR3-3.0bpw
Measured 2026-07-25 on 4x RTX PRO 6000 Blackwell 96 GB
(TP4/DCP4, MTP-3 greedy) with the published sha-pinned image and the shipped
server.sh / docker-compose.yml preset, using llm_decode_bench.py with the same flags as
run_release_benchmarks.sh. Every raw benchmark log is embedded verbatim beneath its table.
Section 8 reports an independent third-party evaluation run by malaiwah/glm52-exl3-vast.
2026-07-25 update (3) โ per-model runtime scoping + per-ordinal arch cache key (v26)
Follow-up to the v20 rebase, from CodeRabbit review on local-inference-lab/vllm#139.
Bug: the EXL3 rank-sliced runtime cache was keyed only on device, dtype, shape,
topk and planner settings. The cached entry owns mutable Trellis/prefill scratch and
parity staging buffers. A target MoE layer and the rank-sliced MTP-78 draft layer
match on every one of those components -- same hidden/intermediate size, same local
expert count, same topk, same planner env, and both resolve max_num_batched_tokens
from the same scheduler config -- so the draft was reusing the target's scratch. That
defeats the target/draft isolation their independently captured CUDA graphs rely on.
Fix: the cache key is now scoped to the owning quant config, so each model gets exactly one runtime. This is deliberately coarser than per-layer: the prefill arena is ~1054 MiB, so per-layer runtimes would need tens of GiB per rank across 75+ layers.
Runtime: verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a (sha256:8753406fโฆ).
Measured effect (the fix is observable in memory, not in the log -- the planner line
uses info_once and is deduplicated):
| Rank-sliced runtimes | GPU KV cache | Concurrency @524K | |
|---|---|---|---|
| pre-fix (shared scratch) | 1 (target+draft collide) | 1,115,904 tokens | 2.13x |
| v26 (scoped) | 2 (isolated) | 998,656 tokens | 1.90x |
The 119,552-token KV reduction corresponds to ~1056 MiB, matching the 1054.2 MiB arena -- direct evidence that a second runtime is allocated. Lower KV capacity is the intended cost of the isolation.
Second fix in this runtime (v26). CodeRabbit correctly rejected a first attempt at
the compile-cache device key: torch.cuda.get_device_capability() /
get_device_name() already resolve against torch.cuda.current_device(), so passing
that ordinal explicitly was a no-op and left the process-wide key free to freeze
whichever GPU was current on the first call. The real fix memoizes the architecture key
per device ordinal and threads the ordinal through _static_compile_cache_context,
which is lru_cached on the compile callable and would otherwise have re-frozen the
identity at that layer. The returned key still omits the ordinal, so GPUs of the same
architecture keep sharing compiled artifacts. This matters on this rig specifically:
its four boards report two different device names (Max-Q and non-Max-Q) while sharing
compute capability 12.0.
Quality on v26:
| Suite | Config | Result |
|---|---|---|
| Estonia | c2, 5 runs | 5/5 pass, 0 fail, correct rate 1.00 |
| LAVD | c5, 5 runs | 3 EXACT / 2 NEAR / 0 FAIL, correct rate 1.00 |
For continuity: the pre-fix runtime measured LAVD 2E/3N/0F. An intermediate scoped build measured 2E/2N/1F on one 5-run sample and 2E/3N/0F on a second; that single failure was a wrong ledger total at 7,114 completion tokens against a 24,576 cap, i.e. an answer-quality miss rather than truncation, and within this profile's run-to-run spread. This runtime shows no failures on either suite.
2026-07-25 update (2) โ rebased onto the FINAL Gilded Gnosis v20 base
The runtime is rebased onto the finalized v20 common base
voipmonitor/vllm:gilded-gnosis-v20-vllm0c79e41-sic3828fd-fi801d57a-cu132-20260727
(vLLM 5517197, Sparkinfer be0edca, FlashInfer 801d57a, CUTLASS 4.6.0).
New runtime: verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a
(sha256:da185fe8โฆ).
Why: the finalized base consolidates the DCP prefill auto-policy and the
corrected workspace accounting, resolving the >8k DCP prefill collapse present
in earlier v20 candidates, plus long-context MTP alignment and deterministic
dynamic-MoE output. EXL3 is enabled by rebasing the EXL3 source layer onto that
pinned stack (the base ships no EXL3/Trellis loader), keeping EXL3 quantization
separate while sharing the corrected runtime. Of the 12 runtime overlay files only
models/deepseek_v2.py and v1/attention/backends/mla/indexer.py differ in the new
base; both were re-derived from the v20-final versions with the EXL3 edits replayed
on top, so v20's DCP/indexer work is preserved rather than overwritten by the overlay.
Config change: server.sh / docker-compose.yml now set the DCP policy that the
base launcher resolves from DCP_*=auto for TP4/DCP4 โ query split 1, full-CKV gather 1,
top-k owner merge 1, indexer shards 0, CKV prefetch depth 1, prefetch workspace 1024 MiB,
and DCP_PREFILL_WORKSPACE=1 (VLLM_DCP_PROJECT_BEFORE_MERGE=1 +
VLLM_B12X_MLA_DCP_GATHER_IN_WORKSPACE=1). These are set explicitly because the preset
calls vllm serve directly and bypasses /usr/local/bin/serve-gilded-gnosis.sh.
Note VLLM_DCP_QUERY_SPLIT moves from 0 to 1 versus the previous preset.
Boot assertions observed: engine
v0.11.2.dev280+gilded.gnosis.v20.vllm5517197.sibe0edca.fi801d57a.cu132.20260725,
vLLM is using nccl==2.30.4, EXL3 rank-sliced runtime planned: Trellis m=1..32 block_m=8, prefill trellis block_m=64 arena=1054.2MiB capacity=3072 chunk=128 topk=8,
Preallocated 30.8 MiB for 2 persistent CKV execution lane(s), Using native CKV layer prefetch with depth=1 and 2 workspace slots, GPU KV cache 1,115,904 tokens (2.13x at
524,288). The base's InstantTensor loader line does not appear on this path because the
EXL3 checkpoint loads through the EXL3 rank-sliced loader.
Regression (no degradation):
| Suite | Config | Result |
|---|---|---|
| Estonia | c2, 5 runs | 5/5 pass, 0 fail, correct rate 1.00, 70.0 tok/s aggregate |
| LAVD | c5, 5 runs | 2 EXACT / 3 NEAR / 0 FAIL, correct rate 1.00, 55.6 tok/s aggregate |
All sections below were measured on the previous (v21-mtp78tr3) image and remain
valid for the checkpoint itself; only the runtime base changed.
2026-07-25 update โ MTP layer-78 is now EXL3 tr3
This checkpoint now ships the MTP (layer 78) routed experts in EXL3 Trellis tr3 (3.0 bpw), matching layers 3-77; the previous BF16 MTP head is retired. The layer-78 file drops from 19.9 GB to 4.24 GB (โ15.66 GB), freeing ~3.9 GiB/rank.
Requires the updated runtime:
verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a
(sha256:9b1befc1โฆ) plus VLLM_EXL3_TRELLIS_MIN_M=1 (the compose / server.sh
default in this repo). The prior v20 image cannot load a tr3 MTP layer. The two
loader fixes are in vLLM PR #139 (local-inference-lab/vllm#139); no Sparkinfer
change was needed (validated against #49 / si1a88b38).
Re-measured on the same 4ร RTX PRO 6000 (TP4/DCP4, MTP-3, util 0.96,
VLLM_EXL3_TRELLIS_MIN_M=1, auto-profiled KV):
| Metric | BF16-MTP build (Sections below) | tr3-MTP build (this) |
|---|---|---|
| GPU KV cache @ 0.96 util | 1,132,544 tok (2.16ร @ 524K) | |
| Prefill 8k / 64k / 128k (tok/s) | 2,551 / โ / 1,833 | 2,521 / 1,916 / 1,765 |
| Decode C1 / C4 / C8 (tok/s) | 87.5 / 219.3 / 308.1 | 89.7 / 225.3 / 293.5 |
| Estonia (long-ctx retrieval) | PASS 30/30 | PASS 10/10 |
| LAVD (ledger consistency) | 18E / 11N / 1F | EXACT 5 / NEAR 5 / FAIL 0 |
Decode is neutral (MTP is lossless); the real gain is **+66% KV-cache /
concurrency headroom** from the freed VRAM. Accuracy is unchanged โ the detailed
Sections 1-8 below were measured on the prior BF16-MTP build and remain
representative for quality; only the image, the layer-78 format, and the
KV/serving preset changed.
What this quantization costs
| BF16 full precision | 1,506 GB |
| EXL3 TR3 3.0bpw | 316.5 GB (tr3-MTP build; was 332.2 GB with the BF16 MTP head) |
| Size vs BF16 | 21.0% (a 79.0% reduction) |
| Effective whole-model rate | ~3.45 bpw (routed experts, now including MTP layer-78, are a flat 3.0; attention, shared experts, embeddings, LM head, and the non-expert parts of layer 78 stay BF16) |
BF16 weights alone would need 16x 96 GB cards. This fits on 4 with room for the KV cache.
Image
verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff
This is the current runtime. The per-benchmark sections further down were
measured on the earlier v21-mtp78tr3 image and are retained as the
checkpoint-level record; see the dated update sections above for what changed
in the runtime since, and for the re-verification run on this image.
Serving preset (as shipped)
| Setting | Value |
|---|---|
GPU_MEMORY_UTILIZATION |
0.96 |
MAX_MODEL_LEN |
524288 (tr3-MTP build; was 262144) |
NUM_GPU_BLOCKS_OVERRIDE |
empty โ auto-profile (~1,132,544 KV tokens @ 0.96; was pinned 1024) |
VLLM_EXL3_TRELLIS_MIN_M |
1 (required for the tr3 MTP draft's small-m GEMMs; was 4) |
MAX_NUM_BATCHED_TOKENS |
3072 |
MAX_NUM_SEQS |
8 |
| MTP | enabled, 3 tokens, greedy draft |
ENABLE_ASYNC_SCHEDULING |
0 (correctness guard) |
VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE |
0 (lossless setting) |
| Attention / MoE | B12X_MLA_SPARSE / b12x |
| Quantization | exl3 |
| KV cache | nvfp4_ds_mla (shipped default) and fp8 (comparison arm) |
Only --kv-cache-dtype and (for section 3) --default-chat-template-kwargs were parameterized;
every other flag is byte-identical to the published compose file.
Summary
| Test | Runs | nvfp4_ds_mla | fp8 |
|---|---|---|---|
| Estonia (needle retrieval, 133K ctx) | 30 @ c2 | PASS 30 / FAIL 0 | PASS 30 / FAIL 0 |
| LAVD (ledger consistency) | 30 @ c5 | 18 EXACT / 11 NEAR / 1 FAIL (97%) | 15 EXACT / 13 NEAR / 2 FAIL (93%) |
| Hotel-lights, low tier | 30 @ c5 | 15 EXACT / 15 FAIL (50%) | 18 EXACT / 12 FAIL (60%) |
| Hotel-lights, Max tier | 30 @ c5 | 20 EXACT / 10 FAIL (67%) | 20 EXACT / 10 FAIL (67%) |
| KLD vs BF16 | 5 | 0.116138 | 0.101198 |
| Decode C1 / C8 (tok/s) | โ | 87.5 / 308.1 | 86.6 / 312.2 |
| Prefill 8k / 128k (tok/s) | โ | 2,551 / 1,833 | 2,660 / 1,771 |
1. Estonia โ long-context needle retrieval
133,186-token prompt. 30 runs, concurrency 2, temperature 0, repetition penalty 1.25,
max_tokens 40000, regex-scored on the final answer line.
| KV cache | score | completed | hit max_tokens | tok p50 | avg latency | TTFT | gen tok/s |
|---|---|---|---|---|---|---|---|
| nvfp4_ds_mla | PASS 30 / FAIL 0 | 30/30 | 0 | 2,250 | 38.0 s | 0.61 s | 67.9 |
| fp8 | PASS 30 / FAIL 0 | 30/30 | 0 | 2,370 | 51.8 s | 0.60 s | 71.2 |
100% on both KV formats, no run near the token cap. The repetition penalty matters here: without it the model loops on retrieval phrasing and exhausts the output budget without answering.
Estonia โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Completion Token Statistics Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:estonia โ
โ Concurrency: 2 โ
โ Measured runs: 30 | Max tokens: 40000 โ
โ Scoring: \bestonia\b โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Completion Token Statistics โ
โ One optional prefix-cache scout request is used to populate prefill first. โ
โ Built-in profile run at fixed concurrency C=2. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ estonia โ
โ prompt โ profile:estonia โ
โ prompt chars โ 707,372 โ
โ requested runs โ 30 โ
โ concurrency โ 2 โ
โ max tokens โ 40000 โ
โ scoring โ regex โ
โ prefill scout โ 133,186 prompt tok / 79.72s = 1,671 tok/s โ
โ correct regex โ \bestonia\b โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,105 W | max 1,124 W | limit 1,200 W | over 10m 56s | 273 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโโโฌโโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโโโฌโโโโฎ
โ pโฆ โ done/โฆ โ score โ staโฆ โ outputโฆ โ output โฆ โ aggregaโฆ โ avg reโฆ โ โฆ โ
โโโโโโผโโโโโโโโโผโโโโโโโโโโโโผโโโโโโโผโโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโโโผโโโโค
โ 2 โ 30/30 โ PASS 30 โฆ โ โ
โ
โ
โฆ โ 2,250 โ 3,687 โ 67.9 โ 38.0 โ โฆ โ
โฐโโโโโดโโโโโโโโโดโโโโโโโโโโโโดโโโโโโโดโโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโโโดโโโโฏ
Selected C=2
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ PASS 30 / FAIL 0 โ
โ stars โ โ
โ
โ
โ
โ
โ
โ
โ
โ
โ
๐ โ
โ hit max_tokens โ 0 โ
โ completion tokens avg โ 2,536 โ
โ completion tokens p50 โ 2,250 โ
โ completion tokens p90 โ 3,687 โ
โ completion tokens p99 โ 4,725 โ
โ elapsed avg โ 38.0s โ
โ TTFT avg โ 0.61s โ
โ aggregate gen tok/s โ 67.9 โ
โ mean per-request gen tok/s โ 67.9 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโฏ
Interpretation: completion-token p50/p90/p99 tells how many decode tokens the
model needed to reach its final answer under this engine/config. Correctness is
scored from the final non-empty answer line by default, matching the GLM
dense-MLA vs NSA benchmark methodology. The prefill scout is not a scored
answer; it is the max_tokens=1 prefix-cache warmup and its prompt_tokens/TTFT
value is reported as scout prefill speed. Concurrency Results groups completed
requests by parallelism; Completed Requests shows the latest individual finished
answers.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/estonia-nvfp4_ds_mla-c2-r30-rp125.json
Estonia โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Completion Token Statistics Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:estonia โ
โ Concurrency: 2 โ
โ Measured runs: 30 | Max tokens: 40000 โ
โ Scoring: \bestonia\b โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Completion Token Statistics โ
โ One optional prefix-cache scout request is used to populate prefill first. โ
โ Built-in profile run at fixed concurrency C=2. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ estonia โ
โ prompt โ profile:estonia โ
โ prompt chars โ 707,372 โ
โ requested runs โ 30 โ
โ concurrency โ 2 โ
โ max tokens โ 40000 โ
โ scoring โ regex โ
โ prefill scout โ 133,186 prompt tok / 78.91s = 1,688 tok/s โ
โ correct regex โ \bestonia\b โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,101 W | max 1,127 W | limit 1,200 W | over 14m 28s | 361 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโโโฌโโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโโโฌโโโโฎ
โ pโฆ โ done/โฆ โ score โ staโฆ โ outputโฆ โ output โฆ โ aggregaโฆ โ avg reโฆ โ โฆ โ
โโโโโโผโโโโโโโโโผโโโโโโโโโโโโผโโโโโโโผโโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโโโผโโโโค
โ 2 โ 30/30 โ PASS 30 โฆ โ โ
โ
โ
โฆ โ 2,370 โ 4,513 โ 71.2 โ 51.8 โ โฆ โ
โฐโโโโโดโโโโโโโโโดโโโโโโโโโโโโดโโโโโโโดโโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโโโดโโโโฏ
Selected C=2
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ PASS 30 / FAIL 0 โ
โ stars โ โ
โ
โ
โ
โ
โ
โ
โ
โ
โ
๐ โ
โ hit max_tokens โ 0 โ
โ completion tokens avg โ 3,644 โ
โ completion tokens p50 โ 2,370 โ
โ completion tokens p90 โ 4,513 โ
โ completion tokens p99 โ 26,158 โ
โ elapsed avg โ 51.8s โ
โ TTFT avg โ 0.60s โ
โ aggregate gen tok/s โ 71.2 โ
โ mean per-request gen tok/s โ 67.3 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโฏ
Interpretation: completion-token p50/p90/p99 tells how many decode tokens the
model needed to reach its final answer under this engine/config. Correctness is
scored from the final non-empty answer line by default, matching the GLM
dense-MLA vs NSA benchmark methodology. The prefill scout is not a scored
answer; it is the max_tokens=1 prefix-cache warmup and its prompt_tokens/TTFT
value is reported as scout prefill speed. Concurrency Results groups completed
requests by parallelism; Completed Requests shows the latest individual finished
answers.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/estonia-fp8-c2-r30-rp125.json
2. LAVD โ long-context ledger consistency
48,302-char structured ledger; the model must find human data-entry errors, apply the repair rule,
and return ticket count and hours. Ground truth 72, 46.0. 30 runs, concurrency 5,
temperature 0, repetition penalty 1.15, max_tokens 24576.
| KV cache | EXACT | NEAR | FAIL | pass (E+N) | hit max_tokens | tok p50 | avg latency | gen tok/s |
|---|---|---|---|---|---|---|---|---|
| nvfp4_ds_mla | 18 | 11 | 1 | 29/30 (97%) | 0 | 9,206 | 180.6 s | 53.3 |
| fp8 | 15 | 13 | 2 | 28/30 (93%) | 0 | 8,908 | 190.4 s | 52.2 |
All three failures are single-axis near-misses, not parse errors or truncation:
65, 42.5 (count -7) ยท 72, 41.75 (count exact, hours -4.25) ยท 66, 45.25 (count -6).
Each used fewer tokens than the ~9K median, i.e. they stopped searching early rather than looping.
LAVD โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LAVD Context Consistency Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:lavd-test โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 24576 โ
โ Scoring: EXACT / NEAR / FAIL numeric pair โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LAVD Context Consistency Test โ
โ Arithmetic is intentionally simple; the test checks whether the model keeps โ
โ a long structured context consistent, finds human data-entry errors, applies โ
โ the repair rule, and returns the final ticket count and hours. Built-in โ
โ profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโค
โ profile โ lavd-test โ
โ prompt โ profile:lavd-test โ
โ prompt chars โ 48,302 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 24576 โ
โ scoring โ ledger_lavd โ
โ expected โ 72, 46 โ
โ prompt sha256 โ 5c83674d5f0fd2a7 โ
โ dataset sha256 โ 612f8041bbca048c โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,130 W | max 1,143 W | limit 1,200 W | over 19m 20s | 483 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 18 / NEAR 11โฆ โ โ
โฆ โ 9,206 โ 11,799 โ 53.3 โ 180.6 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 18 / NEAR 11 / FAIL 1 โ
โ stars โ โ
โ
โ
โ
โ
โ
โโโโ ๐ โ
โ hit max_tokens โ 0 โ
โ completion tokens avg โ 9,517 โ
โ completion tokens p50 โ 9,206 โ
โ completion tokens p90 โ 11,799 โ
โ completion tokens p99 โ 14,634 โ
โ elapsed avg โ 180.6s โ
โ TTFT avg โ 1.99s โ
โ aggregate gen tok/s โ 53.3 โ
โ mean per-request gen tok/s โ 53.3 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 14 โ 5 โ 5,721 โ FAIL โ 65, 42.5 (count -7, hours -3.50) | 65, 42.5 โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the parsed final numeric pair is exactly 72, 46.0.
NEAR means both count and hours are within the configured tolerance; FAIL means
the answer was unparseable or outside tolerance. The 10-slot quality bar is a
rounded distribution: โ
=EXACT, โ=NEAR, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/lavd-nvfp4_ds_mla-c5-r30-rp115.json
LAVD โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LAVD Context Consistency Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:lavd-test โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 24576 โ
โ Scoring: EXACT / NEAR / FAIL numeric pair โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LAVD Context Consistency Test โ
โ Arithmetic is intentionally simple; the test checks whether the model keeps โ
โ a long structured context consistent, finds human data-entry errors, applies โ
โ the repair rule, and returns the final ticket count and hours. Built-in โ
โ profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโค
โ profile โ lavd-test โ
โ prompt โ profile:lavd-test โ
โ prompt chars โ 48,302 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 24576 โ
โ scoring โ ledger_lavd โ
โ expected โ 72, 46 โ
โ prompt sha256 โ 5c83674d5f0fd2a7 โ
โ dataset sha256 โ 612f8041bbca048c โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,127 W | max 1,138 W | limit 1,200 W | over 19m 45s | 493 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 15 / NEAR 13โฆ โ โ
โฆ โ 8,908 โ 13,173 โ 52.2 โ 190.4 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 15 / NEAR 13 / FAIL 2 โ
โ stars โ โ
โ
โ
โ
โ
โโโโโ ๐ โ
โ hit max_tokens โ 0 โ
โ completion tokens avg โ 9,831 โ
โ completion tokens p50 โ 8,908 โ
โ completion tokens p90 โ 13,173 โ
โ completion tokens p99 โ 14,658 โ
โ elapsed avg โ 190.4s โ
โ TTFT avg โ 1.95s โ
โ aggregate gen tok/s โ 52.2 โ
โ mean per-request gen tok/s โ 52.3 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 13 โ 5 โ 9,982 โ FAIL โ 72, 41.75 (count +0, hours -4.25) | 72, 41.75 โ
โ 23 โ 5 โ 7,131 โ FAIL โ 66, 45.25 (count -6, hours -0.75) | 66, 45.25 โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the parsed final numeric pair is exactly 72, 46.0.
NEAR means both count and hours are within the configured tolerance; FAIL means
the answer was unparseable or outside tolerance. The 10-slot quality bar is a
rounded distribution: โ
=EXACT, โ=NEAR, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/lavd-fp8-c5-r30-rp115.json
3. Hotel-lights โ reasoning, low tier vs Max tier
100-room light-cycling puzzle, expected answer 48. 30 runs each, concurrency 5, temperature 0, exact numeric scoring. Run at both reasoning tiers.
| Tier | KV cache | EXACT | FAIL | pass | hit max_tokens | avg completion tok | wall |
|---|---|---|---|---|---|---|---|
| low | nvfp4_ds_mla | 15 | 15 | 50% | 2 | 18,903 | 39 min |
| low | fp8 | 18 | 12 | 60% | 3 | 20,806 | 44 min |
| Max | nvfp4_ds_mla | 20 | 10 | 67% | 4 | 37,389 | 77 min |
| Max | fp8 | 20 | 10 | 67% | 6 | 38,346 | 81 min |
Max tier buys +7 to +17 points for 2x the reasoning tokens and 2x the wall time. Both KV formats converge to the same 67% at Max. Some runs still exhaust even a 60,000-token budget, so this task can absorb unbounded reasoning.
Reasoning-tier gotcha, worth knowing before reproducing. This model's chat_template.jinja
line 2 reads:
{%- set effective_reasoning_effort = 'high' if reasoning_effort is defined and reasoning_effort == 'high' else 'max' -%}
Only the literal string high selects the lower tier. Anything else โ including max, or
omitting the field โ selects Max. The shipped compose sends --default-chat-template-kwargs '{"reasoning_effort":"high"}', so the default
preset runs the lower tier. The Max rows above were produced by setting that server default to
max; each run logged the resolved value from the live container to confirm the tier applied.
Hotel-lights low tier โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:hotel-lights โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 40000 โ
โ Scoring: EXACT / FAIL final number โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ A compact reasoning profile with a known numeric answer. It checks whether โ
โ the model handles repeated toggles plus the cat reset rule and returns 48. โ
โ Built-in profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ hotel-lights โ
โ prompt โ profile:hotel-lights โ
โ prompt chars โ 385 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 40000 โ
โ scoring โ numeric_exact โ
โ expected โ 48 โ
โ prefill scout โ 102 prompt tok / 0.19s = 526 tok/s โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,129 W | max 1,145 W | limit 1,200 W | over 38m 36s | 964 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 15 / NEAR 0 โฆ โ โ
โฆ โ 18,048 โ 29,997 โ 52.4 โ 361.2 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 15 / NEAR 0 / FAIL 15 โ
โ stars โ โ
โ
โ
โ
โ
โโโโโ ๐ โ
โ hit max_tokens โ 2 โ
โ completion tokens avg โ 18,903 โ
โ completion tokens p50 โ 18,048 โ
โ completion tokens p90 โ 29,997 โ
โ completion tokens p99 โ 40,000 โ
โ elapsed avg โ 361.2s โ
โ TTFT avg โ 0.31s โ
โ aggregate gen tok/s โ 52.4 โ
โ mean per-request gen tok/s โ 51.8 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 4 โ 5 โ 11,559 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 5 โ 5 โ 40,000 โ FAIL โ 31 (expected 48, got 31) | Let's check $k=2 \times โ
โ โ โ โ โ 3 \times 5 \times 7 \times 11 \times 13 \times 17 โ
โ โ โ โ โ \times 19 \times 23 \times 29 \times 31 \ โ
โ 7 โ 5 โ 18,848 โ FAIL โ 52 (expected 48, got 52) | 52 โ
โ 8 โ 5 โ 9,706 โ FAIL โ 45 (expected 48, got 45) | 45 โ
โ 13 โ 5 โ 15,266 โ FAIL โ 47 (expected 48, got 47) | 47 โ
โ 14 โ 5 โ 17,935 โ FAIL โ 50 (expected 48, got 50) | 50 โ
โ 17 โ 5 โ 25,984 โ FAIL โ 47 (expected 48, got 47) | 47 โ
โ 18 โ 5 โ 6,827 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 20 โ 5 โ 20,794 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 22 โ 5 โ 18,403 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 24 โ 5 โ 11,375 โ FAIL โ 42 (expected 48, got 42) | 42 โ
โ 26 โ 5 โ 40,000 โ FAIL โ 46 (expected 48, got 46) | Let me review k=46 โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ
=EXACT, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotel-nvfp4_ds_mla-c5-r30.json
Hotel-lights low tier โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:hotel-lights โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 40000 โ
โ Scoring: EXACT / FAIL final number โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ A compact reasoning profile with a known numeric answer. It checks whether โ
โ the model handles repeated toggles plus the cat reset rule and returns 48. โ
โ Built-in profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ hotel-lights โ
โ prompt โ profile:hotel-lights โ
โ prompt chars โ 385 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 40000 โ
โ scoring โ numeric_exact โ
โ expected โ 48 โ
โ prefill scout โ 102 prompt tok / 0.19s = 524 tok/s โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,121 W | max 1,136 W | limit 1,200 W | over 43m 45s | 1,092 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 18 / NEAR 0 โฆ โ โ
โฆ โ 20,108 โ 30,229 โ 52.0 โ 400.1 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 18 / NEAR 0 / FAIL 12 โ
โ stars โ โ
โ
โ
โ
โ
โ
โโโโ ๐ โ
โ hit max_tokens โ 3 โ
โ completion tokens avg โ 20,806 โ
โ completion tokens p50 โ 20,108 โ
โ completion tokens p90 โ 30,229 โ
โ completion tokens p99 โ 40,000 โ
โ elapsed avg โ 400.1s โ
โ TTFT avg โ 0.31s โ
โ aggregate gen tok/s โ 52.0 โ
โ mean per-request gen tok/s โ 51.6 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 5 โ 5 โ 9,772 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 6 โ 5 โ 8,645 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 9 โ 5 โ 9,616 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 11 โ 5 โ 40,000 โ FAIL โ 17801 (expected 48, got 17801) | What about $m=37 โ
โ โ โ โ โ \times 37 \times 13 = 17801 > โ
โ 12 โ 5 โ 23,780 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 14 โ 5 โ 17,984 โ FAIL โ 46 (expected 48, got 46) | 46 โ
โ 15 โ 5 โ 40,000 โ FAIL โ 0 (expected 48, got 0) | So $M(70) = 70$. Count = โ
โ โ โ โ โ 0. โ
โ 16 โ 5 โ 14,794 โ FAIL โ 47 (expected 48, got 47) | 47 โ
โ 19 โ 5 โ 16,012 โ FAIL โ 47 (expected 48, got 47) | 47 โ
โ 23 โ 5 โ 11,105 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 26 โ 5 โ 8,702 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 30 โ 5 โ 40,000 โ FAIL โ 2 (expected 48, got 2) | So the elements in $D_2$ โ
โ โ โ โ โ that are $\ge โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ
=EXACT, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotel-fp8-c5-r30.json
Hotel-lights Max tier โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:hotel-lights โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 60000 โ
โ Scoring: EXACT / FAIL final number โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ A compact reasoning profile with a known numeric answer. It checks whether โ
โ the model handles repeated toggles plus the cat reset rule and returns 48. โ
โ Built-in profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ hotel-lights โ
โ prompt โ profile:hotel-lights โ
โ prompt chars โ 385 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 60000 โ
โ scoring โ numeric_exact โ
โ expected โ 48 โ
โ prefill scout โ 102 prompt tok / 0.20s = 507 tok/s โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,130 W | max 1,144 W | limit 1,200 W | over 77m 13s | 1,927 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 20 / NEAR 0 โฆ โ โ
โฆ โ 34,282 โ 60,000 โ 51.6 โ 725.4 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 20 / NEAR 0 / FAIL 10 โ
โ stars โ โ
โ
โ
โ
โ
โ
โ
โโโ ๐ โ
โ hit max_tokens โ 4 โ
โ completion tokens avg โ 37,389 โ
โ completion tokens p50 โ 34,282 โ
โ completion tokens p90 โ 60,000 โ
โ completion tokens p99 โ 60,000 โ
โ elapsed avg โ 725.4s โ
โ TTFT avg โ 0.31s โ
โ aggregate gen tok/s โ 51.6 โ
โ mean per-request gen tok/s โ 51.2 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 1 โ 5 โ 24,913 โ FAIL โ 50 (expected 48, got 50) | 50 โ
โ 7 โ 5 โ 60,000 โ FAIL โ 283 (expected 48, got 283) | Let's test $k' = 2 โ
โ โ โ โ โ \cdot 191 \cdot 283 โ
โ 11 โ 5 โ 32,608 โ FAIL โ 40 (expected 48, got 40) | 40 โ
โ 14 โ 5 โ 60,000 โ FAIL โ 8 (expected 48, got 8) | Divisors of 32: 1, 2, 4, โ
โ โ โ โ โ 8, โ
โ 16 โ 5 โ 30,373 โ FAIL โ 52 (expected 48, got 52) | 52 โ
โ 21 โ 5 โ 31,931 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 23 โ 5 โ 36,285 โ FAIL โ 47 (expected 48, got 47) | 47 โ
โ 25 โ 5 โ 25,730 โ FAIL โ 32 (expected 48, got 32) | 32 โ
โ 28 โ 5 โ 60,000 โ FAIL โ 62 (expected 48, got 62) | m=62: 62, โ
โ 29 โ 5 โ 60,000 โ FAIL โ 83 (expected 48, got 83) | Let's re-verify Room 83 โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ
=EXACT, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotelmax-nvfp4_ds_mla-c5-r30.json
Hotel-lights Max tier โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Prompt: profile:hotel-lights โ
โ Concurrency: 5 โ
โ Measured runs: 30 | Max tokens: 60000 โ
โ Scoring: EXACT / FAIL final number โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Completion Stats โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Hotel Lights Reasoning Test โ
โ A compact reasoning profile with a known numeric answer. It checks whether โ
โ the model handles repeated toggles plus the cat reset rule and returns 48. โ
โ Built-in profile run at fixed concurrency C=5. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Profile
โญโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ field โ value โ
โโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ profile โ hotel-lights โ
โ prompt โ profile:hotel-lights โ
โ prompt chars โ 385 โ
โ requested runs โ 30 โ
โ concurrency โ 5 โ
โ max tokens โ 60000 โ
โ scoring โ numeric_exact โ
โ expected โ 48 โ
โ prefill scout โ 102 prompt tok / 0.19s = 525 tok/s โ
โฐโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,124 W | max 1,138 W | limit 1,200 W | over 80m 42s | 2,012 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Concurrency Results
โญโโโฌโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโฌโโโโโฌโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโฌโโโโโโโโฌโโโโฎ
โ โ donโฆ โ score โ sโฆ โ outputโฆ โ outpuโฆ โ aggregaโฆ โ avg โฆ โ โฆ โ
โโโโผโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโผโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโผโโโโโโโโผโโโโค
โ โ 30/โฆ โ EXACT 20 / NEAR 0 โฆ โ โ
โฆ โ 34,446 โ 60,000 โ 51.7 โ 742.4 โ โฆ โ
โฐโโโดโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโดโโโโโโโโดโโโโฏ
Selected C=5
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ metric โ value โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ completed โ 30/30 โ
โ score โ EXACT 20 / NEAR 0 / FAIL 10 โ
โ stars โ โ
โ
โ
โ
โ
โ
โ
โโโ ๐ โ
โ hit max_tokens โ 6 โ
โ completion tokens avg โ 38,346 โ
โ completion tokens p50 โ 34,446 โ
โ completion tokens p90 โ 60,000 โ
โ completion tokens p99 โ 60,000 โ
โ elapsed avg โ 742.4s โ
โ TTFT avg โ 0.31s โ
โ aggregate gen tok/s โ 51.7 โ
โ mean per-request gen tok/s โ 51.5 โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Failed Final Answers
โญโโโโโฌโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ # โ C โ tokens โ score โ final answer โ
โโโโโโผโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 3 โ 5 โ 31,254 โ FAIL โ 45 (expected 48, got 45) | 45 โ
โ 5 โ 5 โ 60,000 โ FAIL โ 31 (expected 48, got 31) | What about $n=31 \ โ
โ 13 โ 5 โ 60,000 โ FAIL โ 33 (expected 48, got 33) | D=11: 11, 33, โ
โ 16 โ 5 โ 60,000 โ FAIL โ 4 (expected 48, got 4) | What about $m=25 \cdot 4 โ
โ โ โ โ โ = โ
โ 17 โ 5 โ 60,000 โ FAIL โ 38903 (expected 48, got 38903) | $38903 / โ
โ 18 โ 5 โ 17,288 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 19 โ 5 โ 19,525 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 24 โ 5 โ 60,000 โ FAIL โ 3 (expected 48, got 3) | Guest 7: 7 times. $7 โ
โ โ โ โ โ \equiv 1 \pmod 3$. R -> G -> R. โ
โ 27 โ 5 โ 26,455 โ FAIL โ 49 (expected 48, got 49) | 49 โ
โ 30 โ 5 โ 60,000 โ FAIL โ 5 (expected 48, got 5) | m=35: divs=[1,5,7,35]. โ
โ โ โ โ โ n=1: 1->0. n=5: โ
โฐโโโโโดโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Interpretation: EXACT means the final parsed number matches the expected answer.
FAIL means the answer was unparseable or a different number. The 10-slot quality
bar is a rounded distribution: โ
=EXACT, โ=FAIL.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/hotelmax-fp8-c5-r30.json
4. KLD vs BF16 reference
5 cold-reload runs per dtype (fresh container each run), 2,047 scored positions,
WikiText-2 @ ctx 2048, speculative_config=None. Lower is better.
| run | nvfp4_ds_mla | fp8 |
|---|---|---|
| 1 | 0.115928 | 0.100616 |
| 2 | 0.115569 | 0.102350 |
| 3 | 0.117059 | 0.100503 |
| 4 | 0.115955 | 0.100163 |
| 5 | 0.116181 | 0.102360 |
| mean | 0.116138 | 0.101198 |
| sd | 0.000559 | 0.001069 |
fp8 shows 0.0149 (-12.9%) lower divergence from the BF16 reference. nvfp4_ds_mla remains the shipped default because it costs roughly half the KV bytes per token โ that is what provides the context headroom at this preset โ and it matched or beat fp8 on the retrieval and ledger tasks.
This is the honest measure of what quantization costs: divergence is small but not zero.
5. Prefill throughput
--standalone-prefill --prefill-only --prefill-contexts 8k,64k,128k.
| Context | nvfp4 TTFT | nvfp4 tok/s | fp8 TTFT | fp8 tok/s |
|---|---|---|---|---|
| 8k | 3.21 s | 2,551 | 3.08 s | 2,660 |
| 64k | 32.96 s | 1,957 | 33.80 s | 1,909 |
| 128k | 70.31 s | 1,833 | 72.78 s | 1,771 |
Prefill โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1, 2, 4, 8, 16, 32, 64, 128] โ
โ Decode contexts: ['0', '16k', '32k', '64k', '128k'] โ
โ Decode: skipped (--prefill-only) | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill-only: standalone cold profile (auto) | Sustained decode: 0 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: standalone cold profile ['8k', '64k', '128k']
Calibrating padding text (run=zamzztpcfkqj, up to 128k)...
Token targeting: single-point estimate from 8k (use --token-targeting exact
for /tokenize binary search)
Calibrated: 6.18 chars/token (cached, source=8k)
8k: 50,601 chars (~8,191 tokens)
16k: 101,202 chars (~16,383 tokens)
32k: 202,404 chars (~32,767 tokens)
64k: 404,809 chars (~65,535 tokens)
128k: 809,618 chars (~131,071 tokens)
Done.
llm-decode-bench v0.4.29
Prefill Speed (C=1, client ISL / TTFT)
PCIe rx/tx
Context Tokens TTFT (s) Client tok/s Server tok/s avg N
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
8k 8,202 3.21 2,551 2,565 (2) 21578/21298 2
64k 64,516 32.96 1,957 1,963 (1) 77950/79445 1
128k 128,890 70.31 1,833 1,838 (1) 88481/90763 1
Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the
prefix-cache scout request that decode needs anyway. Server tok/s is optional
Prometheus validation when the engine exports prefill counters and the exact
counter delta is uncontaminated; for vLLM this uses newly computed KV tokens,
not request prompt tokens.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/prefill-nvfp4_ds_mla.json
Prefill โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1, 2, 4, 8, 16, 32, 64, 128] โ
โ Decode contexts: ['0', '16k', '32k', '64k', '128k'] โ
โ Decode: skipped (--prefill-only) | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill-only: standalone cold profile (auto) | Sustained decode: 0 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: standalone cold profile ['8k', '64k', '128k']
Calibrating padding text (run=vzteflzjnjfk, up to 128k)...
Token targeting: single-point estimate from 8k (use --token-targeting exact
for /tokenize binary search)
Calibrated: 6.18 chars/token (cached, source=8k)
8k: 50,601 chars (~8,191 tokens)
16k: 101,202 chars (~16,383 tokens)
32k: 202,404 chars (~32,767 tokens)
64k: 404,809 chars (~65,535 tokens)
128k: 809,618 chars (~131,071 tokens)
Done.
llm-decode-bench v0.4.29
Prefill Speed (C=1, client ISL / TTFT)
PCIe rx/tx
Context Tokens TTFT (s) Client tok/s Server tok/s avg N
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
8k 8,201 3.08 2,660 2,674 (2) 25261/22621 2
64k 64,515 33.80 1,909 1,915 (1) 80688/78187 1
128k 128,889 72.78 1,771 1,776 (1) 85527/85895 1
Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the
prefix-cache scout request that decode needs anyway. Server tok/s is optional
Prometheus validation when the engine exports prefill counters and the exact
counter delta is uncontaminated; for vLLM this uses newly computed KV tokens,
not request prompt tokens.
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/prefill-fp8.json
6. Decode throughput, C1-C8
--concurrency 1,2,3,4,5,6,7,8 --contexts 0 --duration 20 --max-tokens 8192 --skip-prefill --temperature 0.
Aggregate tokens/sec:
| concurrency | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| nvfp4_ds_mla | 87.5 | 147.4 | 183.7 | 219.3 | 242.4 | 274.9 | 291.1 | 308.1 |
| fp8 | 86.6 | 143.3 | 184.7 | 217.2 | 241.2 | 270.3 | 289.6 | 312.2 |
Per-user decode (nvfp4): 87.5 / 73.7 / 61.2 / 54.8 / 48.5 / 45.8 / 41.6 / 38.5 tok/s โ 3.5x
aggregate scaling C1 to C8. Single-stream figures are for the shipped lossless configuration
(VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE=0); the lossy alternative trades correctness for about
+10 tok/s at C1 and is deliberately disabled.
Decode C1-C8 โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1, 2, 3, 4, 5, 6, 7, 8] โ
โ Decode contexts: ['0'] โ
โ Duration: 20.0s per decode test | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill: skipped | Sustained decode: 8 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.
llm-decode-bench v0.4.29
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 2 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Sustained Decode โ
โ Steady-state decode throughput after the engine has admitted the requested โ
โ concurrency and passed warmup. Use this as the main tuning/regression signal โ
โ for kernels, NCCL, DCP, MTP, and scheduler changes. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate tok/s + TTFT/ITL
โญโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโฌโโโโโโโโฎ
โ ctx \ โ โ โ โ โ โ โ โ โ
โ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโผโโโโโโโโค
โ 0 โ 87.5 โ 147.4 โ 183.7 โ 219.3 โ 242.4 โ 274.9 โ 291.1 โ 308.1 โ
โ โ 161/11 โ 253/13 โ 363/16 โ 398/18 โ 442/20 โ 472/โฆ โ 499/23 โ 542/โฆ โ
โฐโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโดโโโโโโโโฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default
(continuous completion_tokens when the server supports it). Prometheus is kept
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโค
โ 0 โ 87.5 โ 73.7 โ 61.2 โ 54.8 โ 48.5 โ 45.8 โ 41.6 โ 38.5 โ
โฐโโโโโโโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโฏ
Client request latency: p50 / p90 ms
โญโโโโโโโโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโค
โ 0 โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ
โฐโโโโโโโโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc
coordinate. ITL is computed from observed generated tokens, including streams
stopped at the measurement boundary; a missing ITL means no stream produced at
least two measured output tokens. Per-request tok/s and request latency are
shown in separate per-cell matrices. Completion/sample counts and full
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate
tok/s remains the primary throughput signal.
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary
โญโโโโฌโโฌโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโฌโโโโโโฌโโโโโโโฌโโโโโโฌโโโโโโโโโโโโโโโโฎ
โ โฆ โ โ mode โ GPU avg/โฆ โ Mem โฆ โ W avg/โฆ โ T โฆ โ CPUโฆ โ VRโฆ โ PCIe rx/tx aโฆ โ
โโโโโผโโผโโโโโโโโผโโโโโโโโโโโโผโโโโโโโโผโโโโโโโโโโผโโโโโโผโโโโโโโผโโโโโโผโโโโโโโโโโโโโโโโค
โ 0 โ โ sustโฆ โ 100/100% โ 48% โ 1067/1โฆ โ 86C โ 73C โ 94โฆ โ 10945/10913 โ
โ 0 โ โ sustโฆ โ 100/100% โ 49% โ 1104/1โฆ โ 89C โ 73C โ 94โฆ โ 17508/17924 โ
โ 0 โ โ sustโฆ โ 100/100% โ 48% โ 1126/1โฆ โ 90C โ 73C โ 94โฆ โ 22034/21826 โ
โ 0 โ โ sustโฆ โ 100/100% โ 49% โ 1136/1โฆ โ 89C โ 73C โ 94โฆ โ 24354/24152 โ
โ 0 โ โ sustโฆ โ 99/99% โ 47% โ 1137/1โฆ โ 90C โ 73C โ 94โฆ โ 26736/26290 โ
โ 0 โ โ sustโฆ โ 99/99% โ 47% โ 1141/1โฆ โ 90C โ 73C โ 94โฆ โ 29873/29618 โ
โ 0 โ โ sustโฆ โ 99/99% โ 47% โ 1144/1โฆ โ 90C โ 79C โ 94โฆ โ 31554/31721 โ
โ 0 โ โ sustโฆ โ 99/100% โ 46% โ 1144/1โฆ โ 90C โ 73C โ 94โฆ โ 33035/33073 โ
โฐโโโโดโโดโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโดโโโโโโโโโโดโโโโโโดโโโโโโโดโโโโโโดโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,062 W | max 1,147 W | limit 1,200 W | over 3m 51s | 97 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Hardware summary is sampled from nvidia-smi during the measured part of each
cell. Whole-run GPU power is the sampled sum of GPU power draw across the
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is
a coarse live diagnostic, not a per-kernel NCCL profiler.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 3 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Burst / E2E Decode โ
โ Not run. Re-run with --run-burst to append a finite client-facing request โ
โ burst after Sustained Decode. This is intentionally disabled by default โ
โ because it adds another full decode matrix. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Primary Summary โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Primary matrices repeated last so the important numbers are visible without โ
โ scrolling back through diagnostics. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate decode tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโค
โ 0 โ 87.5 โ 147.4 โ 183.7 โ 219.3 โ 242.4 โ 274.9 โ 291.1 โ 308.1 โ
โฐโโโโโโโโโโโโโดโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโฏ
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1c8-nvfp4_ds_mla.json
Decode C1-C8 โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1, 2, 3, 4, 5, 6, 7, 8] โ
โ Decode contexts: ['0'] โ
โ Duration: 20.0s per decode test | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill: skipped | Sustained decode: 8 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.
llm-decode-bench v0.4.29
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 2 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Sustained Decode โ
โ Steady-state decode throughput after the engine has admitted the requested โ
โ concurrency and passed warmup. Use this as the main tuning/regression signal โ
โ for kernels, NCCL, DCP, MTP, and scheduler changes. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate tok/s + TTFT/ITL
โญโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโฌโโโโโโโโฎ
โ ctx \ โ โ โ โ โ โ โ โ โ
โ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโผโโโโโโโโโผโโโโโโโโค
โ 0 โ 86.6 โ 143.3 โ 184.7 โ 217.2 โ 241.2 โ 270.3 โ 289.6 โ 312.2 โ
โ โ 153/11 โ 244/14 โ 363/16 โ 393/18 โ 443/20 โ 474/โฆ โ 510/23 โ 538/โฆ โ
โฐโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโโดโโโโโโโโฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default
(continuous completion_tokens when the server supports it). Prometheus is kept
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโผโโโโโโโค
โ 0 โ 86.6 โ 71.7 โ 61.6 โ 54.3 โ 48.2 โ 45.0 โ 41.4 โ 39.0 โ
โฐโโโโโโโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโดโโโโโโโฏ
Client request latency: p50 / p90 ms
โญโโโโโโโโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฌโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโผโโโโโโค
โ 0 โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ โ/โ โ
โฐโโโโโโโโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโดโโโโโโฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc
coordinate. ITL is computed from observed generated tokens, including streams
stopped at the measurement boundary; a missing ITL means no stream produced at
least two measured output tokens. Per-request tok/s and request latency are
shown in separate per-cell matrices. Completion/sample counts and full
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate
tok/s remains the primary throughput signal.
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary
โญโโโโฌโโฌโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโฌโโโโโโฌโโโโโโโฌโโโโโโฌโโโโโโโโโโโโโโโโฎ
โ โฆ โ โ mode โ GPU avg/โฆ โ Mem โฆ โ W avg/โฆ โ T โฆ โ CPUโฆ โ VRโฆ โ PCIe rx/tx aโฆ โ
โโโโโผโโผโโโโโโโโผโโโโโโโโโโโโผโโโโโโโโผโโโโโโโโโโผโโโโโโผโโโโโโโผโโโโโโผโโโโโโโโโโโโโโโโค
โ 0 โ โ sustโฆ โ 100/100% โ 47% โ 1067/1โฆ โ 86C โ 74C โ 95โฆ โ 10754/10642 โ
โ 0 โ โ sustโฆ โ 100/100% โ 48% โ 1100/1โฆ โ 88C โ 73C โ 95โฆ โ 17246/17439 โ
โ 0 โ โ sustโฆ โ 100/100% โ 47% โ 1117/1โฆ โ 89C โ 73C โ 95โฆ โ 21808/21260 โ
โ 0 โ โ sustโฆ โ 100/100% โ 47% โ 1126/1โฆ โ 90C โ 73C โ 95โฆ โ 24263/24104 โ
โ 0 โ โ sustโฆ โ 99/99% โ 46% โ 1131/1โฆ โ 90C โ 74C โ 95โฆ โ 25999/25642 โ
โ 0 โ โ sustโฆ โ 99/99% โ 46% โ 1134/1โฆ โ 90C โ 74C โ 95โฆ โ 28964/28378 โ
โ 0 โ โ sustโฆ โ 99/100% โ 46% โ 1138/1โฆ โ 90C โ 74C โ 95โฆ โ 30996/31210 โ
โ 0 โ โ sustโฆ โ 99/100% โ 45% โ 1137/1โฆ โ 90C โ 73C โ 95โฆ โ 33155/32814 โ
โฐโโโโดโโดโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโดโโโโโโโโโโดโโโโโโดโโโโโโโดโโโโโโดโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 1,060 W | max 1,139 W | limit 1,200 W | over 3m 50s | 96 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Hardware summary is sampled from nvidia-smi during the measured part of each
cell. Whole-run GPU power is the sampled sum of GPU power draw across the
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is
a coarse live diagnostic, not a per-kernel NCCL profiler.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 3 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Burst / E2E Decode โ
โ Not run. Re-run with --run-burst to append a finite client-facing request โ
โ burst after Sustained Decode. This is intentionally disabled by default โ
โ because it adds another full decode matrix. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Primary Summary โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Primary matrices repeated last so the important numbers are visible without โ
โ scrolling back through diagnostics. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate decode tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฌโโโโโโโโฎ
โ ctx \ conc โ 1 โ 2 โ 3 โ 4 โ 5 โ 6 โ 7 โ 8 โ
โโโโโโโโโโโโโโผโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโผโโโโโโโโค
โ 0 โ 86.6 โ 143.3 โ 184.7 โ 217.2 โ 241.2 โ 270.3 โ 289.6 โ 312.2 โ
โฐโโโโโโโโโโโโโดโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโโโฏ
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1c8-fp8.json
Decode C1 dedicated 30 s โ nvfp4_ds_mla log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1] โ
โ Decode contexts: ['0'] โ
โ Duration: 30.0s per decode test | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill: skipped | Sustained decode: 1 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.
llm-decode-bench v0.4.29
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 2 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Sustained Decode โ
โ Steady-state decode throughput after the engine has admitted the requested โ
โ concurrency and passed warmup. Use this as the main tuning/regression signal โ
โ for kernels, NCCL, DCP, MTP, and scheduler changes. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate tok/s + TTFT/ITL
โญโโโโโโโโโโโโโฌโโโโโโโโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโโค
โ 0 โ 83.5 160/12 โ
โฐโโโโโโโโโโโโโดโโโโโโโโโโโโโโฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default
(continuous completion_tokens when the server supports it). Prometheus is kept
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโค
โ 0 โ 83.5 โ
โฐโโโโโโโโโโโโโดโโโโโโโฏ
Client request
latency: p50 / p90
ms
โญโโโโโโโโโโโโโฌโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโค
โ 0 โ โ/โ โ
โฐโโโโโโโโโโโโโดโโโโโโฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc
coordinate. ITL is computed from observed generated tokens, including streams
stopped at the measurement boundary; a missing ITL means no stream produced at
least two measured output tokens. Per-request tok/s and request latency are
shown in separate per-cell matrices. Completion/sample counts and full
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate
tok/s remains the primary throughput signal.
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary
โญโโโโฌโโฌโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโฌโโโโโโฌโโโโโโโฌโโโโโโฌโโโโโโโโโโโโโโโโฎ
โ โฆ โ โ mode โ GPU avg/โฆ โ Mem โฆ โ W avg/โฆ โ T โฆ โ CPUโฆ โ VRโฆ โ PCIe rx/tx aโฆ โ
โโโโโผโโผโโโโโโโโผโโโโโโโโโโโโผโโโโโโโโผโโโโโโโโโโผโโโโโโผโโโโโโโผโโโโโโผโโโโโโโโโโโโโโโโค
โ 0 โ โ sustโฆ โ 100/100% โ 47% โ 1071/1โฆ โ 86C โ 74C โ 94โฆ โ 10791/10730 โ
โฐโโโโดโโดโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโดโโโโโโโโโโดโโโโโโดโโโโโโโดโโโโโโดโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 996 W | max 1,099 W | limit 1,200 W | over 48s | 21 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Hardware summary is sampled from nvidia-smi during the measured part of each
cell. Whole-run GPU power is the sampled sum of GPU power draw across the
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is
a coarse live diagnostic, not a per-kernel NCCL profiler.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 3 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Burst / E2E Decode โ
โ Not run. Re-run with --run-burst to append a finite client-facing request โ
โ burst after Sustained Decode. This is intentionally disabled by default โ
โ because it adds another full decode matrix. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Primary Summary โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Primary matrices repeated last so the important numbers are visible without โ
โ scrolling back through diagnostics. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate decode
tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโค
โ 0 โ 83.5 โ
โฐโโโโโโโโโโโโโดโโโโโโโฏ
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1-nvfp4_ds_mla.json
Decode C1 dedicated 30 s โ fp8 log
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Configuration โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ LLM Inference Benchmark โ
โ Model: GLM-5.2-EXL3-TR3-3.0bpw @ localhost:8000 โ
โ Decode concurrency: [1] โ
โ Decode contexts: ['0'] โ
โ Duration: 30.0s per decode test | Max tokens: 8192 โ
โ Pre-decode warmup: C=1 max-runnable context for 3s โ
โ Prefill: skipped | Sustained decode: 1 cells โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Engine: vLLM
0.11.2.dev280+gilded.gnosis.v20.vllm3e731bc.si1a88b38.fi801d57a.cu132.20260722
Models: ['GLM-5.2-EXL3-TR3-3.0bpw']
KV cache budget (vLLM metrics): 262,144 tokens (1024 blocks ร 64; local 65,536 ร
CP 4; CP source: local process)
Model context length: 262,144 tokens
Prefill tests: skipped
Done.
llm-decode-bench v0.4.29
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 2 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Sustained Decode โ
โ Steady-state decode throughput after the engine has admitted the requested โ
โ concurrency and passed warmup. Use this as the main tuning/regression signal โ
โ for kernels, NCCL, DCP, MTP, and scheduler changes. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate tok/s + TTFT/ITL
โญโโโโโโโโโโโโโฌโโโโโโโโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโโค
โ 0 โ 82.4 153/12 โ
โฐโโโโโโโโโโโโโดโโโโโโโโโโโโโโฏ
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default
(continuous completion_tokens when the server supports it). Prometheus is kept
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโค
โ 0 โ 82.4 โ
โฐโโโโโโโโโโโโโดโโโโโโโฏ
Client request
latency: p50 / p90
ms
โญโโโโโโโโโโโโโฌโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโค
โ 0 โ โ/โ โ
โฐโโโโโโโโโโโโโดโโโโโโฏ
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc
coordinate. ITL is computed from observed generated tokens, including streams
stopped at the measurement boundary; a missing ITL means no stream produced at
least two measured output tokens. Per-request tok/s and request latency are
shown in separate per-cell matrices. Completion/sample counts and full
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate
tok/s remains the primary throughput signal.
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary
โญโโโโฌโโฌโโโโโโโโฌโโโโโโโโโโโโฌโโโโโโโโฌโโโโโโโโโโฌโโโโโโฌโโโโโโโฌโโโโโโฌโโโโโโโโโโโโโโโโฎ
โ โฆ โ โ mode โ GPU avg/โฆ โ Mem โฆ โ W avg/โฆ โ T โฆ โ CPUโฆ โ VRโฆ โ PCIe rx/tx aโฆ โ
โโโโโผโโผโโโโโโโโผโโโโโโโโโโโโผโโโโโโโโผโโโโโโโโโโผโโโโโโผโโโโโโโผโโโโโโผโโโโโโโโโโโโโโโโค
โ 0 โ โ sustโฆ โ 100/100% โ 46% โ 1066/1โฆ โ 86C โ 75C โ 95โฆ โ 10586/10558 โ
โฐโโโโดโโดโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโดโโโโโโโโโโดโโโโโโดโโโโโโโดโโโโโโดโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโ Whole-run GPU Power โโโโโโโโโโโโโโโโโโโโโโโฎ
โ avg 993 W | max 1,099 W | limit 1,200 W | over 48s | 21 samples โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Hardware summary is sampled from nvidia-smi during the measured part of each
cell. Whole-run GPU power is the sampled sum of GPU power draw across the
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is
a coarse live diagnostic, not a per-kernel NCCL profiler.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Phase 3 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Burst / E2E Decode โ
โ Not run. Re-run with --run-burst to append a finite client-facing request โ
โ burst after Sustained Decode. This is intentionally disabled by default โ
โ because it adds another full decode matrix. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Primary Summary โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ Primary matrices repeated last so the important numbers are visible without โ
โ scrolling back through diagnostics. โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Aggregate decode
tok/s
โญโโโโโโโโโโโโโฌโโโโโโโฎ
โ ctx \ conc โ 1 โ
โโโโโโโโโโโโโโผโโโโโโโค
โ 0 โ 82.4 โ
โฐโโโโโโโโโโโโโดโโโโโโโฏ
Results saved to
/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d1737
7be8eb/scratchpad/testsuite/results2/decode-c1-fp8.json
7. Power
Whole-run GPU power across the 4-GPU set (1,200 W aggregate limit):
| Run | avg | max | duration |
|---|---|---|---|
| LAVD nvfp4 | 1,130 W | 1,143 W | 19m 20s |
| LAVD fp8 | 1,127 W | 1,138 W | 19m 45s |
| Estonia nvfp4 | 1,105 W | 1,124 W | 10m 56s |
| Estonia fp8 | 1,101 W | 1,127 W | 14m 28s |
Sustained draw is 92-94% of the aggregate limit. Two of the four cards are 600 W-capable parts software-limited to 300 W, so single-stream decode is partly bounded by the slowest rank's clock.
8. Independent third-party evaluation
Run and published by malaiwah/glm52-exl3-vast
on 2026-07-23/24, independently of this repository's author. The original write-up, the client
harness, and the raw per-question summary are mirrored under
independent-eval/.
Results (pass@1, pooled across repeats)
| Benchmark | Questions x repeats = n | EXL3 3.0bpw | GLM-5.2 BF16 (Z.ai published) | ~95% CI |
|---|---|---|---|---|
| AIME 2026 | 30 x 4 = 120 | 99.2 | 99.2 | ยฑ1.6 |
| HMMT Feb 2026 | 33 x 4 = 132 | 95.5 | 92.5 | ยฑ3.6 |
| GPQA Diamond | 198 x 2 = 396 | 91.4 | 91.2 | ยฑ2.8 |
All three land within sampling noise of the BF16 reference โ no measurable reasoning degradation was detected at 3.0 bits per weight.
GPQA per-question stability across its 2 repeats: 173 of 198 questions correct both times, 16 split 1-of-2, 9 wrong both times. Across all 396 generations: 0 truncations, 0 errors, 1 sample with no extractable answer, 10,561 avg completion tokens, 15.9 h wall time.
How it was run
- Sampling matched to Z.ai's published eval settings:
temperature=1.0,top_p=0.95 - Max generation: 163,840 tokens (math), 131,072 (GPQA). Zero truncations occurred.
- No thinking-effort override โ server default reasoning mode
- Math prompt: Z.ai's
Explanation: / Exact Answer: / Confidence:system prompt - Math datasets:
MathArena/aime_2026,MathArena/hmmt_feb_2026 - Math grading:
math-verifysymbolic equivalence; fallback chainExact Answer:line -> last \boxed{} -> none - GPQA:
Idavidrein/gpqa(gpqa_diamond), simple-evals / Artificial-Analysis MCQ template, options deterministically shuffled per (question, repeat), regex letter extraction - Client: async Python harness, 32 concurrent requests (16 GPQA + 8 AIME + 8 HMMT) with all three benchmarks running simultaneously; ~65 tok/s aggregate under that mixed long-reasoning load; 8.63M completion tokens over ~16 h
- pass@1 computed over all repeats pooled
How this run differed from the shipped preset
Reported by the runner as fp8 KV cache, max_model_len 524288, server version
0.17.0rc1.dev4499+g60c82d972 โ corresponding to the earlier image tag
v1-gg-60c82d972-spi1937274-cu132-sm120a rather than the published v20-gg6722c1d-si1a88b38.
Same model weights, same compose and server script, MTP-3 enabled (speculative decoding affects
throughput, not the output distribution).
Reading these numbers fairly
- The BF16 column is Z.ai's published figures, not a re-measurement on this harness. The math
grading used
math-verifysymbolic equivalence rather than Z.ai's GPT-5.5 judge. This is therefore measured-versus-published, not a controlled head-to-head. - HMMT +3.0 over BF16 should not be read as the quant beating full precision โ a quantization cannot exceed its source in expectation. With 33 questions, a ยฑ3.6 interval and a different grader, that gap is noise plus methodology.
- Confidence intervals are simple binomial approximations. Repeats of the same question are correlated, so true intervals are somewhat wider.
- AIME and HMMT are small sets (30 and 33 questions). GPQA Diamond at 198 x 2 is the most statistically solid of the three.
Known reproducible quirks (seen on multiple quants, likely model-level)
- HMMT Q20: a common reasoning path converges on
1100where the gold answer is20460. Reproduced across different quantizations of this base model; this quant scored 2 of 4 repeats. - GPQA idx 79 (dataset order): triggers unusually long reasoning chains.
Incident note from the runner
Three requests stalled mid-run on dropped server connections โ client sockets stayed ESTABLISHED while the server no longer tracked the request. All three hit the client read timeout, auto-retried and completed. Zero lost or errored samples in the final data. Suggested hardening: TCP keepalives plus a tighter per-request timeout.
Reproduce with the mirrored harness (point --base-url at any OpenAI-compatible endpoint):
python independent-eval/mathbench.py --dataset MathArena/aime_2026 --repeats 4 --concurrency 8
python independent-eval/mathbench.py --dataset MathArena/hmmt_feb_2026 --repeats 4 --concurrency 8
python independent-eval/gpqa_bench.py --repeats 2 --concurrency 16
Reproducing the serving benchmarks
Note (2026-07-25): the
docker-compose.ymlandserver.shembedded further below reproduce the BF16-MTP Sections 1-8. The repo's liveserver.sh/docker-compose.ymlare the tr3-MTP build (v21 image,NUM_GPU_BLOCKS_OVERRIDEempty โ auto-profile,VLLM_EXL3_TRELLIS_MIN_M=1,MAX_MODEL_LEN=524288);./server.sh startbelow pulls and runs that current preset.
hf download brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw --local-dir "$HOME/models/GLM-5.2-EXL3-TR3-3.0bpw"
cd "$HOME/models/GLM-5.2-EXL3-TR3-3.0bpw"
chmod +x server.sh && ./server.sh start
Full docker-compose.yml (all serve flags)
services:
glm52:
image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff}
container_name: glm52-exl3-sparkinfer
ports:
- "${BIND_ADDRESS:-127.0.0.1}:${PORT:-8000}:8000"
gpus: all
shm_size: "32g"
ipc: host
ulimits:
memlock: -1
nofile: 1048576
environment:
CUDA_VISIBLE_DEVICES: "${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
CUDA_DEVICE_ORDER: PCI_BUS_ID
CUDA_DEVICE_MAX_CONNECTIONS: "32"
CUTE_DSL_ARCH: sm_120a
TORCH_CUDA_ARCH_LIST: 12.0a
FLASHINFER_CUDA_ARCH_LIST: 12.0f
FLASHINFER_DISABLE_VERSION_CHECK: "1"
OMP_NUM_THREADS: "16"
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
SAFETENSORS_FAST_GPU: "1"
NCCL_IB_DISABLE: "1"
NCCL_P2P_LEVEL: SYS
NCCL_PROTO: LL,LL128,Simple
VLLM_USE_FLASHINFER_SAMPLER: "1"
VLLM_USE_B12X_FP8_GEMM: "1"
VLLM_USE_B12X_SPARSE_INDEXER: "1"
VLLM_USE_B12X_MOE: "1"
VLLM_USE_V2_MODEL_RUNNER: "1"
VLLM_ENABLE_PCIE_ALLREDUCE: "1"
VLLM_PCIE_ALLREDUCE_BACKEND: b12x
VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE: "${VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE:-64KB}"
VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE: "${VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE:-84KB}"
VLLM_PCIE_DMA_FP8: ag
B12X_PCIE_DMA_FP8: ag
VLLM_CPP_AR_1STAGE_NCCL_CUTOFF: 56KB
VLLM_CPP_AR_IGNORE_CUTOFF_MAX_ROWS: "0"
VLLM_RTX6K_FUSED_ALLREDUCE_ADD: "0"
VLLM_RTX6K_FUSED_ALLREDUCE_ADD_END_BARRIER: "0"
VLLM_USE_AOT_COMPILE: "1"
VLLM_USE_BREAKABLE_CUDAGRAPH: "0"
VLLM_USE_FUSED_MOE_GROUPED_TOPK: "1"
VLLM_USE_B12X_MHC: "1"
B12X_MHC_MAX_TOKENS: "16384"
VLLM_USE_B12X_WO_PROJECTION: "1"
B12X_MLA_SM120_UNIFIED: "1"
B12X_DENSE_SPLITK_TURBO: "1"
B12X_W4A16_TC_DECODE: "1"
B12X_MOE_FORCE_A16: "1"
VLLM_DISABLE_SHARED_EXPERTS_STREAM: "${VLLM_DISABLE_SHARED_EXPERTS_STREAM:-1}"
VLLM_DISABLED_KERNELS: MarlinFP8ScaledMMLinearKernel
VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE: "${VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE:-0}"
VLLM_B12X_MLA_SPEC_DECODE_MAX_Q: "8"
VLLM_USE_B12X_DCP_A2A: "1"
VLLM_DCP_A2A_MAX_TOKENS: "16"
VLLM_DCP_A2A_LARGE_BACKEND: ag_rs
VLLM_DCP_GLOBAL_TOPK: "${VLLM_DCP_GLOBAL_TOPK:-1}"
VLLM_DCP_SHARD_DRAFT: "${VLLM_DCP_SHARD_DRAFT:-1}"
VLLM_DCP_QUERY_SPLIT: "0"
VLLM_B12X_MLA_CKV_GATHER: "1"
VLLM_B12X_MLA_CKV_GATHER_MIN_TOKENS: "512"
VLLM_B12X_MLA_CKV_GATHER_MAX_TOKENS: "16384"
ENABLE_MTP: "${ENABLE_MTP:-1}"
MTP_TOKENS: "${MTP_TOKENS:-3}"
MTP_DRAFT_SAMPLE_METHOD: "${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
ENABLE_ASYNC_SCHEDULING: "${ENABLE_ASYNC_SCHEDULING:-0}"
GLM52_INDEX_TOPK_PATTERN: "${GLM52_INDEX_TOPK_PATTERN:-FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS}"
NUM_GPU_BLOCKS_OVERRIDE: "${NUM_GPU_BLOCKS_OVERRIDE:-1024}"
MAX_NUM_BATCHED_TOKENS: "${MAX_NUM_BATCHED_TOKENS:-3072}"
VLLM_EXL3_TRELLIS_MIN_M: "4"
VLLM_EXL3_TRELLIS_MAX_M: "32"
VLLM_EXL3_TRELLIS_BLOCK_M: "8"
VLLM_EXL3_PREFILL_CHUNK: "128"
VLLM_CACHE_DIR: /cache/jit/vllm
TRITON_CACHE_DIR: /cache/jit/triton
TORCH_EXTENSIONS_DIR: /cache/jit/torch_extensions
TORCHINDUCTOR_CACHE_DIR: /cache/jit/torchinductor
FLASHINFER_WORKSPACE_BASE: /cache/jit/flashinfer
XDG_CACHE_HOME: /cache/jit
TVM_FFI_CACHE_DIR: /cache/jit/tvm-ffi
VLLM_MEMORY_PROFILE_INCLUDE_ATTN: "1"
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS: "1"
VLLM_DEBUG_WORKSPACE: "${VLLM_DEBUG_WORKSPACE:-0}"
volumes:
- ${MODEL_DIR:-/home/brandonmusic/models/GLM-5.2-EXL3-TR3-3.0bpw}:/model:ro
- ${CACHE_DIR:-/home/brandonmusic/.cache/glm52-tr3-release}:/cache:rw
entrypoint:
- /bin/bash
- -lc
command:
- |
unset NCCL_GRAPH_FILE NCCL_GRAPH_DUMP_FILE VLLM_B12X_MLA_EXTEND_MAX_CHUNKS
spec_args=()
if [[ "$${ENABLE_MTP:-1}" == "1" ]]; then
printf -v spec_config '{"method":"mtp","num_speculative_tokens":%s,"moe_backend":"triton","draft_sample_method":"%s"}' \
"$${MTP_TOKENS:-3}" "$${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
spec_args=(--speculative-config "$$spec_config")
fi
if [[ "$${ENABLE_ASYNC_SCHEDULING:-0}" == "1" ]]; then
async_args=(--async-scheduling)
else
async_args=(--no-async-scheduling)
fi
index_pattern="$${GLM52_INDEX_TOPK_PATTERN}"
if [[ "$${#index_pattern}" -ne 78 ]]; then
printf 'GLM-5.2 index_topk_pattern must cover all 78 layers (got %s)\n' "$${#index_pattern}" >&2
exit 2
fi
printf -v hf_overrides '{"use_index_cache":true,"index_topk_pattern":"%s"}' "$${index_pattern}"
block_args=()
if [[ -n "$${NUM_GPU_BLOCKS_OVERRIDE:-}" ]]; then
block_args=(--num-gpu-blocks-override "$${NUM_GPU_BLOCKS_OVERRIDE}")
fi
exec vllm serve /model \
--served-model-name GLM-5.2-EXL3-TR3-3.0bpw \
--host 0.0.0.0 --port 8000 --trust-remote-code \
--tensor-parallel-size 4 \
--decode-context-parallel-size 4 \
--dcp-comm-backend a2a \
--dcp-kv-cache-interleave-size ${DCP_KV_CACHE_INTERLEAVE_SIZE:-64} \
--seed 0 \
--quantization exl3 \
--kv-cache-dtype ${KV_CACHE_DTYPE:-nvfp4_ds_mla} \
--attention-backend B12X_MLA_SPARSE \
--moe-backend b12x \
--load-format safetensors \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[4,8,12,16,20,24,28,32],"custom_ops":["all"],"pass_config":{"fuse_allreduce_rms":true}}' \
--gpu-memory-utilization ${GPU_MEMORY_UTILIZATION:-0.96} \
--max-model-len ${MAX_MODEL_LEN:-262144} \
--max-num-seqs 8 \
--max-num-batched-tokens $${MAX_NUM_BATCHED_TOKENS:-3072} \
--max-cudagraph-capture-size 32 \
--enable-chunked-prefill \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--default-chat-template-kwargs '{"reasoning_effort":"high"}' \
--hf-overrides "$${hf_overrides}" \
"$${block_args[@]}" \
"$${async_args[@]}" \
"$${spec_args[@]}"
Full server.sh
#!/usr/bin/env bash
set -Eeuo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
export IMAGE="${IMAGE:-verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a@sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff}"
export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
export PORT="${PORT:-8000}"
export BIND_ADDRESS="${BIND_ADDRESS:-127.0.0.1}"
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.96}"
export MAX_MODEL_LEN="${MAX_MODEL_LEN:-262144}"
export DCP_KV_CACHE_INTERLEAVE_SIZE="${DCP_KV_CACHE_INTERLEAVE_SIZE:-64}"
export VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE="${VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE:-0}"
export VLLM_DCP_SHARD_DRAFT="${VLLM_DCP_SHARD_DRAFT:-1}"
export VLLM_DCP_GLOBAL_TOPK="${VLLM_DCP_GLOBAL_TOPK:-1}"
export VLLM_DISABLE_SHARED_EXPERTS_STREAM="${VLLM_DISABLE_SHARED_EXPERTS_STREAM:-1}"
export VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE="${VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE:-64KB}"
export VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE="${VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE:-84KB}"
export ENABLE_MTP="${ENABLE_MTP:-1}"
export MTP_TOKENS="${MTP_TOKENS:-3}"
export MTP_DRAFT_SAMPLE_METHOD="${MTP_DRAFT_SAMPLE_METHOD:-greedy}"
export ENABLE_ASYNC_SCHEDULING="${ENABLE_ASYNC_SCHEDULING:-0}"
export GLM52_INDEX_TOPK_PATTERN="${GLM52_INDEX_TOPK_PATTERN:-FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS}"
export NUM_GPU_BLOCKS_OVERRIDE="${NUM_GPU_BLOCKS_OVERRIDE:-1024}"
export MAX_NUM_BATCHED_TOKENS="${MAX_NUM_BATCHED_TOKENS:-3072}"
export COMPOSE_PROJECT_NAME="${COMPOSE_PROJECT_NAME:-glm52-exl3-sparkinfer}"
COMPOSE_FILE="${COMPOSE_FILE:-/tmp/claude-1000/-home-brandonmusic-KLC-SANDBOXES/50980f6d-56ae-4115-a6bf-0d17377be8eb/scratchpad/testsuite/config/docker-compose.yml}"
COMPOSE=(docker compose -f "$COMPOSE_FILE")
usage() {
cat <<'EOF'
Usage: ./server.sh [start|stop|restart|logs|status|pull]
Environment overrides:
IMAGE, MODEL_DIR, CACHE_DIR, PORT, BIND_ADDRESS, CUDA_VISIBLE_DEVICES,
GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, DCP_KV_CACHE_INTERLEAVE_SIZE,
VLLM_B12X_MLA_SPEC_EXTEND_AS_DECODE, VLLM_DCP_SHARD_DRAFT,
VLLM_DCP_GLOBAL_TOPK,
VLLM_DISABLE_SHARED_EXPERTS_STREAM,
VLLM_PCIE_ONESHOT_ALLREDUCE_MAX_SIZE,
VLLM_PCIE_ONESHOT_FUSED_ADD_RMS_NORM_MAX_SIZE,
ENABLE_MTP, MTP_TOKENS, MTP_DRAFT_SAMPLE_METHOD, ENABLE_ASYNC_SCHEDULING,
GLM52_INDEX_TOPK_PATTERN,
NUM_GPU_BLOCKS_OVERRIDE,
MAX_NUM_BATCHED_TOKENS,
COMPOSE_PROJECT_NAME, COMPOSE_FILE
EOF
}
require_runtime() {
command -v docker >/dev/null 2>&1 || {
echo "docker is required" >&2
exit 1
}
docker compose version >/dev/null
[[ -f "$COMPOSE_FILE" ]] || {
echo "Compose file not found: $COMPOSE_FILE" >&2
exit 1
}
}
require_model() {
[[ -f "$MODEL_DIR/config.json" ]] || {
echo "Model config not found: $MODEL_DIR/config.json" >&2
exit 1
}
[[ -f "$MODEL_DIR/model.safetensors.index.json" ]] || {
echo "Model index not found: $MODEL_DIR/model.safetensors.index.json" >&2
exit 1
}
mkdir -p "$CACHE_DIR"
}
action="${1:-start}"
require_runtime
case "$action" in
start)
require_model
docker pull "$IMAGE"
"${COMPOSE[@]}" up -d --force-recreate
echo "Starting on http://localhost:$PORT"
echo "Follow startup with: $0 logs"
;;
stop)
"${COMPOSE[@]}" down
;;
restart)
require_model
docker pull "$IMAGE"
"${COMPOSE[@]}" up -d --force-recreate
echo "Restarting on http://localhost:$PORT"
;;
logs)
"${COMPOSE[@]}" logs --tail 100 -f glm52
;;
status)
"${COMPOSE[@]}" ps
curl -fsS "http://localhost:$PORT/v1/models" || true
printf '\n'
;;
pull)
docker pull "$IMAGE"
;;
-h|--help|help)
usage
;;
*)
usage >&2
exit 2
;;
esac
Long-context needle-in-a-haystack (v28, 2026-07-26)
Run on 4x RTX PRO 6000 Blackwell (SM120), TP4 / DCP4, nvfp4_ds_mla KV,
MTP-3 draft at layer 78, VLLM_EXL3_TRELLIS_MIN_M=1, max-model-len 524288.
A unique authorization code is planted at three depths (0.1 / 0.5 / 0.9) in a
filler document and requested back with greedy decoding.
nvfp4_ds_mla KV
| target ctx | real prompt tokens | depth 0.1 | depth 0.5 | depth 0.9 |
|---|---|---|---|---|
| 8k | 4,844 | HIT | HIT | HIT |
| 32k | 19,316 | HIT | HIT | HIT |
| 65k | 39,188 | HIT | HIT | HIT |
| 128k | 77,168 | HIT | HIT | HIT |
| 200k | 199,783 | HIT | HIT | HIT |
| 300k | 299,648 | HIT | HIT | HIT |
| 400k | 399,512 | HIT | HIT | HIT |
| 480k | 479,396 | HIT | HIT | HIT |
24/24 needles recovered. GPU KV cache 959,744-998,400 tokens. Deepest 480k probe 315 s.
fp8 KV
Identical checkpoint, weights, draft and flags; only --kv-cache-dtype changed.
| target ctx | real prompt tokens | depth 0.1 | depth 0.5 | depth 0.9 |
|---|---|---|---|---|
| 8k | 8,010 | HIT | HIT | HIT |
| 65k | 64,965 | HIT | HIT | HIT |
| 128k | 127,854 | HIT | HIT | HIT |
| 200k | 199,784 | HIT | HIT | HIT |
| 300k | 299,648 | HIT | HIT | HIT |
| 480k | 479,396 | HIT | HIT | HIT |
18/18 needles recovered. GPU KV cache 648,192 tokens (8-bit vs 4-bit, so a
smaller pool than nvfp4_ds_mla at the same utilization). Deepest 480k probe
311 s.
Combined: 42/42 across both KV dtypes, three depths each, to ~480k real
prompt tokens -- within ~45k of the 524,288 max-model-len ceiling. No garbled
output and zero engine restarts in either lane.
Context: vLLM issue #183 reported that VLLM_EXL3_TRELLIS_MIN_M=1 silently
corrupts long-context output. That did not reproduce here in either KV dtype. Independently, the
fused Trellis MoE was verified bitwise-correct at m=1,2,3 at this exact geometry
(tile 64x256x64x256, block_size_m=8, capacity 32, topk=8) with the scratch
arena NaN-poisoned. Note that MIN_M=1 widens the Trellis window so the draft's
m=1..3 GEMMs stay on the fused, graph-capturable path; it is not a per-token
capture and carries no throughput penalty.
Boot without the MIN_M workaround (v29, 2026-07-27)
Historically, serving an EXL3 rank-sliced tr3 MTP draft required setting
VLLM_EXL3_TRELLIS_MIN_M=1 by hand; without it the engine could not start
(vLLM issue #183):
RuntimeError: EXL3 eager parity path entered during CUDA graph capture (m=3);
capture sizes must lie inside the Trellis window [4, 32]
v29 removes the requirement. The backend stamps each layer's draft/target role
at construction (runner_type == "draft"), where the vllm-config context is
live, and defaults draft layers' Trellis window to MIN_CAPTURABLE_TRELLIS_M=1
automatically. Target layers keep the historical default of 4; an explicit
VLLM_EXL3_TRELLIS_MIN_M still overrides both.
Validation on this rig (4x RTX PRO 6000 SM120, TP4/DCP4, tr3 MTP-78, MTP-3),
with VLLM_EXL3_TRELLIS_MIN_M entirely unset:
| gate | result |
|---|---|
| engine boot + serve | PASS (previously guaranteed startup failure) |
| capture-time window error in logs | 0 occurrences |
blank-env int('') startup crash |
0 occurrences (blank now means unset) |
| greedy inference | PASS |
| needle 8k/65k/128k x depths 0.1/0.5/0.9 | 9/9 recovered |
The compose files in this repo now leave VLLM_EXL3_TRELLIS_MIN_M unset by
default. Fix commits: vLLM PR #139 239ba678b5 + 796ea923f1.
v30 (2026-07-27): env-knob registration + sparkinfer PR#79 module
Delta vs v29 (which carries all correctness fixes): (1) the nine EXL3 env knobs are registered in vllm envs.py -- startup 'Unknown vLLM environment variable' warnings drop 15 -> 8 (the remaining 8 are base-runtime-owned), and the knobs join the torch.compile cache-key factors (one-time ~70 s recompile after changing one); (2) the SparkInfer wheel is rebuilt from the PR#49 branch rebased onto master AFTER PR#79 ("perf(pcie): add exact DCP top-k owner exchange"), so the CUDA-IPC owner-exchange module ships in the image. It is DORMANT here: PR#79 has zero overlap with the EXL3/MoE lane (only +1 line outside its new files), and the vLLM-side owner algorithm is not yet in this image's pinned base. Gates on the pinned v30: boot with VLLM_EXL3_TRELLIS_MIN_M unset PASS, 0 capture-window errors, warnings 15->8 confirmed, greedy inference PASS, tool-calls 4/4.
v31 (2026-07-27): unified v20 base refresh (SparkInfer c3828fd)
Base bump only on the SparkInfer axis (vLLM pin unchanged at 0c79e41, so the 13-file vLLM overlay is byte-identical to v30). Wheel rebuilt from PR#49 (11 commits) on the new integration pin c3828fd and verified a strict superset of the base's canonical SparkInfer (168/168 source files present). Gates on the pinned digest: boot with VLLM_EXL3_TRELLIS_MIN_M unset PASS, warnings 8, 0 capture errors, inference PASS, tool-calls 4/4, KV 963,840.
Correction note for v30: its wheel was built from SparkInfer master rather than the base's integration pin, so the pip install replaced the base's canonical SparkInfer with a tree missing the integration-only PCIe calibration commits. No effect on the published serving configs (they pin DCP controls explicitly, and calibration only engages on 'auto'), but helper/auto-calibration users should prefer v31.