# GLM-5.2 EXL3 release validation These artifacts were collected on four RTX PRO 6000 Blackwell 96 GB GPUs with the final release image, TP4 + DCP4 A2A, MTP1 greedy, asynchronous scheduling disabled, and NVFP4 DeepSeek-MLA KV cache unless a row says otherwise. Tested local image ID: ```text sha256:bfd6d6670db37b04e9cbef7375722e3f71d66745abf1714c05cc5b71fd126715 ``` Published immutable image: ```text verdictai/glm52-exl3-sparkinfer:v2-gg-1043999ab-spi879ca0ad-cu132-sm120a@sha256:bfd6d6670db37b04e9cbef7375722e3f71d66745abf1714c05cc5b71fd126715 ``` Source revisions: - Gilded Gnosis vLLM: `1043999ab5f13350aaacc188e713d6c0928387e9` - Sparkinfer: `879ca0ad878958a0fbb1d6a1393c1bb40ec54790` The vLLM banner in the logs retains the base package's `g60c82d972` version string. The `ai.verdict.vllm.revision` and `ai.verdict.sparkinfer.revision` OCI labels in `image-inspect.json` identify the source actually installed in the tested image. ## Correctness | Evaluation | Result | Concurrency | Max output | Cap hits | | --- | ---: | ---: | ---: | ---: | | LAVD pass 1 | 10/10 (5 exact, 5 near) | 5 | 20,000 | 0 | | LAVD pass 2 | 10/10 (3 exact, 7 near) | 5 | 20,000 | 0 | | Estonia | 5/5 pass | 5 | 5,000 | 0 | The two LAVD passes are independent full runs, for a combined 20/20 scorer pass. `near` is a passing LAVD score, not a failure. LAVD used temperature 0 and repetition penalty 1.15; Estonia used temperature 0 and repetition penalty 1.25. ## KV-cache KLD Each KLD figure is the mean of five independent runs against the same verified BF16 reference-logit artifact. Every run evaluated all 2,047 positions from a 2,048-token sample with TP4 + DCP4 A2A, the 78-layer sparse-index pattern, asynchronous scheduling disabled, and eager execution. KLD isolates model and KV-cache drift; the serving evaluations above validate MTP1 and CUDA graphs. The reference artifact SHA256 is `87f992a689c054a0548a4b3863da6c809f9239beacd5786d0401e45904fec063`. | KV cache | Mean KLD | Sample SD | Min | Max | Runs | | --- | ---: | ---: | ---: | ---: | ---: | | NVFP4 DeepSeek MLA | 0.1124021 | 0.0025948 | 0.1086084 | 0.1156108 | 5 | | FP8 | 0.1036666 | 0.0018374 | 0.1016535 | 0.1066077 | 5 | ## Prefill Cold standalone prefill uses `prompt_tokens / TTFT`; the client figure is the headline and the uncontaminated vLLM Prometheus counter is shown as validation. | Requested context | Prompt tokens | TTFT | Client tok/s | Server tok/s | | ---: | ---: | ---: | ---: | ---: | | 8K | 8,201 | 5.46 s | 1,502 | 1,507 | | 64K | 64,512 | 51.64 s | 1,249 | 1,252 | | 128K | 128,881 | 109.00 s | 1,182 | 1,185 | ## Sustained decode These are 20-second steady-state cells after warmup, at zero input context and with EOS ignored. Aggregate throughput comes from continuous OpenAI stream usage. There were no request errors, underfilled cells, or capacity-limited cells. A separate 30-second C1 run measured 48.5 tok/s. | Concurrency | Aggregate tok/s | Per-request tok/s | | ---: | ---: | ---: | | 1 | 48.9 | 48.9 | | 2 | 112.9 | 56.4 | | 3 | 154.0 | 51.3 | | 4 | 188.2 | 47.0 | | 5 | 218.2 | 43.6 | | 6 | 239.4 | 39.9 | | 7 | 253.9 | 36.3 | | 8 | 266.8 | 33.3 | The release preset allocates 524,288 logical KV-cache tokens: 2,048 blocks x 64 tokens x DCP4, with 131,072 tokens local to each DCP rank. This is the validated allocation and configured per-request context cap. ## Raw artifacts - [C1-C8 decode JSON](decode-c1-c8.json) and [Rich TUI log](decode-c1-c8.log) - [Dedicated C1 JSON](decode-c1-dedicated.json) and [Rich TUI log](decode-c1-dedicated.log) - [Prefill JSON](prefill-8k-64k-128k.json) and [Rich TUI log](prefill-8k-64k-128k.log) - [LAVD pass 1 JSON](lavd-c5-r10.json) and [Rich TUI log](lavd-c5-r10.log) - [LAVD pass 2 JSON](lavd-c5-r10-pass2.json) and [Rich TUI log](lavd-c5-r10-pass2.log) - [Estonia JSON](estonia-c5-r5.json) and [Rich TUI log](estonia-c5-r5.log) - [NVFP4 KLD summary](kld-nvfp4-dcp4-summary.json) - [FP8 KLD summary](kld-fp8-dcp4-summary.json) - [Exact tested-image inspection](image-inspect.json) At C8, the hottest observed GPU reached 90 C. NVIDIA hardware and software thermal-slowdown flags were inactive during the run.