brandonmusic's picture
Publish validated v2 GG/Sparkinfer runtime
c7e592f verified
|
Raw History Blame Contribute Delete
4.17 kB

GLM-5.2 EXL3 release validation

These artifacts were collected on four RTX PRO 6000 Blackwell 96 GB GPUs with the final release image, TP4 + DCP4 A2A, MTP1 greedy, asynchronous scheduling disabled, and NVFP4 DeepSeek-MLA KV cache unless a row says otherwise.

Tested local image ID:

sha256:bfd6d6670db37b04e9cbef7375722e3f71d66745abf1714c05cc5b71fd126715

Published immutable image:

verdictai/glm52-exl3-sparkinfer:v2-gg-1043999ab-spi879ca0ad-cu132-sm120a@sha256:bfd6d6670db37b04e9cbef7375722e3f71d66745abf1714c05cc5b71fd126715

Source revisions:

  • Gilded Gnosis vLLM: 1043999ab5f13350aaacc188e713d6c0928387e9
  • Sparkinfer: 879ca0ad878958a0fbb1d6a1393c1bb40ec54790

The vLLM banner in the logs retains the base package's g60c82d972 version string. The ai.verdict.vllm.revision and ai.verdict.sparkinfer.revision OCI labels in image-inspect.json identify the source actually installed in the tested image.

Correctness

Evaluation Result Concurrency Max output Cap hits
LAVD pass 1 10/10 (5 exact, 5 near) 5 20,000 0
LAVD pass 2 10/10 (3 exact, 7 near) 5 20,000 0
Estonia 5/5 pass 5 5,000 0

The two LAVD passes are independent full runs, for a combined 20/20 scorer pass. near is a passing LAVD score, not a failure. LAVD used temperature 0 and repetition penalty 1.15; Estonia used temperature 0 and repetition penalty 1.25.

KV-cache KLD

Each KLD figure is the mean of five independent runs against the same verified BF16 reference-logit artifact. Every run evaluated all 2,047 positions from a 2,048-token sample with TP4 + DCP4 A2A, the 78-layer sparse-index pattern, asynchronous scheduling disabled, and eager execution. KLD isolates model and KV-cache drift; the serving evaluations above validate MTP1 and CUDA graphs. The reference artifact SHA256 is 87f992a689c054a0548a4b3863da6c809f9239beacd5786d0401e45904fec063.

KV cache Mean KLD Sample SD Min Max Runs
NVFP4 DeepSeek MLA 0.1124021 0.0025948 0.1086084 0.1156108 5
FP8 0.1036666 0.0018374 0.1016535 0.1066077 5

Prefill

Cold standalone prefill uses prompt_tokens / TTFT; the client figure is the headline and the uncontaminated vLLM Prometheus counter is shown as validation.

Requested context Prompt tokens TTFT Client tok/s Server tok/s
8K 8,201 5.46 s 1,502 1,507
64K 64,512 51.64 s 1,249 1,252
128K 128,881 109.00 s 1,182 1,185

Sustained decode

These are 20-second steady-state cells after warmup, at zero input context and with EOS ignored. Aggregate throughput comes from continuous OpenAI stream usage. There were no request errors, underfilled cells, or capacity-limited cells. A separate 30-second C1 run measured 48.5 tok/s.

Concurrency Aggregate tok/s Per-request tok/s
1 48.9 48.9
2 112.9 56.4
3 154.0 51.3
4 188.2 47.0
5 218.2 43.6
6 239.4 39.9
7 253.9 36.3
8 266.8 33.3

The release preset allocates 524,288 logical KV-cache tokens: 2,048 blocks x 64 tokens x DCP4, with 131,072 tokens local to each DCP rank. This is the validated allocation and configured per-request context cap.

Raw artifacts

At C8, the hottest observed GPU reached 90 C. NVIDIA hardware and software thermal-slowdown flags were inactive during the run.