Fidelity measurement: KL(reference || this quant) on frozen tokens, receipt-backed

#3
by malaiwah - opened

What this is. A third-party fidelity measurement of brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw @ f79c9167690ca705e877ae4dc55a841d1aae1247 (exl3-mcg, 3 bits per weight as declared) against its unquantized reference, made with quant-fidelity-suite. Measured by malaiwah, not by the model's author. Every number below is read from a sealed receipt named at the bottom.

KL(reference ‖ candidate), mean tokenwise, nats 0.09094554706733873
top-1 agreement 0.917361993160723
KL median / p95 / p99 / max 0.005630015746389436 / 0.41566289292853176 / 1.4783458837700123 / 11.277942933497664
reference root dataset malaiwah/glm52-fidelity-root-v1 @ 5977559307ee9fb7d6478e81a875faa10ffee9b8 (capture a544e029a0392c2a…)
panel panel--glm53.malaiwah.corpus5x5-v1, 25 contexts, 51175 scored positions
direction / vocabulary / accumulation KL(reference
method decode-and-run, weights only: exl3-trellis-tp-compose-to-bf16 -- each routed-expert module stored as tensor-parallel rank shards (the artifact's own hybrid_tr3_tail declares tp and the slicing axis) decoded to bf16 per rank on the capture device via exllamav3's transcribed codebooks (codebook mcg, declared 3.0 bits average) and composed into the whole weight in ascending rank order; non-routed tensors carried as shipped; same engine, schedule and device as the reference capture
caveat weights-only reconstruction on the HF transformers stack (the trellis payloads decoded to bf16 weights by this suite's transcription of exllamav3's codebooks, not by exllamav3 itself); the served kernel's fp16 activations and on-the-fly dequant are not in this number; a different panel and teacher than the author's own figure, so the two are not comparable. The registry files this row as advisory
determinism two fresh processes captured the candidate; both sealed captures carry content digest f485b230d4545f46… (self-comparison 0.0)
comparability class advisory

Scope (scope_digest): attn.o=native:bf16@16|attn.other=native:bf16@16|attn.qkv=native:bf16@16|embed_tokens=native:bf16@16|lm_head=native:bf16@16|mlp.down=native:bf16@16|mlp.gate=native:bf16@16|mlp.up=native:bf16@16|moe.experts=quantized:exl3-mcg@3|moe.router=native:mixed|moe.shared_expert=native:bf16@16|mtp=quantized:mixed|norm=native:bf16@16|head=native|kv=bf16

Disclosures on the comparison receipt:

  • native_head_replay (info): HEAD-1d: each side replayed through its own sealed head (reference a012be05e771, candidate a012be05e771); head error is inside the measurement, as under HEAD-2, and nothing is substituted. The heads are content-identical.
  • weights_reconstructed (caveat): candidate was captured from a WEIGHTS-ONLY RECONSTRUCTION: 230400 trellis payload group(s) were decoded to bf16 by engines/tools/exl3hf_surface.py:decode_payload_hf, this repository's transcription of exllamav3's mul1/mcg codebooks (exllamav3_ext/quant/codebook.cuh, pack.cu), before the transformers forward (method exl3-trellis-tp-compose-to-bf16; K histogram K3 x 230400). 57600 module(s) stored as tp=4 rank shards were composed in ascending rank order. the decoder has NOT been proven bitwise against exllamav3 itself (no engines/tools/layer-outer-evidence/exl3-decoder-parity-vs-exllamav3.json evidence); it is proven against in-house fp64 routes and real payloads only. the served exllamav3 kernel's fp16 activations and on-the-fly dequant are not in this number. The comparison is advisory.

What this number does not do. It is a same-lane distance from one reference capture on one panel. It does not rank this artifact against numbers measured on another panel, lane or reference, and a same-lane root does not retroactively upgrade rows measured against another teacher. Per-window scatter exceeds the gap between adjacent bit-widths; compare only within a group whose comparability keys match.

Receipts.

  • comparison receipt receipt_sha256 bab61f9e3a8ce550f236fc676a1e879cb905dc751904d935550c559e3c8850d7
  • root-qualification receipt receipt_sha256 6c3b31fc305a245c3eb8d28935f31075a6d75d8d097dc51051cb23f0e41623f7 (canonical dataset_sha256 889c5cb8087e2da8f656dddc2e5f7dcc27415c6689df15dc0da8531ce2354211, repeat 36cc76bf829fa22e63c752aaf6c784c8acd9653b23b5b14b3d64745551d9d6d0)
  • job job_id_full 42592a4a44a0835522938475405a7b61219f8264cdd6578f29154db821571ca7
  • candidate capture dataset published: malaiwah/glm52-fidelity-exl3-tr3-3.0bpw-brandonmusic-v1 @ e4345aa35c1afe5f3ff23aacd7ed35c87d5b0f66

Reproduce: fetch the two datasets named above and run fidelity-dataset compare --reference <root> --candidate <this> --own-heads; the receipt's estimator block is the exact recipe. Questions and corrections are welcome here; the registry files this as a third-party row (measured_by enumerated, never conflated with author-reported numbers).

Sign up or log in to comment