Real KLD testing against Unsloth's gguf BF16 of Qwen 3.8-27B

#54
by mrumel - opened

Ternary Bonsai 2 27B vs Qwen3.8-27B quants: benchmark summary

Hardware: Tesla PG500-216 (V100-class, Volta, sm_70, 32 GB), Dell PowerEdge, Unraid
Runtime: PrismML llama.cpp fork, build bdc23b56b (10728), CUDA 12.8.1, CUDA_DOCKER_ARCH=70
Models tested:

  • Qwen3.8-27B-BF16 (Unsloth): reference
  • Qwen3.8-27B-Q8_0 (Unsloth)
  • Qwen3.8-27B-UD-Q4_K_XL (Unsloth)
  • Ternary-Bonsai-2-27B-PQ2_0 (PrismML): 2.13 bpw, group-128 ternary
  • Ternary-Bonsai-2-27B-PTQ1_0 (PrismML): 1.75 bpw dense ternary (speed only)

Summary

Q8_0 is effectively identical to BF16, and Q4_K_XL is an excellent quant (within about 1% on every measure). Bonsai PQ2_0 is about 38% faster at generation and uses 10 GB less VRAM than Q4_K_XL, but it diverges from the original model far more than normal quantization does: roughly 40× the Q4's KL divergence, and it picks a different top token about 1 in 4.5 times.

Bonsai is a retrained ternary model, not a rounded copy of Qwen3.8-27B, so these numbers show it behaves differently from the original, not necessarily worse on real tasks. Chat-template task evals (gsm8k, mmlu_pro) are the remaining test.

Current pick: Q4_K_XL as the default for fidelity; Bonsai PQ2_0 only where speed or VRAM matters more than matching Qwen's behavior. Avoid PTQ1_0 on Volta.

Speed (llama-bench)

-ngl 99 -fa 1 -p 512 -n 128 -d 0,8192 -r 5

Model Size pp512 (t/s) tg128 (t/s) pp512 @ 8K tg128 @ 8K
Qwen3.8 Q4_K_XL 16.34 GiB 751.6 ± 37.2 36.09 ± 0.06 654.2 ± 33.0 34.04 ± 0.17
Bonsai PQ2_0 6.70 GiB 700.7 ± 46.2 49.87 ± 0.28 634.3 ± 16.3 46.49 ± 0.21
Bonsai PTQ1_0 5.53 GiB 757.8 ± 22.1 32.89 ± 0.16 664.8 ± 15.4 31.42 ± 0.14

generation_speed

Notes:

  • Prompt processing is roughly tied across all three; decode speed is the only real difference.
  • Bonsai PQ2_0 is 2.4× smaller than Q4_K_XL but only 1.38× faster. Effective decode bandwidth is about 335 GB/s vs about 590 GB/s for Q4_K, so the ternary kernels are unpacking-bound on Volta rather than memory-bound.
  • PTQ1_0's denser trit packing costs more ALU per weight, making it slower than both others on this card. Its only advantage is 1.2 GB less VRAM.
  • Both models slow down very little at 8K context (hybrid attention).

Perplexity (wikitext-2 test, full 145 chunks, -c 2048)

Model PPL
Qwen3.8 Q4_K_XL 6.3439 ± 0.0404
Bonsai PQ2_0 8.3536 ± 0.0576

Bonsai is about 32% higher. All models share the same tokenizer (gpt2 / qwen35, 248,320 tokens, BOS 248044, EOS 248046), so these are directly comparable.

KL divergence vs BF16 (wikitext-2, 50 chunks, -c 2048)

Reference logits from BF16 with partial offload (-ngl 34, about 26.9 GB VRAM, 51.5 s/chunk, about 43 min). BF16 PPL over these 50 chunks: 6.1345 ± 0.0654.

Metric Q8_0 Q4_K_XL Bonsai PQ2_0
Mean PPL 6.1396 6.1003 8.0420
PPL ratio vs BF16 1.0008 0.9944 1.3110
Cor(ln PPL) vs BF16 99.96% 99.73% 92.16%
Mean KLD 0.000644 0.008685 0.340281
Median KLD 0.000101 0.001830 0.175541
90% KLD 0.000467 0.010595 0.752498
95% KLD 0.000784 0.019142 1.185927
99% KLD 0.002857 0.065884 2.886118
99.9% KLD 0.041576 0.558065 7.351860
Max KLD 4.84 21.93 21.24
Mean Δp −0.003% −0.029% −4.376%
RMS Δp 0.670% 2.582% 15.952%
Same top token 99.284% 96.882% 77.769%

mean_kld

kld_percentiles

same_top_token

Interpretation

  • Q8_0: Indistinguishable from BF16 for practical purposes (mean KLD 0.0006, 99.3% top-token agreement).
  • Q4_K_XL: A strong quant. Its slightly lower PPL than BF16 is a common artifact, not an improvement; KLD shows small real drift. Greedy output differs from BF16 about once every 32 tokens.
  • Bonsai PQ2_0: Divergence is broad, not limited to a few outlier tokens: its median KLD (0.18) exceeds the Q4's 99th percentile (0.066). It disagrees with BF16's top pick about once every 4.5 tokens. Its negative mean Δp shows it is systematically less confident in the correct next token. The lower per-chunk PPL correlation (92%) means it finds different passages hard, consistent with a retrained sibling model rather than a lossy copy.
  • Ratios: Bonsai mean KLD is about 39× Q4_K_XL and about 530× Q8_0.

Caveats

  • Perplexity and KLD score raw Wikipedia text without a chat template or reasoning. They measure closeness to Qwen3.8-27B's distribution, not task ability.
  • PrismML's "~98% of benchmark performance" claim is based on task benchmarks in thinking mode, which this suite doesn't test.
  • Bonsai's kernels are not tuned or benchmarked by PrismML on Volta; speed results may not reflect newer GPUs.
  • The BF16 base run disabled fused Gated Delta Net ops because of the CPU/GPU split. This affects speed only, not results.

Next steps

  • Run task evals through llama-server with the chat template, e.g. lm_eval --model local-chat-completions --tasks gsm8k,mmlu_pro --apply_chat_template --limit 200, with matching sampling and reasoning effort for Q4_K_XL and Bonsai PQ2_0.
  • Optionally, HellaSwag / Winogrande via llama-perplexity --hellaswag / --winogrande for quick accuracy-based checks.

Commands used

# prism helper
prism() { docker run --rm --init --gpus '"device=1"' -v /mnt/user/appdata/llama_cpp/model:/models local/llama.cpp-prism:full "$@"; }

# speed
prism --bench -m <model> -ngl 99 -fa 1 -p 512 -n 128 -d 0,8192 -r 5

# perplexity
prism --perplexity -m <model> -f /models/wiki.test.raw -ngl 99 -c 2048

# KLD base (BF16, partial offload)
prism --perplexity -m <bf16-00001-of-00002.gguf> -f /models/wiki.test.raw -ngl 34 -c 2048 --chunks 50 \
  --kl-divergence-base /models/qwen-bf16.kld

# KLD compare
prism --perplexity -m <model> -f /models/wiki.test.raw -ngl 99 -c 2048 --chunks 50 \
  --kl-divergence-base /models/qwen-bf16.kld --kl-divergence

Very interesting result. I wonder if the large KL divergence here may connect directly to what was observed in #44.

In #44, plain rotate+RTN already reproduced roughly 92% of the ternary symbols, yet replacing the remaining ~8% with those apparently natural RTN choices caused a catastrophic perplexity collapse. GPTQ/Hessian-based attempts also failed to reproduce that residual structure, which suggested that the important difference was not simply a better post-training quantizer, but trained weight movement before the final ternary projection.

Seen from that perspective, the large KLD here may not be “quantization error” in the usual PTQ sense at all.

It may instead be measuring how far the retrained ternary model has moved into a different low-bit-friendly solution basin while preserving much of the useful task behavior.

That would also explain why the model can differ substantially from BF16 at the token-distribution level, while still retaining surprisingly strong reasoning/coding performance.

Of course, the exact Bonsai 2 training recipe is not public, so this is only an interpretation. But the combination of:

  • ~92% symbol agreement with simple rotate+RTN,
  • extreme sensitivity to the remaining ~8% in #44,
  • failure of ordinary PTQ/Hessian methods to reproduce those choices,
  • and the much larger KLD measured here,

seems consistent with QAT/KD or some other training process moving selected weights across ternary decision boundaries rather than merely minimizing local rounding error.

In other words, perhaps the interesting quantity here is not just “how much information was lost by quantization,” but “how far training intentionally moved the model before quantization became cheap.”

That distinction may be important when interpreting Bonsai against conventional Q4/Q8 quantization.

Yes, that is a good point mktnhr.

I have also seen though "real word tasks" do diverge from Q4 quantizations.

Please see linked: https://www.youtube.com/watch?v=NZMUtVPxJvQ

Could you please include the comparison with UD-IQ2_XXS or UD-IQ2_S as well? Since they are in the same "weight class"?

Sign up or log in to comment