Polaris2 Fable B F451
MoQ vs DavidAU Fable Fusion 711

Actual GGUF size, payload BPW, WikiText-2 PPL, Mean KLD, and p999 KLD for 9 regular MoQ models, the separately highlighted MoQ-NVFP4 model, and 9 GGUF files published by DavidAU for the Fable Fusion 711 version. Every point uses the same 580-chunk, context-512 evaluation. MoQ-4.85 and MoQ-4.95 are intentionally omitted.

Our regular MoQ models
9
Highlighted NVFP4 model
1
DavidAU Fable Fusion 711 GGUF
9 / 9
BF16 reference PPL
6.320035
Visible series Drag to pan, use the mouse wheel to zoom, or use the chart toolbar. Click a legend item to toggle it in one chart.

Near-size comparisons

Our modelDavidAU modelGB (ours / DavidAU)Size smallerMean KLD lowerp999 lowerPPL lower
MoQ-3.6 IQ2_M 12.134 / 12.125 -0.08% +50.41% +52.82% +6.18%
MoQ-4.1 IQ3_M 14.266 / 14.532 +1.83% +42.88% +45.75% +1.50%
MoQ-4.8 IQ4_XS 16.150 / 17.034 +5.19% +8.01% +17.69% +0.13%
MoQ-5.1 Q4_K_S 17.452 / 17.537 +0.49% +31.16% +33.90% +0.12%

Positive percentages mean our model is smaller or has a lower metric. The pairs are selected by close file size, not by nominal quantization label.

Size-aligned RTX 5090 throughput

Modelpp512 tok/stg128 tok/spg32768,256 tok/s
MoQ-NVFP4 3256.97 +/- 407.68 78.92 +/- 0.13 2377.72 +/- 10.42
MoQ-4.8 2669.45 +/- 227.45 78.65 +/- 0.29 2356.68 +/- 9.67
MoQ-4.6 2731.86 +/- 191.67 82.87 +/- 0.16 2400.43 +/- 5.78
NVFP4 vs MoQ-4.8 +22.01% +0.34% +0.89%

MoQ-NVFP4 and MoQ-4.8 are the same 16.150 GB size; MoQ-4.6 remains as an additional reference. Mean +/- sample standard deviation over 3 repetitions. f16 KV, -ngl 999 -fa on -b 2048 -ub 512 -t 16. The delta row uses MoQ-4.8 as the baseline; positive means NVFP4 is faster.

Complete measured results

SeriesModelGBGiBPayload BPWPPLMean KLDp999 KLDSame top
Our MoQ MoQ-3.2 10.810 10.068 3.1621 6.805998 0.097714 2.751078 86.459%
DavidAU Fable Fusion 711 DavidAU-IQ2_M 12.125 11.292 3.5471 7.100246 0.140480 3.807066 83.864%
Our MoQ MoQ-3.6 12.134 11.301 3.5500 6.661694 0.069665 1.796203 88.594%
Our MoQ MoQ-3.8 12.840 11.958 3.7566 6.518863 0.044993 1.291825 90.505%
Our MoQ MoQ-4.1 14.266 13.286 4.1740 6.447097 0.027318 0.807045 92.266%
DavidAU Fable Fusion 711 DavidAU-IQ3_M 14.532 13.534 4.2520 6.545219 0.047823 1.487676 90.751%
Our MoQ MoQ-4.3 15.002 13.972 4.3897 6.421027 0.019253 0.591232 93.300%
Our MoQ MoQ-4.6 15.242 14.195 4.4599 6.392239 0.015145 0.505663 94.425%
Our MoQ MoQ-4.8 16.150 15.041 4.7258 6.379456 0.012836 0.423877 94.828%
MoQ-NVFP4 MoQ-NVFP4 16.150 15.041 4.7258 6.397915 0.015615 0.518097 94.393%
Our MoQ MoQ-4.9 16.524 15.389 4.8353 6.384560 0.012218 0.402534 94.935%
DavidAU Fable Fusion 711 DavidAU-IQ4_XS 17.034 15.864 4.9846 6.387851 0.013953 0.514973 94.774%
Our MoQ MoQ-5.1 17.452 16.253 5.1069 6.371239 0.009613 0.325677 95.480%
DavidAU Fable Fusion 711 DavidAU-Q4_K_S 17.537 16.333 5.1321 6.378748 0.013965 0.492694 94.785%
DavidAU Fable Fusion 711 DavidAU-IQ4_NL 17.753 16.534 5.1952 6.386419 0.013701 0.492569 94.926%
DavidAU Fable Fusion 711 DavidAU-Q4_K_M 18.499 17.228 5.4135 6.365021 0.011404 0.412386 95.299%
DavidAU Fable Fusion 711 DavidAU-Q5_K_S 20.631 19.214 6.0379 6.341159 0.005466 0.211335 96.862%
DavidAU Fable Fusion 711 DavidAU-Q5_K_M 21.182 19.728 6.1993 6.337103 0.004809 0.179674 97.087%
DavidAU Fable Fusion 711 DavidAU-Q6_K 24.034 22.383 7.0343 6.324763 0.001467 0.059772 98.402%

GB is decimal. Payload BPW excludes the GGUF header and divides the data payload by the tensor parameter count. MoQ recipes were audited tensor by tensor before evaluation.

Evaluation conditions

-c 512 -b 8192 -ub 4096 -ngl auto --fit on -fitt 1024 --op-offload -fa on -t 16 -tb 16 --kl-divergence

Corpus: WikiText-2 raw wiki.test.raw. KLD reference: supplied BF16 logits file. Benchmark: -p 512 -n 128 -pg 32768,256 -ctk f16 -ctv f16 -ngl 999 -fa on -r 3. All GPU jobs ran serially.