Polaris2 Fable B F451
MoQ vs DavidAU Fable Fusion 711
Actual GGUF size, payload BPW, WikiText-2 PPL, Mean KLD, and p999 KLD for 9 regular MoQ models, the separately highlighted MoQ-NVFP4 model, and 9 GGUF files published by DavidAU for the Fable Fusion 711 version. Every point uses the same 580-chunk, context-512 evaluation. MoQ-4.85 and MoQ-4.95 are intentionally omitted.
Near-size comparisons
| Our model | DavidAU model | GB (ours / DavidAU) | Size smaller | Mean KLD lower | p999 lower | PPL lower |
|---|---|---|---|---|---|---|
| MoQ-3.6 | IQ2_M | 12.134 / 12.125 | -0.08% | +50.41% | +52.82% | +6.18% |
| MoQ-4.1 | IQ3_M | 14.266 / 14.532 | +1.83% | +42.88% | +45.75% | +1.50% |
| MoQ-4.8 | IQ4_XS | 16.150 / 17.034 | +5.19% | +8.01% | +17.69% | +0.13% |
| MoQ-5.1 | Q4_K_S | 17.452 / 17.537 | +0.49% | +31.16% | +33.90% | +0.12% |
Positive percentages mean our model is smaller or has a lower metric. The pairs are selected by close file size, not by nominal quantization label.
Size-aligned RTX 5090 throughput
| Model | pp512 tok/s | tg128 tok/s | pg32768,256 tok/s |
|---|---|---|---|
| MoQ-NVFP4 | 3256.97 +/- 407.68 | 78.92 +/- 0.13 | 2377.72 +/- 10.42 |
| MoQ-4.8 | 2669.45 +/- 227.45 | 78.65 +/- 0.29 | 2356.68 +/- 9.67 |
| MoQ-4.6 | 2731.86 +/- 191.67 | 82.87 +/- 0.16 | 2400.43 +/- 5.78 |
| NVFP4 vs MoQ-4.8 | +22.01% | +0.34% | +0.89% |
MoQ-NVFP4 and MoQ-4.8 are the same 16.150 GB size; MoQ-4.6 remains as an additional reference. Mean +/- sample standard deviation over 3 repetitions. f16 KV, -ngl 999 -fa on -b 2048 -ub 512 -t 16. The delta row uses MoQ-4.8 as the baseline; positive means NVFP4 is faster.
Complete measured results
| Series | Model | GB | GiB | Payload BPW | PPL | Mean KLD | p999 KLD | Same top |
|---|---|---|---|---|---|---|---|---|
| Our MoQ | MoQ-3.2 | 10.810 | 10.068 | 3.1621 | 6.805998 | 0.097714 | 2.751078 | 86.459% |
| DavidAU Fable Fusion 711 | DavidAU-IQ2_M | 12.125 | 11.292 | 3.5471 | 7.100246 | 0.140480 | 3.807066 | 83.864% |
| Our MoQ | MoQ-3.6 | 12.134 | 11.301 | 3.5500 | 6.661694 | 0.069665 | 1.796203 | 88.594% |
| Our MoQ | MoQ-3.8 | 12.840 | 11.958 | 3.7566 | 6.518863 | 0.044993 | 1.291825 | 90.505% |
| Our MoQ | MoQ-4.1 | 14.266 | 13.286 | 4.1740 | 6.447097 | 0.027318 | 0.807045 | 92.266% |
| DavidAU Fable Fusion 711 | DavidAU-IQ3_M | 14.532 | 13.534 | 4.2520 | 6.545219 | 0.047823 | 1.487676 | 90.751% |
| Our MoQ | MoQ-4.3 | 15.002 | 13.972 | 4.3897 | 6.421027 | 0.019253 | 0.591232 | 93.300% |
| Our MoQ | MoQ-4.6 | 15.242 | 14.195 | 4.4599 | 6.392239 | 0.015145 | 0.505663 | 94.425% |
| Our MoQ | MoQ-4.8 | 16.150 | 15.041 | 4.7258 | 6.379456 | 0.012836 | 0.423877 | 94.828% |
| MoQ-NVFP4 | MoQ-NVFP4 | 16.150 | 15.041 | 4.7258 | 6.397915 | 0.015615 | 0.518097 | 94.393% |
| Our MoQ | MoQ-4.9 | 16.524 | 15.389 | 4.8353 | 6.384560 | 0.012218 | 0.402534 | 94.935% |
| DavidAU Fable Fusion 711 | DavidAU-IQ4_XS | 17.034 | 15.864 | 4.9846 | 6.387851 | 0.013953 | 0.514973 | 94.774% |
| Our MoQ | MoQ-5.1 | 17.452 | 16.253 | 5.1069 | 6.371239 | 0.009613 | 0.325677 | 95.480% |
| DavidAU Fable Fusion 711 | DavidAU-Q4_K_S | 17.537 | 16.333 | 5.1321 | 6.378748 | 0.013965 | 0.492694 | 94.785% |
| DavidAU Fable Fusion 711 | DavidAU-IQ4_NL | 17.753 | 16.534 | 5.1952 | 6.386419 | 0.013701 | 0.492569 | 94.926% |
| DavidAU Fable Fusion 711 | DavidAU-Q4_K_M | 18.499 | 17.228 | 5.4135 | 6.365021 | 0.011404 | 0.412386 | 95.299% |
| DavidAU Fable Fusion 711 | DavidAU-Q5_K_S | 20.631 | 19.214 | 6.0379 | 6.341159 | 0.005466 | 0.211335 | 96.862% |
| DavidAU Fable Fusion 711 | DavidAU-Q5_K_M | 21.182 | 19.728 | 6.1993 | 6.337103 | 0.004809 | 0.179674 | 97.087% |
| DavidAU Fable Fusion 711 | DavidAU-Q6_K | 24.034 | 22.383 | 7.0343 | 6.324763 | 0.001467 | 0.059772 | 98.402% |
GB is decimal. Payload BPW excludes the GGUF header and divides the data payload by the tensor parameter count. MoQ recipes were audited tensor by tensor before evaluation.
Evaluation conditions
-c 512 -b 8192 -ub 4096 -ngl auto --fit on -fitt 1024 --op-offload -fa on -t 16 -tb 16 --kl-divergence
Corpus: WikiText-2 raw wiki.test.raw. KLD reference: supplied BF16 logits file. Benchmark: -p 512 -n 128 -pg 32768,256 -ctk f16 -ctv f16 -ngl 999 -fa on -r 3. All GPU jobs ran serially.