YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Performance Summary: Qwen3-30B-A3B REAP and FP8-DYNAMIC RTN Ablations Study
Overview
This report compares 4 compressed/quantized variants of Qwen3-30B-A3B-Instruct-2507 against the unquantized baseline across 5 evaluation benchmarks.
- FP8-DYNAMIC: Quantized with FP8-DYNAMIC RTN only
- REAP-25pct: Compressed with REAP (25% sparsity) only
- FP8-DYNAMIC + REAP-25pct: Quantized with FP8-DYNAMIC RTN first, then compressed with REAP (25% sparsity)
- REAP-25pct + FP8-DYNAMIC: Compressed with REAP (25% sparsity) first, then quantized with FP8-DYNAMIC RTN
Important Note: The baseline model was served with tp=2 across 2 GPUs, so its server statistics may appear different from the single-GPU variants. The baseline model size (57 GB) is accurate, but its KV cache size and token capacity may be inflated relative to a single-GPU deployment.
Model Configurations
| Configuration | Model Size | Size Reduction | Session ID |
|---|---|---|---|
| Baseline (Unquantized) | 57.00 GB | — | 07102026028 |
| FP8-DYNAMIC | 29.21 GB | 48.8% | 07102026-f817 |
| REAP-25pct | 43.43 GB | 23.8% | 07102026-aedc |
| FP8-DYNAMIC + REAP-25pct | 22.43 GB | 60.6% | 07102026-da72 |
| REAP-25pct + FP8-DYNAMIC | 22.43 GB | 60.6% | 07102026-05ea |
Accuracy Scores
| Task | Baseline | FP8-DYNAMIC | REAP-25pct | FP8-DYNAMIC + REAP-25pct | REAP-25pct + FP8-DYNAMIC |
|---|---|---|---|---|---|
| gsm8k_platinum_cot_llama | 0.9680 | 0.9763 | 0.9586 | 0.9573 | 0.9570 |
| ifeval | 0.8945 | 0.8921 | 0.8705 | 0.8733 | 0.8713 |
| math_500|0 | 0.8887 | 0.8913 | 0.8893 | 0.8740 | 0.8813 |
| mmlu_cot_llama | 0.8422 | 0.8397 | 0.7930 | 0.7910 | 0.7921 |
| mmlu_pro_chat | 0.7881 | 0.7848 | 0.7264 | 0.7237 | 0.7253 |
| Average | 0.8763 | 0.8768 | 0.8476 | 0.8439 | 0.8454 |
Accuracy Recovery vs. Baseline
| Task | FP8-DYNAMIC | REAP-25pct | FP8-DYNAMIC + REAP-25pct | REAP-25pct + FP8-DYNAMIC |
|---|---|---|---|---|
| gsm8k_platinum_cot_llama | 100.86% (+0.0083) | 99.03% (-0.0094) | 98.89% (-0.0107) | 98.86% (-0.0110) |
| ifeval | 99.73% (-0.0024) | 97.32% (-0.0240) | 97.63% (-0.0212) | 97.41% (-0.0232) |
| math_500|0 | 100.29% (+0.0026) | 100.07% (+0.0006) | 98.35% (-0.0147) | 99.17% (-0.0074) |
| mmlu_cot_llama | 99.70% (-0.0025) | 94.16% (-0.0492) | 93.92% (-0.0512) | 94.05% (-0.0501) |
| mmlu_pro_chat | 99.58% (-0.0033) | 92.17% (-0.0617) | 91.83% (-0.0644) | 92.03% (-0.0628) |
| Average | 100.06% (+0.0005) | 96.72% (-0.0287) | 96.30% (-0.0324) | 96.47% (-0.0309) |
Variant Ranking (by Average Recovery)
| Rank | Configuration | Avg Recovery | Avg Δ |
|---|---|---|---|
| 1 | FP8-DYNAMIC | 100.06% | +0.0005 |
| 2 | REAP-25pct | 96.72% | -0.0287 |
| 3 | REAP-25pct + FP8-DYNAMIC | 96.47% | -0.0309 |
| 4 | FP8-DYNAMIC + REAP-25pct | 96.30% | -0.0324 |
Resource Utilization
| Metric | Baseline | FP8-DYNAMIC | REAP-25pct | FP8-DYNAMIC + REAP-25pct | REAP-25pct + FP8-DYNAMIC |
|---|---|---|---|---|---|
| KV Cache Size (GB) | 40.29* | 39.90 | 25.86 | 46.57 | 46.57 |
| KV Cache Tokens | 880,192* | 435,776 | 282,448 | 508,672 | 508,672 |
| Recommended Concurrency | 27.51* | 13.62 | 8.83 | 15.90 | 15.90 |
* Baseline was served with tp=2 across 2 GPUs; KV cache size and token capacity are inflated relative to single-GPU variants and should not be directly compared.
Concurrency Issues
| Configuration | Task | Issues |
|---|---|---|
| REAP-25pct | mmlu_pro_chat | max KV cache: 91.2% |
Evaluation Date: 2026-07-10 vLLM Version: 0.24.0 Max Model Length: 32,000 tokens Source PRs: #37, #38, #39, #40, #41
- Downloads last month
- 90