| # MTP Test Results |
|
|
| Raw results from the MTP viability test plan. Updated as tests complete. |
|
|
| ## Phase 1: Task Type Effect (temp=0.7, MTP N=3, -n 2500) |
|
|
| ### Acceptance Rate |
|
|
| | Quant | Code | Factual | Analysis | Essay* | Creative | |
| |-----------|--------|---------|----------|---------|----------| |
| | Q4_K_M | 82.3% | 69.4% | 54.2% | 55.8% | 47.2% | |
| | Q5_K_M | 84.2% | 67.5% | 56.1% | 55.1% | 43.4% | |
| | Q6_K | 83.9% | 69.0% | 55.4% | 54.8% | 43.9% | |
| | F16 | 80.8% | β | β | β | β | |
| | Q8_0 | 86.8% | 69.8% | 55.6% | 55.7% | 46.8% | |
|
|
| *Essay data from previous long-gen tests (temp 0.7, different prompt). F16 Code at N=3 (optimal is N=4, see Phase 6). |
| |
| ### Decode Speed (tok/s) |
| |
| | Quant | Code | Factual | Analysis | Essay* | Creative | |
| |-----------|--------|---------|----------|---------|----------| |
| | Q4_K_M | 19.66 | 17.51 | 14.90 | 14.95 | 13.74 | |
| | Q5_K_M | 19.23 | 16.54 | 14.68 | 13.83 | 12.64 | |
| | Q6_K | 20.11 | 17.59 | 15.24 | 14.97 | 13.36 | |
| | F16 | 16.22 | β | β | β | β | |
| | Q8_0 | 25.36 | 21.67 | 18.63 | 18.33 | 16.88 | |
|
|
| --- |
|
|
| ## Phase 2: Temperature x Task Interaction (MTP N=3, -n 2500) |
|
|
| ### Q4_K_M: Temperature Effect |
|
|
| | Temp | Code (tok/s) | Code Acc% | Creative (tok/s) | Creative Acc% | |
| |-------|-------------|-----------|-------------------|---------------| |
| | 0.0 | 20.59 | 87.1% | 13.93 | 47.8% | |
| | 0.3 | 20.84 | 88.5% | 13.55 | 45.6% | |
| | 0.7 | 19.66 | 82.3% | 13.74 | 47.2% | |
|
|
| ### Q8_0: Temperature Effect |
| |
| | Temp | Code (tok/s) | Code Acc% | Creative (tok/s) | Creative Acc% | |
| |-------|-------------|-----------|-------------------|---------------| |
| | 0.0 | 25.46 | 86.8% | 17.04 | 47.3% | |
| | 0.3 | 25.10 | 85.4% | 17.36 | 48.7% | |
| | 0.7 | 25.36 | 86.8% | 16.88 | 46.8% | |
| |
| --- |
| |
| ## Phase 3: Baselines (no MTP) at temp=0.7 |
| |
| | Quant | Speed (tok/s) | Tokens Gen | ms/token | |
| |-----------|---------------|------------|----------| |
| | Q4_K_M | 15.07 | ~2500 | 66.36 | |
| | Q5_K_M | 13.12 | 1757 | 76.24 | |
| | Q6_K | 13.44 | 1820 | 74.38 | |
| | Q8_0 | 11.39 | 1955 | 87.78 | |
| | F16 | 6.60 | 1808 | 151.51 | |
| |
| --- |
| |
| ## Speedup Summary (MTP vs Baseline) |
| |
| Speedup = (MTP_speed / Baseline_speed - 1) Γ 100% |
| |
| ### Phase 1: temp=0.7 (essay baseline) |
| |
| | Quant | Code | Factual | Analysis | Essay | Creative | |
| |-----------|---------|---------|----------|---------|----------| |
| | Q4_K_M | +30.5% | +16.2% | -1.1% | -0.8% | -8.8% | |
| | Q5_K_M | +46.6% | +26.1% | +11.9% | +5.4% | -3.7% | |
| | Q6_K | +49.6% | +30.7% | +13.4% | +11.2% | -0.6% | |
| | Q8_0 | +122.7% | +90.3% | +63.6% | +60.9% | +48.2% | |
| |
| ### F16: temp=0.0, N=4 (F16 baseline 6.60 tok/s) |
| |
| | Quant | Code | Factual | Analysis | Essay | Creative | |
| |-----------|---------|---------|----------|---------|----------| |
| | F16 | +171% | +125% | +91% | +77% | +67% | |
| |
| ### Phase 2: temp effect on code |
| |
| | Quant | 0.0 | 0.3 | 0.7 | |
| |---------|--------|--------|--------| |
| | Q4_K_M | +36.6% | +38.3% | +30.5% | |
| | Q8_0 | +123.5%| +120.5%| +122.7%| |
|
|
| ### Phase 2: temp effect on creative |
|
|
| | Quant | 0.0 | 0.3 | 0.7 | |
| |---------|--------|--------|--------| |
| | Q4_K_M | -7.6% | -10.1% | -8.8% | |
| | Q8_0 | +49.6% | +52.4% | +48.2% | |
| |
| --- |
| |
| ## Key Findings |
| |
| 1. **Task type is the dominant factor**: 40-pt acceptance spread (code 87% β creative 43%) vs ~1-pt quant effect |
| 2. **Temperature barely matters for acceptance**: Code stays at 82-89% across all temps. Creative stays at 43-48%. |
| 3. **Quant level matters for speed, not acceptance**: All quants have similar acceptance, but Q8_0's MTP speed is 25+ tok/s vs 19-20 for Q4_K_M due to less dequant overhead. The bigger the model, the more MTP helps because baseline is more bandwidth-bound. |
| 4. **Q8_0 + MTP + code = +123% speedup**: The best case β 25.36 tok/s vs 11.39 baseline |
| 5. **Creative writing is the worst case for all quants**: ~47% acceptance at all temperatures, below breakeven for Q4_K_M (net slowdown) |
| 6. **Q8_0 MTP helps even for creative**: +48% speedup despite only 47% acceptance, because Q8_0 baseline is so bandwidth-starved that any accepted drafts help |
| |
| --- |
| |
| ## Phase 6: Draft Token Count Effect (temp=0.7, -n 2500) |
| |
| Testing N=1 through N=5 on 4 key scenarios. N=3 data from Phase 1 included for completeness. |
| |
| ### Q8_0 Code (best case β high determinism) |
|
|
| | N | Speed (tok/s) | Acceptance | vs Baseline (11.39) | |
| |---|---------------|------------|----------------------| |
| | 1 | 17.81 | 96.8% | +56% | |
| | 2 | 22.24 | 91.9% | +95% | |
| | 3 | 25.36 | 86.8% | +123% | |
| | 4 | 23.74 | 79.0% | +108% | |
| | 5 | 21.36 | 71.9% | +88% | |
|
|
| **Optimal: N=3** (25.36 tok/s). N=4 and N=5 lose more to rejection than they gain from extra drafts. |
|
|
| ### Q8_0 Creative (low determinism) |
| |
| | N | Speed (tok/s) | Acceptance | vs Baseline (11.39) | |
| |---|---------------|------------|----------------------| |
| | 1 | 15.54 | 71.6% | +36% | |
| | 2 | 16.48 | 54.1% | +45% | |
| | 3 | 16.88 | 46.8% | +48% | |
| | 4 | 13.90 | 35.3% | +22% | |
| | 5 | 11.70 | 29.7% | +3% | |
| |
| **Optimal: N=3** (16.88 tok/s). N=1 and N=2 are decent alternatives. N=5 barely breaks even. |
| |
| ### Q4_K_M Code (moderate determinism, bandwidth-efficient) |
| |
| | N | Speed (tok/s) | Acceptance | vs Baseline (15.07) | |
| |---|---------------|------------|----------------------| |
| | 1 | 16.92 | 93.0% | +12% | |
| | 2 | 17.74 | 91.3% | +18% | |
| | 3 | 19.66 | 82.3% | +31% | |
| | 4 | 19.61 | 79.6% | +30% | |
| | 5 | 16.29 | 64.9% | +8% | |
| |
| **Optimal: N=3** (19.66 tok/s). N=4 is nearly identical. N=5 collapses from low acceptance. |
| |
| ### Q4_K_M Creative (worst case β low determinism, bandwidth-efficient) |
| |
| | N | Speed (tok/s) | Acceptance | vs Baseline (15.07) | |
| |---|---------------|------------|----------------------| |
| | 1 | 15.01 | 73.6% | -0.4% | |
| | 2 | 13.33 | 57.9% | -12% | |
| | 3 | 13.74 | 47.2% | -9% | |
| | 4 | 11.37 | 36.0% | -25% | |
| | 5 | 9.35 | 28.5% | -38% | |
| |
| **Optimal: N=1** (15.01 tok/s, essentially break-even). Even N=1 only matches baseline β MTP has no viable configuration for Q4_K_M creative. N=2 through N=5 are all net slowdowns. |
| |
| ### F16 Code (temp=0.7, extremely bandwidth-starved) |
| |
| | N | Speed (tok/s) | Acceptance | vs Baseline (6.60) | |
| |---|---------------|------------|----------------------| |
| | 3 | 16.22 | 80.8% | +146% | |
| | 4 | 17.86 | 78.8% | +171% | |
| | 5 | 17.23 | 70.9% | +161% | |
| |
| **Optimal: N=4** (17.86 tok/s). F16 benefits from N=4 (not N=3 like other quants) because its 6.60 tok/s baseline is so bandwidth-starved that the extra draft token's marginal bandwidth savings outweigh the lower acceptance. N=5 drops below N=4. |
| |
| ### N-Draft Conclusions |
| |
| 1. **N=3 is the optimal default** for quantized models (Q4-Q8). It maximizes throughput before diminishing returns. |
| 2. **F16 shifts the optimum to N=4** β its extremely slow baseline (6.60 tok/s) means each extra accepted draft token saves proportionally more bandwidth, justifying the slight acceptance drop from N=3βN=4. |
| 3. **For creative/low-determinism tasks, lower N helps acceptance but rarely enough** β Q4_K_M creative at N=1 still can't beat baseline. Q8_0 creative benefits from N=1-3 but N=4-5 are counterproductive. |
| 4. **N=5 is almost never optimal** β acceptance drops too steeply. Only viable for Q8_0 code where even 71.9% acceptance yields +88% speedup. |
| 5. **The acceptance cliff is sharper than expected** β each additional draft token has progressively lower acceptance. The first draft token (N=1) has 71-97% acceptance, but the 5th often falls below 30%. |
| 6. **For Q4_K_M creative, no N value makes MTP worthwhile** β even N=1 (73.6% acceptance, 15.01 tok/s) only matches the 15.07 baseline. The draft generation overhead eats the entire gain. |
| |
| --- |
| |
| ## Phase 5: Final Recommendations |
| |
| ### MTP Viability Matrix (Qwen3.6-27B, M2 Max 96GB) |
| |
| | Use Case | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | F16 | |
| |----------|--------|--------|------|------|-----| |
| | **Coding assistant** (temp 0.0-0.3) | π’ +37% | π’ ~+47% | π’ ~+50% | π’ +123% | π’ +171% | |
| | **Factual / technical QA** (temp 0.3-0.7) | π‘ +16% | π’ +26% | π’ +31% | π’ +90% | π’ +125% | |
| | **Technical analysis** (temp 0.7) | π΄ -1% | π‘ +12% | π‘ +13% | π’ +64% | π’ +91% | |
| | **General chat / essay** (temp 0.7) | π΄ -1% | π‘ +5% | π‘ +11% | π’ +61% | π’ +77% | |
| | **Creative writing** (temp 0.7) | π΄ -9% | π΄ -4% | π΄ -1% | π’ +48% | π’ +67% | |
| |
| ### Legend |
| - π’ **Recommended** β clear speedup |
| - π‘ **Marginal** β helps slightly, task-dependent |
| - π΄ **Not recommended** β net slowdown or break-even |
| |
| ### Decision Rules |
| |
| 1. **Always enable MTP for F16 and Q8_0** β F16 sees +67% to +171% across all tasks (6.60 tok/s baseline is extremely bandwidth-starved). Q8_0 sees +48% to +123%. Any accepted drafts help when the baseline is this constrained. |
|
|
| 2. **Enable MTP for code at any quant** β code generation is deterministic enough (79-89% acceptance) that MTP always wins, giving +30% to +171%. |
|
|
| 3. **Enable MTP for factual/technical at Q5_K_M+** β Q4_K_M is marginal (+16%), but Q5_K_M and above get meaningful speedups. |
|
|
| 4. **Disable MTP for creative writing at Q4_K_M through Q6_K** β acceptance is 43-47%, below breakeven for these quants. Net slowdown of -1% to -9%. |
| |
| 5. **Temperature doesn't matter** β acceptance is remarkably stable across temp 0.0-0.7 for both code and creative. The sampling chain mismatch is a fixed penalty; temperature doesn't amplify or reduce it. |
| |
| 6. **For mixed workloads** (coding + creative): Enable MTP for Q5_K_M+ and F16 (net positive across tasks). For Q4_K_M, only enable if code-heavy. |
| |
| 7. **F16 optimal N=4 (not N=3)** β the extremely slow baseline means N=4's extra draft token is worth more than the slight acceptance drop. All other quants use N=3. |
| |
| ### Why F16 and Q8_0 Benefit Most |
| |
| F16 has the slowest baseline (6.60 tok/s) because the 51GB model is extremely bandwidth-starved β it's nearly 2Γ the size of Q8_0 and 3Γ Q4_K_M. MTP draft tokens are "free" bandwidth-wise (they reuse already-loaded weights), so every accepted draft token saves a full decode pass. Even at 39% acceptance (creative), the net effect is +67% because the baseline is so constrained. |
| |
| Q8_0 follows the same pattern at 11.39 tok/s baseline. Conversely, Q4_K_M's baseline (15.07 tok/s) is faster because the 16GB model has better bandwidth utilization. MTP has less room to improve and the overhead (draft generation + verification) is proportionally larger, making the breakeven acceptance higher. |
| |
| This creates a clear hierarchy: the larger the model (more bandwidth-starved), the more MTP helps regardless of task type. |
| |
| ### Caveats |
| |
| - All tests on M2 Max 96GB with Metal GPU. Results may differ on CUDA or other hardware. |
| - Baseline speeds measured with essay task; other tasks may have slightly different baselines. |
| - N=3 draft tokens (`--spec-draft-n-max 3`). N=5 may shift breakeven thresholds. |
| - Non-thinking mode (`<|think_off|>`) used for Phases 1-6. Phase 7 tested thinking mode: acceptance drops to ~73% for code (vs ~87% non-thinking) but still above breakeven for Q8_0. |
| |
| --- |
| |
| ## Phase 7: Thinking Mode Effect (Q8_0, temp=0.7, N=3, -n 2500) |
| |
| Testing whether chain-of-thought (thinking) mode changes MTP acceptance. Same code prompt, same settings, just without `<|think_off|>`. |
| |
| | Mode | Speed (tok/s) | Acceptance | Accepted/Generated | |
| |------|---------------|------------|---------------------| |
| | Non-thinking | 25.36 | 86.8% | 1778/2160 | |
| | Thinking | 22.04 | 73.1% | 1716/2346 | |
| |
| Thinking mode drops acceptance by ~14 points (86.8% β 73.1%). The model's CoT output is less deterministic than expected β the reasoning steps have multiple valid paths, so the MTP head and verification disagree more often. Speed still good at 22.04 tok/s vs 11.39 baseline (+94%), but lower than non-thinking's +123%. |
| |
| Raw log: `/tmp/mtp-q8_0-thinking-code-0.7.log` |
| |
| --- |
| |
| ## Notes |
| |
| - Breakeven acceptance for N=3 β 59% for low quants (less MTP overhead), lower for Q8_0, much lower for F16 |
| - Temperature mismatch is root cause: MTP drafts greedily (top_k=1), verification uses user sampling chain |
| - All quants use identical Q8_0 MTP layers (blk.64 tensors) |
| - Baseline used essay task; actual baseline varies slightly by task but is a reasonable proxy |
| - Creative acceptance at temp 0.0 is still only ~39-48% β the model itself has diverse top-token distributions for creative content regardless of sampling temperature |
| - Thinking mode lowers acceptance (~73% vs ~87% for code) but still well above breakeven for Q8_0 |
| - F16 optimal N=4 (vs N=3 for all other quants) β the 6.60 tok/s baseline makes each extra draft token more valuable |
| - F16 task tests run at temp 0.0, N=4; all other quants tested at temp 0.7, N=3 |
| |