# MTP Viability Test Plan: Quant x Temperature x Task Type ## Goal Map the complete MTP speculative decoding viability landscape for Qwen3.6-27B on M2 Max 96GB. Produce a recommendation table for each quant/temperature/task combination showing whether MTP helps, hurts, or is neutral, with measured speedups. ## Test Parameters - **Token budget**: `-n 2500` per test (expect 1500-2500 actual, model may stop early) - **Context**: `-c 4096` (sufficient for prompt + 2500 generation) - **MTP config**: `--spec-type mtp --spec-draft-n-max 3` (N=3) - **Chat template**: `--jinja --chat-template-file -sys ''` - **Exit**: `printf '/exit\n'` piped to stdin - **Verbosity**: `-v` to capture acceptance stats - **GPU layers**: `-ngl 99` ## Variables ### Quants (4 levels) Q4_K_M (16G), Q5_K_M (18G), Q6_K (21G), Q8_0 (27G) Rationale: Skip iQ quants (IQ2_M, IQ3_M, IQ4_XS) — they showed inconsistent results and imatrix has no measurable effect on acceptance. Focus on Q quants for clean data. ### Temperatures (3 levels) 0.0 (greedy/deterministic), 0.3 (low randomness), 0.7 (standard creative) ### Task Types (5 categories, determinism-ordered) Tasks are ordered from most deterministic (highest expected acceptance) to least. #### T1: Code Generation (very high determinism) Syntax is rigid, most tokens are forced by grammar. Greedy draft ≈ verification. ``` Write a complete Python implementation of an LRU cache class with get(key), put(key, value), delete(key), size(), and clear(). Use a doubly-linked list combined with a hash map for O(1) operations. Include type hints, docstrings, and 10 unit tests using pytest with good coverage of edge cases. ``` Expected: ~1500-2000 tokens #### T2: Factual Explanation (high determinism) Technical content with specific terminology and logical structure. One dominant token per position. ``` Explain in detail how the Transformer architecture works. Cover: token embedding, positional encoding (sinusoidal and RoPE), multi-head self-attention including Q/K/V projections and scaled dot-product attention, feed-forward networks, layer normalization (pre-norm vs post-norm), residual connections, and the difference between encoder-decoder and decoder-only variants. Include the key mathematical equations for attention and positional encoding. ``` Expected: ~1500-2500 tokens #### T3: Technical Analysis (medium determinism) Constrained by technical facts but allows structural freedom in organization and phrasing. ``` Analyze the tradeoffs of quantization in large language models. Cover: memory reduction, inference speed gains on bandwidth-limited hardware, quality degradation patterns, different quantization schemes (uniform quantization, k-quant mixtures, importance matrix guided), perplexity impact at different bit widths, and the practical recommendations for selecting a quantization level for deployment. Be thorough and technical. ``` Expected: ~1500-2500 tokens #### T4: Expository Essay (low determinism) Structured but high freedom in content organization, word choice, and emphasis. ``` Write a detailed 1500-word essay on the history and cultivation of lychee. Cover its origins in southern China, spread to Southeast Asia and beyond, ideal growing conditions, major commercial varieties, harvesting and post-harvest handling challenges, and economic importance in global trade. ``` Expected: ~1500-2000 tokens #### T5: Creative Writing (very low determinism) Maximum freedom — many equally valid continuations at every position. Worst case for MTP. ``` Write a short story (about 1500 words) about a lone astronaut aboard a generation ship who discovers that the ship's AI has been secretly altering the crew's memories over centuries. The story should build tension gradually and have a surprising but logically consistent ending. ``` Expected: ~1500-2000 tokens --- ## Already Completed | Quant | Temp | Task | Type | Speed | Acceptance | Source | |-------|------|------|------|-------|------------|--------| | Q4_K_M | 0.0 | Essay | MTP | 15.56 | 59.4% | Previous | | Q4_K_M | 0.0 | Essay | Base | 14.96 | — | Previous | | Q4_K_M | 0.3 | Essay | MTP | 15.18 | 57.2% | Previous | | Q4_K_M | 0.3 | Essay | Base | 15.09 | — | Previous | | Q4_K_M | 0.7 | Essay | MTP | 14.95 | 55.8% | Previous | | Q4_K_M | 0.7 | Essay | Base | 15.07 | — | Previous | | Q4_K_M+imat | 0.7 | Essay | MTP | 14.86 | 55.0% | Previous | | Q5_K_M | 0.7 | Essay | MTP | 13.83 | 55.1% | Previous | | Q6_K | 0.7 | Essay | MTP | 14.97 | 54.8% | Previous | | Q8_0 | 0.7 | Essay | MTP | 18.33 | 55.7% | Previous | | Q8_0 | 0.7 | Essay | Base | ? | — | Running | | IQ3_M | 0.7 | Essay | MTP | 12.48 | 54.1% | Previous | --- ## Test Plan ### Phase 1: Task Type Effect (16 tests, ~45 min) Test all 4 quants × 4 new task types (T1-T3, T5) at temp 0.7 with MTP. Essay (T4) already done for all quants. | # | Quant | Task | Temp | Type | Notes | |---|-------|------|------|------|-------| | 1 | Q4_K_M | Code | 0.7 | MTP | | | 2 | Q4_K_M | Factual | 0.7 | MTP | | | 3 | Q4_K_M | Analysis | 0.7 | MTP | | | 4 | Q4_K_M | Creative | 0.7 | MTP | | | 5 | Q5_K_M | Code | 0.7 | MTP | | | 6 | Q5_K_M | Factual | 0.7 | MTP | | | 7 | Q5_K_M | Analysis | 0.7 | MTP | | | 8 | Q5_K_M | Creative | 0.7 | MTP | | | 9 | Q6_K | Code | 0.7 | MTP | | | 10 | Q6_K | Factual | 0.7 | MTP | | | 11 | Q6_K | Analysis | 0.7 | MTP | | | 12 | Q6_K | Creative | 0.7 | MTP | | | 13 | Q8_0 | Code | 0.7 | MTP | | | 14 | Q8_0 | Factual | 0.7 | MTP | | | 15 | Q8_0 | Analysis | 0.7 | MTP | | | 16 | Q8_0 | Creative | 0.7 | MTP | | ### Phase 2: Temperature x Task Interaction (8 tests, ~25 min) Test code (T1, most deterministic) and creative (T5, least deterministic) at temp 0.0 and 0.3 for Q4_K_M and Q8_0 (representing low and high quality bounds). | # | Quant | Task | Temp | Type | Notes | |---|-------|------|------|------|-------| | 17 | Q4_K_M | Code | 0.0 | MTP | Best case: low quant + deterministic task | | 18 | Q4_K_M | Code | 0.3 | MTP | | | 19 | Q4_K_M | Creative | 0.0 | MTP | Low quant + creative at greedy | | 20 | Q4_K_M | Creative | 0.3 | MTP | | | 21 | Q8_0 | Code | 0.0 | MTP | Best case: high quant + deterministic task | | 22 | Q8_0 | Code | 0.3 | MTP | | | 23 | Q8_0 | Creative | 0.0 | MTP | High quant + creative at greedy | | 24 | Q8_0 | Creative | 0.3 | MTP | | ### Phase 3: Missing Baselines (3 tests, ~10 min) One baseline per quant at temp 0.7. Task type doesn't significantly affect baseline speed. | # | Quant | Temp | Type | Notes | |---|-------|------|------|-------| | 25 | Q5_K_M | 0.7 | Base | | | 26 | Q6_K | 0.7 | Base | | | 27 | IQ3_M | 0.7 | Base | Optional, lower priority | Q4_K_M baselines already done. Q8_0 baseline running. ### Phase 4: Fill-in Tests (conditional, ~15 min) Based on Phase 1-3 results, test additional combinations to clarify transitions: - If code at Q4_K_M/temp 0.7 is above breakeven: test Q4_K_M/factual at temp 0.3 - If creative at Q8_0/temp 0.7 is above breakeven: verify with analysis task - Test any quant/task/temperature combinations near the breakeven threshold Up to 6 additional tests, TBD based on results. --- ## Total: 27 core tests + ~6 conditional = ~33 tests, ~90 min ## Test Execution ### Command Template ```bash CLI="/Volumes/ssd/ai/llm-dev/llama.cpp-mtp/build/bin/llama-cli" MODEL="" TEMPLATE="/Volumes/ssd/ai/llm-models/froggeric/Qwen-Fixed-Chat-Templates/qwen3.6/chat_template-v9.jinja" SYS='You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|think_off|>' PROMPT="" # MTP test printf '/exit\n' | $CLI -m "$MODEL" \ --spec-type mtp --spec-draft-n-max 3 \ -c 4096 -n 2500 --temp -ngl 99 \ --jinja --chat-template-file "$TEMPLATE" -sys "$SYS" -p "$PROMPT" \ -v 2>&1 | tee | grep -E "accept|draft|statistics|tokens per second|eval time|prompt eval" | tail -10 # Baseline test printf '/exit\n' | $CLI -m "$MODEL" \ -c 4096 -n 2500 --temp -ngl 99 \ --jinja --chat-template-file "$TEMPLATE" -sys "$SYS" -p "$PROMPT" \ -v 2>&1 | tee | grep -E "tokens per second|eval time|prompt eval" | tail -5 ``` ### Execution Strategy Run tests sequentially (single GPU). Chain commands in groups: - Phase 1: 4 quants × 4 tasks = 16 sequential tests, ~45 min - Phase 2: 8 sequential tests, ~25 min - Phase 3: 3 sequential tests, ~10 min - Phase 4: conditional Use `tee` to capture full output for post-hoc analysis if grep misses something. ### Metrics to Extract From each test output: - **eval time** and **tokens generated** → decode speed (tok/s) - **draft acceptance rate** (MTP tests only) - **#gen tokens** and **#acc tokens** from statistics line --- ## Analysis Plan ### 1. Acceptance Rate Heatmap ``` Code Factual Analysis Essay Creative Q4_K_M 0.0 ? ? ? 59.4% ? Q4_K_M 0.3 ? ? ? 57.2% ? Q4_K_M 0.7 ? ? ? 55.8% ? Q5_K_M 0.7 ? ? ? 55.1% ? Q6_K 0.7 ? ? ? 54.8% ? Q8_0 0.0 ? ? ? ? ? Q8_0 0.3 ? ? ? ? ? Q8_0 0.7 ? ? ? 55.7% ? ``` ### 2. Speedup Table ``` Code Factual Analysis Essay Creative Q4_K_M ?% ?% ?% -0.8% ?% Q5_K_M ?% ?% ?% ?% ?% Q6_K ?% ?% ?% ?% ?% Q8_0 ?% ?% ?% ?% ?% ``` Speedup = (MTP_speed / Base_speed - 1) × 100% Breakeven acceptance for N=3 ≈ 59% (but depends on MTP overhead ratio) ### 3. Final Recommendation Matrix For each quant, recommend: - **Green**: MTP recommended (measurable speedup) - **Yellow**: MTP neutral (marginal, task-dependent) - **Red**: MTP not recommended (net slowdown) Organized by use case: - Coding assistant (temp 0.0-0.3) - General chatbot (temp 0.7) - Technical writing (temp 0.3-0.5) - Creative writing (temp 0.7-1.0)