--- license: apache-2.0 language: - en - zh base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 tags: - gguf - quantized - qwen3 - qwen3.8 - hybrid - ssm - gated-delta-net - unsloth-dynamic - imatrix - mtp - vision model_name: Qwen3.8-27B-AEON-Ultimate-Uncensored-UD-GGUF pipeline_tag: text-generation --- # Qwen3.8-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants) **Unsloth Dynamic-style (UD)** GGUF quantizations of [AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16](https://huggingface.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16) Every quant uses **per-tensor overrides** (sensitivity-driven) + **importance matrix** (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved. --- ## Quant Comparison | File | Quant | Size | tg t/s | PPL | KL mean | KL max | KL p99.9 | |------|-------|------|--------|-----|---------|--------|----------| | F16 | F16 | 50.9 GB | 30.0 | 5.7102 | — | — | — | | UD-Q8_0 | Q8_0 | 34.7 GB | 39.9 | 5.7181 | 0.0020 | 1.83 | 0.25 | | **UD-Q6_K** | **Q6_K** | **30.6 GB** | **48.7** | **5.7288** | **0.0042** | **7.20** | **0.38** | | UD-Q5_K_M | Q5_K_M | 28.7 GB | 53.0 | 5.7243 | 0.0087 | 6.20 | 0.98 | | UD-IQ4_XS | IQ4_XS | 25.9 GB | 55.3 | 5.7288 | 0.0240 | 9.37 | 3.26 | Benchmarked on **NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM)**, llama.cpp fork ([a4501150/llama.cpp](https://github.com/a4501150/llama.cpp)), pp=512, tg=128. --- ## What Makes These Different ### SSM Recurrence Preservation Qwen3.8 is a **hybrid GatedDeltaNet + attention** model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at **source precision (F16)** — never quantized. | Tensor | Count | Precision | Rationale | |--------|-------|-----------|-----------| | `ssm_alpha`, `ssm_beta` | 96 | F16 | State update projections — error accumulates in recurrence | | `ssm_out` | 48 | F16 | Output projection feeds directly into residual stream | | `ssm_a`, `ssm_conv1d`, `ssm_dt`, `ssm_norm` | 192 | F32 | Small state tensors (llama-quantize keeps 1D/small tensors at F32) | | `attn_qkv` (SSM input projection) | 48 | F16 | Highest measured KL sensitivity | | `attn_gate` (SSM gate projection) | 48 | F16 | Second-highest measured KL sensitivity | ### Per-Tensor Sensitivity Analysis Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity: | Precision | Tensor Groups | Override Count | |-----------|--------------|----------------| | F16 | SSM recurrence, norms, biases, MTP layer | 512 | | F16 | All attention tensors (`attn_qkv`, `attn_gate`, `attn_v`, `attn_q`, `attn_k`, `attn_output`), `ffn_down` edge | 173 | | Base quant | FFN middle layers, FFN edge gate/up, embeddings | ~181 | 685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules. ### Multi-Domain Calibration + GPU Imatrix Calibrated on a balanced mix across 4 domains from 13 HF datasets: | Domain | Token Budget | Sources | |--------|-------------|---------| | General | 1M | ultrachat, OpenHermes, COIG-CQIA, LongAlpaca, pg19, froggeric/imatrix | | Code | 750K | Magicoder-Evol-Instruct-110K | | Reasoning | 750K | OpenMathInstruct-2, OpenR1-Math-220k | | Agentic | 500K | glaive-function-calling-v2, xlam-function-calling-60k, hermes-function-calling-v1 | Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation. The importance matrix is generated with a **PyTorch GPU-native generator** (`src/generate_imatrix.py`) at **65,536 context** — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via `device_map="auto"`. Per-domain imatrices are merged with equal weights (DI-MATRIX approach). ### MTP + Vision Preserved - **MTP (Multi-Token Prediction)**: Draft head (blk.64) pinned at F16. Use `--spec-type draft-mtp --spec-draft-n-max 3` for ~1.5-2x faster generation. - **Vision**: mmproj file contains the full vision encoder. Use `--mmproj` flag with llama-server for image/video understanding. --- ## Files | File | Description | Size | |------|-------------|------| | `Qwen3.8-27B-AEON-UD-Q8_0.gguf` | Highest quality quantization | 34.7 GB | | `Qwen3.8-27B-AEON-UD-Q6_K.gguf` | **Recommended** — best quality/size | 30.6 GB | | `Qwen3.8-27B-AEON-UD-Q5_K_M.gguf` | Balanced | 28.7 GB | | `Qwen3.8-27B-AEON-UD-IQ4_XS.gguf` | Smallest, for constrained VRAM | 25.9 GB | | `Qwen3.8-27B-AEON-mmproj-F16.gguf` | Vision encoder (use with `--mmproj`) | 885 MB | | `imatrix_merged.dat` | Importance matrix for requantization | 13 MB | --- ## Usage ### llama-server (recommended) ```bash # Q6_K with YaRN 512k context, 5 concurrent slots, MTP + vision llama-server \ -m Qwen3.8-27B-AEON-UD-Q6_K.gguf \ --mmproj Qwen3.8-27B-AEON-mmproj-F16.gguf \ -ngl 99 \ --flash-attn \ -c 524288 \ --parallel 5 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ -kvu \ --cache-ram -1 \ --rope-scaling yarn \ --rope-scale 2.0 \ --yarn-orig-ctx 262144 \ --override-kv "qwen35.context_length=int:524288" \ --spec-type draft-mtp \ --spec-draft-n-max 3 \ --jinja \ --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}' \ --host 0.0.0.0 --port 8080 ``` > **Note:** `--spec-type draft-mtp` requires llama.cpp b9375+. A [custom fork](https://github.com/a4501150/llama.cpp) adds DFlash speculative decoding and Blackwell-tuned flash attention. ### llama-cli ```bash llama-cli \ -m Qwen3.8-27B-AEON-UD-Q6_K.gguf \ -ngl 99 \ --flash-attn \ -c 524288 \ --rope-scaling yarn \ --rope-scale 2.0 \ --yarn-orig-ctx 262144 \ --jinja \ --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}' ``` ### Chat Template Notes - `enable_thinking` activates reasoning mode (chain-of-thought in `` blocks) - `preserve_thinking` retains reasoning blocks in conversation history - **No spaces** after colons in the JSON — Qwen3.8's template parser is whitespace-sensitive --- ## Architecture Qwen3.8-27B is a **hybrid SSM-attention** model: - 64 transformer layers + 1 MTP layer (blk.0-64) - 48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer) - 27B parameters, 24 attention heads, 4 KV heads, head dim 256 - Vocab: 248,320 tokens, native context: 262,144 tokens --- ## Quantization Pipeline Built with super-quant: 1. Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision) 2. Multi-domain calibration data from 13 HF datasets, special tokens stripped 3. GPU-native importance matrix generation (PyTorch, 65k context) + weighted merge 4. Per-tensor sensitivity analysis (KL divergence probing against F16 logits) 5. Hybrid override generation — SSM at source precision, sensitivity-driven for the rest 6. Quantize with per-tensor overrides + imatrix 7. Benchmark: throughput + perplexity + KL divergence vs F16 --- ## Links - **Base model**: [AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16](https://huggingface.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16) - **Quantization pipeline**: super-quant - **llama.cpp fork**: [a4501150/llama.cpp](https://github.com/a4501150/llama.cpp) (DFlash, MTP fixes, Blackwell FA4) ## Credits - Base model: [AEON-7](https://huggingface.co/AEON-7) - Architecture: [Qwen Team](https://huggingface.co/Qwen) - Quantization: [llama.cpp](https://github.com/ggml-org/llama.cpp) - Sensitivity methodology inspired by [APEX quant research](https://github.com/APEX-Quantization) - Calibration datasets: [HuggingFaceH4](https://huggingface.co/HuggingFaceH4), [teknium](https://huggingface.co/teknium), [NousResearch](https://huggingface.co/NousResearch), [nvidia](https://huggingface.co/nvidia), [open-r1](https://huggingface.co/open-r1), [Salesforce](https://huggingface.co/Salesforce), [glaiveai](https://huggingface.co/glaiveai), [froggeric](https://huggingface.co/froggeric) --- **License:** Apache-2.0