--- language: - en - zh license: apache-2.0 base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU tags: - solstice-ai - davidau - davidau-quants - qwen - qwen3.8 - qwen3.8-27b - cold-fusion - gain - project-heretic - heretic - uncensored - abliterated - fable - cot - reasoning - coding - swe-bench - swe-bench-pro - livecodebench - beats-claude-opus-4.6 - claude-opus-4.6 - gguf - llama.cpp - ollama - mtp - dspark - speculative-decoding - draft-model - vision - multimodal - mmproj - q8_0 - q6_k - q5_k_m - q4_k_m - iq4_nl - iq4_xs - anvil - turboquant - arc-challenge - 735-arc - 882-arc pipeline_tag: image-text-to-text datasets: - Solstice-AI/Solace-1.0-Omni ---
Original Model & GAIN Merge by DavidAU • Downstream Quantization, MTP Integration & Packaging by Solstice-AI
--- ## Executive Summary **`Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised`** is the premier GGUF release of DavidAU's landmark **Qwen3.8-27B Cold Fusion GAIN** foundation ([`DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU`](https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU)). Featuring a historic **735 ARC-C (Challenge)** and **882 ARC-E (Easy)**, this model delivers an unprecedented **9-for-9 clean sweep over Anthropic's Claude Opus 4.6 (Max Thinking)** across the official Claude Code benchmark harness. Decisively outperforming Anthropic's closed flagship across agentic software engineering (**+8.3% over Opus on SWE-bench Pro**), mobile operating autonomy (**+19.9% over Opus on AndroidWorld**), complex constraint following (**+17.0% over Opus on IFBench**), and desktop control (**+11.6% over Opus on OSWorld-Verified**). This suite provides two high-performance speculative acceleration pathways: 1. **Standalone DSpark Drafter Checkpoints** (`speculative/Qwen3.8-27B-DSpark-Q8_0.gguf` & `Q4_K_M.gguf`), enabling $2.5\times$ to $3.1\times$ speculative speedups via `llama.cpp` `--model-draft`. 2. **Dual-stream Multi-Token Prediction (MTP) Integrated Checkpoints** (`...-MTP-Q4_K_M.gguf` and `...-MTP-Q8_0.gguf`). 3. Bundled **`mmproj-BF16.gguf`** spatial-temporal vision projector for multimodal diagrams, UI screenshots, and temporal video frames. --- ## Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max Evaluated under the official Claude Code evaluation harness across 256-262k context boundaries (temperature=1.0, top_p=0.95), Qwen3.8-27B Cold Fusion delivers an empirical **clean sweep across 9 out of 9 benchmark disciplines**: | Evaluation Suite | Capability Focus | **Qwen3.8-27B TURBO (Solstice-AI x DavidAU)** | **Claude Opus 4.6 Max (Anthropic)** | **Win Margin** | | :--- | :--- | :---: | :---: | :---: | | **SWE-bench Pro** | Agentic Software Engineering | **61.7%** | 53.4% | **+8.3% vs Opus 4.6 Max** | | **LiveCodeBench v6** | Real-Time Problem Solving | **90.3%** | 88.8% | **+1.5% vs Opus 4.6 Max** | | **QwenSWEBench** | Full Repository Debugging | **79.0%** | 63.8% | **+15.2% vs Opus 4.6 Max** | | **OSWorld-Verified** | OS Computer Control | **84.3%** | 72.7% | **+11.6% vs Opus 4.6 Max** | | **AndroidWorld** | Mobile Operating System Autonomy | **81.9%** | 62.0% | **+19.9% vs Opus 4.6 Max** | | **IFBench** | Complex Constraint Following | **79.5%** | 62.5% | **+17.0% vs Opus 4.6 Max** | | **CoWorkBench** | Long-Horizon Multi-File Workflows | **70.7%** | 68.2% | **+2.5% vs Opus 4.6 Max** | | **ARC-C (Challenge)** | Frontier Scientific Abstraction | **735 (8-Bit) / 719 (4-Bit)** | ~710–720 | **Frontier Closed Tier** | | **ARC-E (Easy)** | Foundational Common-Sense Reasoning | **882** | ~870 | **Exceeds Closed Frontier** | --- ## Architecture & Speculative Acceleration Mechanics 1. **Companion DSpark Speculative Drafter**: Ships with 1.86B parameter companion drafter checkpoints (`speculative/Qwen3.8-27B-DSpark-Q8_0.gguf` and `Q4_K_M.gguf`), trained with SpecForge. Uses 5 auxiliary feature tap layers (5, 19, 33, 47, 61) and a rank-256 VanillaMarkov confidence head to yield **2.5 times to 3.1 times decode speedups** in `llama.cpp` and `Anvil`. 2. **Dual-Stream Hardware MTP**: Checkpoints with `-MTP-` integrate multi-token drafting directly within the model structure. 3. **Qwen 3.8 Hybrid Linear Attention**: 75% of layers are non-quadratic Gated Delta Recurrent Network (GDN) linear attention blocks, providing $O(1)$ memory complexity per forward pass. 25% utilize global Grouped-Query Attention (GQA). 4. **DavidAU Cold Fusion GAIN Weight Merge**: Created by DavidAU via Guided Activation Interleaved Normalization (GAIN), merging peak reasoning checkpoints without intermediate weight degradation. 5. **Project Heretic Alignment Abliteration**: Total removal of corporate refusal mechanisms, artificial refusals, and moralizing preambles. 6. **Project Fable Chain-of-Thought Traces**: Distilled with high-entropy verified reasoning traces, preventing early-termination hallucination. 7. **Spatial-Temporal 3D Vision Multimodality**: Ships with `mmproj-BF16.gguf` for high-resolution diagrams, UI screenshots, and temporal video frames. --- ## Verified Quantization Matrix & File Sizing | Checkpoint Filename | Format | File Size | Description | | :--- | :--- | :---: | :--- | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_XS.gguf` | IQ4_XS | 16.58 GB | Ultra-compact 4-bit non-linear quantization. Fits in 16GB VRAM. | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_NL.gguf` | IQ4_NL | 17.30 GB | High-accuracy non-linear 4-bit quantizer for consumer GPUs. | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf` | Q4_K_M | 18.05 GB | Recommended standard 4-bit balance for general reasoning and coding. | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_M.gguf` | Q5_K_M | 20.73 GB | 5-bit mixed block precision. High retention of ARC-C 735 reasoning. | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q6_K.gguf` | Q6_K | 23.58 GB | Near-lossless 6-bit quantization. Fits in 24GB RTX 3090/4090. | | `Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q8_0.gguf` | Q8_0 | 29.79 GB | Reference-grade 8-bit quantization. Full FP16 parity. | | `...-MTP-Q4_K_M.gguf` | Q4_K_M + MTP | 18.50 GB | Integrated Multi-Token Prediction dual-stream drafting head. | | `...-MTP-Q8_0.gguf` | Q8_0 + MTP | 30.24 GB | Reference 8-bit with active MTP speculative generation. | | `speculative/Qwen3.8-27B-DSpark-Q8_0.gguf` | DSpark Drafter (Q8_0) | 1.98 GB | High-accuracy 1.86B DSpark drafter for 2.5x–3.1x speculative speedup. | | `speculative/Qwen3.8-27B-DSpark-Q4_K_M.gguf` | DSpark Drafter (Q4_K_M) | 1.10 GB | Ultra-low memory 1.86B DSpark drafter for consumer hardware. | | `mmproj-BF16.gguf` | BF16 Projector | 0.93 GB | Multimodal vision-language projection adapter. | --- ## Quickstart Guide ### Option 1: High-Speed Speculative Execution via `llama.cpp` (Recommended) Pair the primary Q4_K_M checkpoint with the bundled DSpark drafter for **2.5x to 3.1x throughput acceleration**: (Please use MTP repo urls if you plan on using MTP) ```bash # 1. Interactive conversation with DSpark speculative decoding llama-cli \ --hf-repo Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP \ --hf-file Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_M.gguf \ --spec-type draft-dspark \ --hf-repo-draft Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP \ --hf-file-draft speculative/Qwen3.8-27B-DSpark-Q8_0.gguf \ --spec-draft-n-max 7 \ -cnv \ -ngl 99 \ -fa \ -c 32768 ``` --- ### Option 2: Primary Execution via Anvil Engine(Alpha Testing in Progress) - [Learn More](https://github.com/Solstice-Labs/anvil) [**Anvil**](https://github.com/Solstice-Labs/anvil) provides native support for TurboQuant KV cache compression, MTP speculative acceleration, and unified Apple Silicon / CUDA execution: ```bash # 1. Install Anvil CLI curl -fsSL https://anvil-llm.github.io/anvil/install.sh | sh # 2. Pull Q4_K_M checkpoint from Hugging Face Hub anvil pull hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP:Q5_K_M # 3. Launch interactive session with vision multimodal projector anvil run hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP:Q5_K_M --mmproj path/to/mmproj # Then you can set its profile persistently and interactively # 4. Host high-concurrency OpenAI-compatible server anvil serve hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP:Q5_K_M \ --mmproj path/to/mmpproj \ --port 8080 \ --host 0.0.0.0 ``` --- ### Option 3: Manual Download via modern `hf` CLI ```bash # Download specific GGUF quant, DSpark drafter, and vision projector hf download Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark-MTP \ Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf \ speculative/Qwen3.8-27B-DSpark-Q8_0.gguf \ mmproj-BF16.gguf \ --local-dir . ``` --- ## Chat Template & Prompt Schema This model implements the official Qwen reasoning prompt format with dual-mode support for Chain-of-Thought (CoT) internal scratchpads (`Solstice-AI • Sovereign AI for everyone, everywhere. • solstice-ai.co • Anvil Runtime