--- language: - en - ru license: apache-2.0 library_name: mlx base_model: - WaveCut/Qwythos-9B-v2-Heretic base_model_relation: quantized tags: - heretic - uncensored - abliteration - mlx - apple-silicon - qwen3.5 pipeline_tag: text-generation --- # Qwythos-9B-v2-Heretic-MLX-8bit **MLX 8-bit quantization** of [`WaveCut/Qwythos-9B-v2-Heretic`](https://huggingface.co/WaveCut/Qwythos-9B-v2-Heretic) — the Heretic-decensored version of [`empero-ai/Qwythos-9B-v2`](https://huggingface.co/empero-ai/Qwythos-9B-v2). Built for **Apple Silicon** (M1/M2/M3/M4). ## Specs | Field | Value | |---|---| | Bits/weight | **8.501** | | File size | ~9.5 GB | | Minimum RAM | ~11 GB unified memory | ## Quantization | Step | Tool | Version | |---|---|---| | Convert + quantize | `mlx_lm.convert` | **mlx-lm 0.31.3** (mlx 0.31.x) | | Quant mode | `affine` (default) | `-q --q-bits 8` | ```bash python -m mlx_lm.convert \ --hf-path WaveCut/Qwythos-9B-v2-Heretic \ -q --q-bits 8 \ --upload-repo WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit ``` ## Usage ```python from mlx_lm import load, generate model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit") response = generate(model, tokenizer, prompt="Hello", max_tokens=256) print(response) ``` ```bash # CLI mlx_lm.generate --model WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit --prompt "Hello" ``` ## Architecture Qwen3.5 **hybrid** — 32 blocks mixing attention and SSM (Mamba-style) layers. Supported in mlx-lm ≥ 0.31.0. ## Disclaimer Uncensored (safety alignment removed via Heretic). The original `empero-ai/Qwythos-9B-v2` maintainers are not affiliated with this derivative. Use responsibly.