--- license: apache-2.0 library_name: mlx tags: - mlx - omlx - qwen3_5_moe - abliterated - uncensored - 3-bit base_model: Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved pipeline_tag: text-generation --- # Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX) A **sensitivity-guided ~3-bit (oQ3)** quant of the abliterated **Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved)** model, built on-device with **omlx**'s mixed-precision `oQ` quantizer. MoE arch `qwen3_5_moe` (35B total / ~3B active). Apple-Silicon MLX format. - **Effective precision:** ~3.6 bpw (≈16 GB weights) — `oQ` keeps sensitive layers higher-bit, so it punches above its nominal bit-width. - **Abliterated / uncensored.** Use responsibly; you are accountable for your outputs. ## Why this quant Despite being the *smallest/fastest* quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted): | Test | Score | |---|---| | Hard reasoning + code (8 tasks) | 8/8 | | Harder quality (multi-digit math, DP) (6) | 6/6 | | General knowledge (20 facts) | 20/20 | | Multi-step agentic tool-use (5) | 5/5 | | Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) | 7/7 | | Throughput | ~100+ tok/s | ## Run it Serve with **omlx** (or any MLX-LM runtime) on Apple Silicon: ```bash omlx serve --port 8000 # then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3" ``` Best as an **agent default with thinking OFF** (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning. *Private quant for personal use. License inherits from the base model.*