How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX)

A sensitivity-guided ~3-bit (oQ3) quant of the abliterated Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved) model, built on-device with omlx's mixed-precision oQ quantizer. MoE arch qwen3_5_moe (35B total / ~3B active). Apple-Silicon MLX format.

  • Effective precision: ~3.6 bpw (≈16 GB weights) — oQ keeps sensitive layers higher-bit, so it punches above its nominal bit-width.
  • Abliterated / uncensored. Use responsibly; you are accountable for your outputs.

Why this quant

Despite being the smallest/fastest quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted):

Test Score
Hard reasoning + code (8 tasks) 8/8
Harder quality (multi-digit math, DP) (6) 6/6
General knowledge (20 facts) 20/20
Multi-step agentic tool-use (5) 5/5
Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) 7/7
Throughput ~100+ tok/s

Run it

Serve with omlx (or any MLX-LM runtime) on Apple Silicon:

omlx serve --port 8000
# then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"

Best as an agent default with thinking OFF (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning.

Private quant for personal use. License inherits from the base model.

Downloads last month
277
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support