How to use from
MLX LM
Generate or start a chat session
# Install MLX LM
uv tool install mlx-lm
# Interactive chat REPL
mlx_lm.chat --model "rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"
Run an OpenAI-compatible server
# Install MLX LM
uv tool install mlx-lm
# Start the server
mlx_lm.server --model "rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
   -H "Content-Type: application/json" \
   --data '{
     "model": "rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3",
     "messages": [
       {"role": "user", "content": "Hello"}
     ]
   }'
Quick Links

Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX)

A sensitivity-guided ~3-bit (oQ3) quant of the abliterated Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved) model, built on-device with omlx's mixed-precision oQ quantizer. MoE arch qwen3_5_moe (35B total / ~3B active). Apple-Silicon MLX format.

  • Effective precision: ~3.6 bpw (≈16 GB weights) — oQ keeps sensitive layers higher-bit, so it punches above its nominal bit-width.
  • Abliterated / uncensored. Use responsibly; you are accountable for your outputs.

Why this quant

Despite being the smallest/fastest quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted):

Test Score
Hard reasoning + code (8 tasks) 8/8
Harder quality (multi-digit math, DP) (6) 6/6
General knowledge (20 facts) 20/20
Multi-step agentic tool-use (5) 5/5
Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) 7/7
Throughput ~100+ tok/s

Run it

Serve with omlx (or any MLX-LM runtime) on Apple Silicon:

omlx serve --port 8000
# then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"

Best as an agent default with thinking OFF (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning.

Private quant for personal use. License inherits from the base model.

Downloads last month
277
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support