How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "rayray916/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX)

A sensitivity-guided ~3-bit (oQ3) quant of the abliterated Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved) model, built on-device with omlx's mixed-precision oQ quantizer. MoE arch qwen3_5_moe (35B total / ~3B active). Apple-Silicon MLX format.

  • Effective precision: ~3.6 bpw (≈16 GB weights) — oQ keeps sensitive layers higher-bit, so it punches above its nominal bit-width.
  • Abliterated / uncensored. Use responsibly; you are accountable for your outputs.

Why this quant

Despite being the smallest/fastest quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted):

Test Score
Hard reasoning + code (8 tasks) 8/8
Harder quality (multi-digit math, DP) (6) 6/6
General knowledge (20 facts) 20/20
Multi-step agentic tool-use (5) 5/5
Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) 7/7
Throughput ~100+ tok/s

Run it

Serve with omlx (or any MLX-LM runtime) on Apple Silicon:

omlx serve --port 8000
# then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"

Best as an agent default with thinking OFF (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning.

Private quant for personal use. License inherits from the base model.

Downloads last month
277
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support