Instructions to use AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP") config = load_config("AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
AX Qwen3.6 35B-A3B MLX OptiQ 4-bit MTP
Parameter count: approximately 35.11B logical target parameters (35B total, approximately 3B active per token).
4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged MTP sidecar is not included in the target count.
This is a self-contained MLX package for Apple Silicon. It combines the pinned upstream OptiQ mixed-precision MoE target with an AX Engine-compatible multi-token-prediction (MTP) sidecar.
Try this model locally with AX Engine.
Attribution and changes
AutomatosX did not train Qwen3.6-35B-A3B or create its OptiQ quantization.
The target model comes from
mlx-community/Qwen3.6-35B-A3B-OptiQ-4bit
at revision 70a3aa32c7feef511182bf16aa332f37e8d82014.
AutomatosX used AX Engine's prepare_mtp_sidecar.py flow to extract the MTP
head from Qwen/Qwen3.6-35B-A3B
at revision 995ad96eacd98c81ed38be0c5b274b04031597b0, unpack the MoE expert tensors,
apply the required RMSNorm-delta normalization, quantize projections to 4-bit
with group size 64, patch the runtime config, and generate the AX manifests.
The upstream OptiQ card is preserved as UPSTREAM_README.md.
No new training or benchmark results are claimed by AutomatosX.
Package details
| Property | Value |
|---|---|
| Target format | MLX Safetensors |
| Architecture | Mixture of experts; 35B total / 3B active |
| Target quantization | OptiQ mixed 4/8-bit, group size 64 |
| OptiQ allocation | 118 components at 4-bit; 392 at 8-bit |
| Achieved target BPW | 4.5062 |
| Vision tower | Bundled BF16 sidecar |
| AX MTP sidecar | 20 logical tensors; 4-bit projections |
| Maximum draft depth | 1 |
| Configured context | 262,144 tokens |
| Intended hardware | Apple Silicon |
The package retains the upstream optiq/mtp.safetensors file and adds the
AX-prepared root mtp.safetensors. AX Engine uses the root sidecar through the
mlx_lm_extra_tensors config entry.
Download and serve
hf download AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP \
--local-dir ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP
ax-engine doctor \
--mlx-model-artifacts-dir ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP
ax-engine serve ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP --port 31418
The download is approximately 25 GB. The server exposes an OpenAI-compatible API. Consult the AX Engine repository for installation and API examples.
AX-specific files
mtp.safetensors: AX-prepared MTP sidecarmtplx_runtime.json: draft-depth and sampler guidanceax_mtp_sidecar_manifest.json: sanitized, revision-pinned provenancemodel-manifest.json: AX native target manifestconfig.json: upstream target config with the AX sidecar registration
Validation
Validated on macOS arm64 with AX Engine 6.9.0 on 2026-07-20:
- AX artifact doctor:
ready, with no model issues - Safetensors headers, data bounds, and index mappings: passed
- MTP tensor-layout exactness baseline: maximum absolute difference
0.0at context length 2,048 - Source and destination revisions are immutable and recorded in the provenance manifest
Quantization can change model quality, and speculative-decoding acceptance depends on the workload. Evaluate this package on your own tasks.
License
Apache License 2.0. See LICENSE, the original Qwen model card, and the pinned
upstream OptiQ card for limitations and responsible-use guidance.
- Downloads last month
- 217
4-bit
Model tree for AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP
Base model
Qwen/Qwen3.6-35B-A3B