--- license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE base_model: - mlx-community/Qwen3.6-35B-A3B-OptiQ-4bit - Qwen/Qwen3.6-35B-A3B pipeline_tag: image-text-to-text library_name: mlx tags: - mlx - mlx-vlm - apple-silicon - qwen3.6 - mixture-of-experts - vision-language-model - conversational - quantized - mixed-precision - optiq - 4-bit - 8-bit - speculative-decoding - mtp - ax-engine - automatosx --- # AX Qwen3.6 35B-A3B MLX OptiQ 4-bit MTP > **Parameter count:** approximately 35.11B logical target parameters (35B > total, approximately 3B active per token). `4-bit` is the target quantization > precision, not a 4B model-size claim. The separately packaged MTP sidecar is > not included in the target count. This is a self-contained MLX package for Apple Silicon. It combines the pinned upstream OptiQ mixed-precision MoE target with an AX Engine-compatible multi-token-prediction (MTP) sidecar. **Try this model locally with [AX Engine](https://github.com/defai-digital/ax-engine).** ## Attribution and changes AutomatosX did **not** train Qwen3.6-35B-A3B or create its OptiQ quantization. The target model comes from [mlx-community/Qwen3.6-35B-A3B-OptiQ-4bit](https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-OptiQ-4bit) at revision `70a3aa32c7feef511182bf16aa332f37e8d82014`. AutomatosX used AX Engine's `prepare_mtp_sidecar.py` flow to extract the MTP head from [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at revision `995ad96eacd98c81ed38be0c5b274b04031597b0`, unpack the MoE expert tensors, apply the required RMSNorm-delta normalization, quantize projections to 4-bit with group size 64, patch the runtime config, and generate the AX manifests. The upstream OptiQ card is preserved as `UPSTREAM_README.md`. No new training or benchmark results are claimed by AutomatosX. ## Package details | Property | Value | | --- | --- | | Target format | MLX Safetensors | | Architecture | Mixture of experts; 35B total / 3B active | | Target quantization | OptiQ mixed 4/8-bit, group size 64 | | OptiQ allocation | 118 components at 4-bit; 392 at 8-bit | | Achieved target BPW | 4.5062 | | Vision tower | Bundled BF16 sidecar | | AX MTP sidecar | 20 logical tensors; 4-bit projections | | Maximum draft depth | 1 | | Configured context | 262,144 tokens | | Intended hardware | Apple Silicon | The package retains the upstream `optiq/mtp.safetensors` file and adds the AX-prepared root `mtp.safetensors`. AX Engine uses the root sidecar through the `mlx_lm_extra_tensors` config entry. ## Download and serve ```bash hf download AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP \ --local-dir ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP ax-engine doctor \ --mlx-model-artifacts-dir ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP ax-engine serve ./AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP --port 31418 ``` The download is approximately 25 GB. The server exposes an OpenAI-compatible API. Consult the [AX Engine repository](https://github.com/defai-digital/ax-engine) for installation and API examples. ## AX-specific files - `mtp.safetensors`: AX-prepared MTP sidecar - `mtplx_runtime.json`: draft-depth and sampler guidance - `ax_mtp_sidecar_manifest.json`: sanitized, revision-pinned provenance - `model-manifest.json`: AX native target manifest - `config.json`: upstream target config with the AX sidecar registration ## Validation Validated on macOS arm64 with AX Engine 6.9.0 on 2026-07-20: - AX artifact doctor: `ready`, with no model issues - Safetensors headers, data bounds, and index mappings: passed - MTP tensor-layout exactness baseline: maximum absolute difference `0.0` at context length 2,048 - Source and destination revisions are immutable and recorded in the provenance manifest Quantization can change model quality, and speculative-decoding acceptance depends on the workload. Evaluate this package on your own tasks. ## License Apache License 2.0. See `LICENSE`, the original Qwen model card, and the pinned upstream OptiQ card for limitations and responsible-use guidance.