Qwen3.6-35B-A3B - Q4NX for FastFlowLM (AMD Ryzen AI XDNA2)

Qwen3.6-35B-A3B, an MoE model, converted to Q4NX for FastFlowLM. This variant ships vision Q4NX weights alongside the text weights.

What is Q4NX?

Q4NX is FastFlowLM's native packed-quantization format - a rearranged Q4_1 layout tuned for the NPU matrix engine's tile sizes and memory access patterns. It is not a GGUF file and it does not run on llama.cpp or Ollama; it is meant exclusively for the FastFlowLM engine on AMD Ryzen AI NPUs.

Requirements

  • FastFlowLM >= 0.9.45 (flm CLI)
  • AMD Ryzen AI processor with XDNA2 (NPU2) - Strix Point / Ryzen AI 300 series or later
  • Linux with the XRT NPU stack installed
  • ~51 GB of unified system memory (Q4NX weights + activations + KV cache)

Files

File Purpose
model.q4nx Quantized Q4NX weights
config.json FastFlowLM model configuration
tokenizer.json Tokenizer
tokenizer_config.json Special tokens and chat template
chat_template.jinja Chat template (optional)
vision_weight.q4nx Vision tower weights (multimodal input)
flm-add.py Installer script - registers this model with FastFlowLM

Install and run

This repository ships flm-add.py, a small installer that copies the model into the FastFlowLM user directory and registers the tag qwen3.6-a3b:35b. It never modifies the system FastFlowLM install.

# one-time environment (add these to ~/.bashrc)
export FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json"
export FLM_XCLBIN_PATH="$HOME/.config/flm"

git lfs install
git clone https://huggingface.co/Atomic-Germ/Qwen3.6-35B-A3B-NPU2
cd Qwen3.6-35B-A3B-NPU2
python3 ./flm-add.py . --tag qwen3.6-a3b:35b
flm run qwen3.6-a3b:35b

Run python3 ./flm-add.py --help for all options. Without a clone, the same command works against the repo id directly:

python3 ./flm-add.py Atomic-Germ/Qwen3.6-35B-A3B-NPU2 --tag qwen3.6-a3b:35b

Kernels

FastFlowLM's NPU kernels (xclbins) are closed source and are not shipped in this repository. flm-add.py links the kernels of the official qwen3.6-moe:35b-a3b model (Qwen3.6-35B-A3B-NPU2), because this model shares the same engine family (qwen3.6-moe) and architecture.

Serve (OpenAI-compatible)

flm serve qwen3.6-a3b:35b --port 8080
curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.6-a3b:35b","messages":[{"role":"user","content":"Hello!"}],"max_tokens":256}'

Model

  • Registry tag: qwen3.6-a3b:35b
  • Engine family: qwen3.6-moe
  • Kernel source: Qwen3.6-35B-A3B-NPU2
  • Context length: 262,144 tokens (from config)
  • model.q4nx size: 23.24 GB
  • Base model: Qwen/Qwen3.6-35B-A3B
  • License: apache-2.0

Original model card

See the upstream model card for training details, benchmarks, and upstream usage. This repository only contains the Q4NX conversion for FastFlowLM.

Downloads last month
509
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Atomic-Germ/Qwen3.6-35B-A3B-NPU2

Finetuned
(200)
this model

Collection including Atomic-Germ/Qwen3.6-35B-A3B-NPU2