Instructions to use rosingrind/Qwen3.6-35B-A3B-oQ6e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use rosingrind/Qwen3.6-35B-A3B-oQ6e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3.6-35B-A3B-oQ6e-mtp rosingrind/Qwen3.6-35B-A3B-oQ6e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
rosingrind/Qwen3.6-35B-A3B-oQ6e-mtp performance in oMLXL
oMLX - LLM inference, performance benchmark Macbook Pro M4 48 GB unified RAM
Benchmark Model: Qwen3.6-35B-A3B-oQ6e-mtp
Engine: Force mlx-lm
Single Request Results
Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1230.5 22.62 832.2 tok/s 44.6 tok/s 4.113 280.1 tok/s 28.62 GB
pp4096/tg128 4857.4 76.81 843.3 tok/s 19.5 tok/s 5.069 808.7 tok/s 29.45 GB
pp8192/tg128 10053.7 19.15 814.8 tok/s 52.6 tok/s 12.500 665.6 tok/s 29.80 GB
pp16384/tg128 21613.0 20.06 758.1 tok/s 50.2 tok/s 24.175 683.0 tok/s 30.45 GB
pp32768/tg128 52608.2 23.89 622.9 tok/s 42.2 tok/s 55.657 591.0 tok/s 31.79 GB