Qwen3.5-35B-A3B-oQ3-mtp

This model was quantized using oQ (oMLX v0.5.2.dev1) mixed-precision quantization.

Base model: Qwen/Qwen3.5-35B-A3B

Chat template: froggeric/Qwen-Fixed-Chat-Templates

Quantization details

  • Model type: qwen3_5_moe
  • Bits: 3
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.2.dev1
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • Chat template parameter: enable_thinking=false (forced)
    • TurboQuant KV Cache: Enabled (disable before oMLX v0.5.2.dev1)
    • Lightning MTP: Enabled (key speed improvement)

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1079.9 18.31 948.3 tok/s 55.0 tok/s 3.417 337.1 tok/s 16.45 GB
pp4096/tg128 3709.2 28.43 1104.3 tok/s 35.5 tok/s 7.346 575.0 tok/s 17.18 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 55.0 tok/s 1.00x 948.3 tok/s 948.3 tok/s 1079.9 3.417
2x 70.0 tok/s 1.27x 865.9 tok/s 432.9 tok/s 2364.9 6.022
4x 98.8 tok/s 1.80x 871.2 tok/s 217.8 tok/s 4559.2 9.884

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 73.3% 22 30 31.7 No
TRUTHFULQA 83.3% 25 30 18.6 No
GSM8K 96.7% 29 30 162.8 No
MATHQA 46.7% 14 30 25.6 No
HUMANEVAL 86.7% 26 30 142.9 No
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Qwen3.5-35B-A3B-oQ3-mtp