File size: 3,323 Bytes
9427733
 
 
 
 
a764f22
9427733
 
12fc5dc
9427733
 
12fc5dc
9427733
 
4812302
9427733
 
 
ef87e10
 
4812302
 
 
ef87e10
9427733
 
 
 
12fc5dc
9427733
 
 
12fc5dc
 
71aa24c
12fc5dc
 
 
71aa24c
9bc4ede
12fc5dc
9427733
 
 
 
12fc5dc
 
 
 
 
 
9427733
 
 
 
 
 
 
 
 
 
 
 
12fc5dc
9427733
2eb1a6b
 
 
 
40e4b09
 
 
 
 
2eb1a6b
9427733
 
 
 
 
 
12fc5dc
 
9427733
12fc5dc
 
9427733
12fc5dc
9427733
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
license: apache-2.0
language:
  - multilingual
base_model: Qwen/Qwen3.5-122B-A10B
base_model_relation: quantized
tags:
  - mlx
  - quantization
  - mixed-precision
  - vision
  - multimodal
  - qwen3.5
  - moe
  - 122B
  - apple-silicon
library_name: mlx
pipeline_tag: image-text-to-text
model_type: qwen3_5_moe
model_name: Qwen3.5-122B-A10B-MLX-3.7bit-VL
metadata:
  parameters:
    - 122B
extra_gated_heading: "Qwen3.5-122B-A10B (MLX 3.7-bit VL)"
---

# Qwen3.5-122B-A10B-MLX-3.7bit-VL

Mixed-precision MLX quantization of [Qwen3.5-122B-A10B](https://huggingface.co/Qwen/Qwen3.5-122B-A10B) β€” Alibaba's latest MoE model with full vision support preserved in BF16.

- **3.655 BPW** | **52 GB** | Vision preserved (BF16)

## πŸš€ Hardware Optimization

This model brings 122B-class multimodal performance to Apple Silicon. By utilizing advanced mixed-precision quantization, we've compressed the model from uniform 4-bit's 65GB down to **52GB** β€” a **13GB reduction** β€” while preserving the full vision pipeline at BF16 precision for lossless image understanding.

This optimization unlocks two distinct local inference experiences:

- **64GB Unified Memory (Minimum):** The uniform 4-bit quantization weighs 65GB and simply cannot fit in 64GB at all. This quantization breaks that barrier β€” fitting a full 122B multimodal model into 64GB for the first time, pushing the hardware boundaries to make local 122B vision+language inference possible on edge devices.
- **96GB+ Unified Memory (Recommended):** Delivers an uncompromised, buttery-smooth multimodal experience. The efficient footprint frees up massive headroom for the KV cache, completely unlocking ultimate long-context capabilities for both text and vision tasks.

## Quantization

4-tier mixed precision by functional sensitivity:

| Bits | Layers | % Params | Description |
|------|--------|----------|-------------|
| BF16 | β€” | ~2% | Vision tower, norm, router, conv1d β€” preserving full visual fidelity |
| 6-bit | β€” | ~8% | Embeddings, v/o_proj, edge layers, full_attention q/k |
| 4-bit | β€” | ~3% | DeltaNet attention, shared expert |
| 3-bit | β€” | ~87% | Expert FFN (256 experts, 8 active/token) |

## Benchmark (M2 Max 96GB)

| | This (3.7bit) | Uniform 4bit |
|--|---------------|-------------|
| Model size | **52 GB** | 65 GB |
| Peak memory (ctx=4k) | **55.1 GB** | 67.7 GB |
| Prefill (1k ctx) | 219.0 tok/s | 210.5 tok/s |
| Prefill (4k ctx) | 230.6 tok/s | 227.7 tok/s |
| Generation (1k ctx) | 36.2 tok/s | 38.3 tok/s |
| Generation (4k ctx) | 33.6 tok/s | 35.8 tok/s |

> Speed is virtually identical, but the 13GB saved makes the difference between a usable and unusable long-context experience on 96GB machines.

## Quality (WikiText-2 Perplexity)

Lower is better. Evaluated on 128 sequences Γ— 2048 tokens.

| Metric | Value |
|--------|-------|
| Mean Perplexity | 5.3536 |
| Median Perplexity | 5.3639 |
| Trimmed Mean Perplexity | 5.6631 |

## Usage

```python
from mlx_vlm import load, generate

model, processor = load("MoringLabs/Qwen3.5-122B-A10B-MLX-3.7bit-VL")

# ζ–‡ζœ¬ε―Ήθ―
response = generate(model, processor, prompt="Hello!", max_tokens=200)

# 图像理解
response = generate(model, processor, prompt="Describe this image", image="photo.jpg", max_tokens=200)
print(response)
```

## License

Apache 2.0