File size: 9,748 Bytes
b9f4ef9
 
 
 
 
 
 
 
 
 
 
 
62dba24
 
b9f4ef9
 
 
ffe28e2
b9f4ef9
 
62dba24
b9f4ef9
62dba24
 
02fe557
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62dba24
 
 
02fe557
1baab62
02fe557
b9f4ef9
02fe557
b9f4ef9
02fe557
b9f4ef9
02fe557
 
 
 
 
62dba24
02fe557
 
 
 
62dba24
02fe557
 
62dba24
02fe557
 
 
 
62dba24
02fe557
62dba24
 
 
 
 
02fe557
 
62dba24
02fe557
62dba24
02fe557
 
62dba24
 
 
 
02fe557
62dba24
02fe557
 
 
 
62dba24
02fe557
62dba24
 
 
02fe557
62dba24
02fe557
 
 
 
 
62dba24
02fe557
 
 
b9f4ef9
02fe557
b9f4ef9
02fe557
b9f4ef9
1baab62
 
 
e22fd85
02fe557
 
dfa3dbf
62dba24
 
 
 
1baab62
02fe557
b9f4ef9
62dba24
b9f4ef9
62dba24
02fe557
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
  - minimax-h3
  - comfyui
  - quantization
  - int8
  - w4
  - nvfp4
  - video
  - audio
  - fl2va
  - ref2va
---

# MiniMax-H3 Quants for ComfyUI

Community quantized diffusion-transformer checkpoints for
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
This repository provides the same eight precision profiles for both **FL2VA**
and **Ref2VA**. All files use the stock ComfyUI fused-QKV and time-table layout:
no custom node or core patch is required.

These are community conversions, not official MiniMax or ComfyOrg releases.

## 1. Choose FL2VA or Ref2VA

- **FL2VA** — text-to-audio-video, optionally conditioned by a first frame,
  last frame, or both.
- **Ref2VA** — reference-to-audio-video using reference images, video, and/or
  audio.

Download one diffusion checkpoint from the matching column below.

## 2. Choose a quant

| Profile | Direct downloads | Size | What it contains and when to use it |
|---|---|---:|---|
| **INT8 ConvRot balanced** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-balanced.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-balanced.safetensors?download=true) | 20.940 GiB | **Recommended for RTX 3090/4090 24 GiB.** 170 INT8 + 38 BF16 semantic matrices. Best tested quality/VRAM balance and fully resident in the RTX 4090 loader test. |
| **INT8 ConvRot safe** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-safe.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-safe.safetensors?download=true) | 20.330 GiB | 185 INT8 + 23 BF16 semantic matrices. Choose this on a 24 GiB RTX 30/40 card when the rest of the workflow needs more VRAM. |
| **INT8 ConvRot max** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-max.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-max.safetensors?download=true) | 21.908 GiB | 145 INT8 + 63 BF16 semantic matrices. Largest BF16 quality island; recommended for 32 GiB or more. About 0.955 GiB was offloaded in the 24 GiB RTX 4090 loader test. |
| **W8/W4 ConvRot balanced** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w8w4-convrot-balanced.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w8w4-convrot-balanced.safetensors?download=true) | 13.565 GiB | 86 W8 + 114 W4 main matrices; BF16 token refiner. Starting point for 16 GiB RTX 30/40 cards. |
| **W4 ConvRot compact** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w4-convrot-compact.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w4-convrot-compact.safetensors?download=true) | 10.067 GiB | 200 W4 main matrices + 8 INT8 token-refiner matrices. Recommended starting point for 12 GiB cards. |
| **W4 ConvRot offload** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w4-convrot-offload.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w4-convrot-offload.safetensors?download=true) | 9.708 GiB | All 208 main/refiner matrices use W4. Smallest portable profile; intended for 8 GiB cards with CPU offload. |
| **NVFP4 quality** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-nvfp4-quality.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-nvfp4-quality.safetensors?download=true) | 13.597 GiB | 170 NVFP4 + 30 BF16 main matrices; BF16 token refiner. Recommended for RTX 50/Blackwell 16–24 GiB when NVFP4 support is available. |
| **NVFP4 compact** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-nvfp4-compact.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-nvfp4-compact.safetensors?download=true) | 10.862 GiB | All 208 main/refiner matrices use block-scaled NVFP4. Smallest Blackwell-specific profile for 8–12 GiB cards. |

### Short answer

- **RTX 4090 24 GiB:** start with `int8-convrot-balanced`; use `safe` if the
  workflow needs more activation memory.
- **RTX 5090 / Blackwell 16–24 GiB:** start with `nvfp4-quality` for headroom,
  or INT8 balanced when portability matters.
- **RTX 30/40 16 GiB:** start with `w8w4-convrot-balanced`.
- **12 GiB:** start with `w4-convrot-compact`.
- **8 GiB:** use `w4-convrot-offload` and CPU offload.
- **32 GiB or more:** `int8-convrot-max` has the largest BF16 island.

Checkpoint size is not full-workflow peak VRAM. Resolution, frame count,
attention backend, text encoder, VAE, and ComfyUI offload settings also matter.
RTX 50 recommendations are architecture-based; no RTX 5090 generation run was
performed on this machine. NVFP4 here is plain block-scaled NVFP4, not AWQ.

## Measured RTX 4090 loader results

FL2VA and Ref2VA were tested independently through stock ComfyUI. Each test
also executed a real quantized INT8 projection.

| Profile | Loaded weights | Peak reserved | Free after load | Result |
|---|---:|---:|---:|---|
| `int8-convrot-safe` | 100% | 20.424 GiB | 2.072 GiB | PASS |
| `int8-convrot-balanced` | 100% | 21.025 GiB | 1.471 GiB | PASS |
| `int8-convrot-max` | 95.6% | 21.002 GiB | 1.494 GiB | PASS; about 0.955 GiB offloaded |

These are loader/kernel measurements, not full prompt-to-decoded-video peaks.

## Compatibility

All 16 files in this repository:

- retain all 50 transformer blocks;
- use fused `qkv_proj = cat(Q,K,V)` tensors expected by stock ComfyUI;
- use a rank-16 FP32, 4,097-point time table;
- retain 51 independent FP32 AdaLN projections;
- load without a custom loader or core patch in the tested ComfyUI revision.

For the original runtime FP32 time MLP and physically separate Q/K/V modules,
use the patch-required
[MiniMax-H3-DynTime-sQKV](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV)
repository instead.

<details>
<summary><strong>Advanced: exact INT8 BF16 islands and time/QKV layout</strong></summary>

INT8 weights use ConvRot/Hadamard rotation with group size 256, per-row FP32
scales, and deterministic scale search. Norms, conditioning projections, patch
projections, output heads, and other small or sensitive tensors retain their
source precision.

| Profile | BF16 attention-output blocks | BF16 MLP `fc2` blocks | Other main semantic matrices |
|---|---|---|---|
| `int8-convrot-safe` | 0, 1, 2, 3, 4, 5, 6, 7, 9, 15, 19, 38, 45, 49 | 49 | INT8 ConvRot |
| `int8-convrot-balanced` | 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 19, 20, 27, 38, 43, 44, 45, 46, 47, 49 | 39, 45, 49 | INT8 ConvRot |
| `int8-convrot-max` | all blocks 0–49 | 29, 39, 44, 45, 49 | INT8 ConvRot |

The eight token-refiner semantic matrices remain BF16 in all three INT8
profiles.

| Feature | This stock repository | Patched DynTime `s-QKV` repository |
|---|---|---|
| Attention | One fused projection call | Separate Q, K, and V calls |
| Original FP32 `time_embedder` | Absent | Present |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| `adaln_curve_basis` | Absent | FP32 `[2688,16]` |
| `adaln_curve_mean` | Absent | FP32 `[2688]` |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Core patch required |

The time table does not remove timestep conditioning. It interpolates a compact
representation of the original measured time curve. Maximum measured table
interpolation error is below `0.001%`; sampled end-to-end AdaLN relative error
is approximately `3e-7` to `4e-7` across 19 timesteps.

</details>

## Validation

Every released checkpoint passed:

1. exact key, shape, dtype, and quantization-inventory checks;
2. sampled reconstruction against its original FL2VA or Ref2VA HF shards;
3. a 19-timestep FP32 AdaLN numerical comparison;
4. complete CPU load as `MiniMaxH3Model` in clean ComfyUI commit `14b05228`;
5. remote byte-size and LFS SHA-256 verification.

Reports are stored under `reports/release_matrix/`. BF16 samples are checked
bit-for-bit. A representative INT8 QKV sample has relative L2 error `0.008814`.
A prompt-to-decoded-video perceptual A/B score has not yet been measured.

## Installation and required components

Place one selected FL2VA or Ref2VA checkpoint in:

```text
ComfyUI/models/diffusion_models/
```

A complete workflow also needs the separately maintained Qwen3-VL MiniMax-H3
text encoder and these shared VAE files:

| File | Role |
|---|---|
| `vae/minimax_h3_video_vae_fp16.safetensors` | Video latent encode/decode |
| `vae/minimax_h3_audio_vae_fp32.safetensors` | Audio latent encode/decode |

No text encoder is included in this repository.

## License and attribution

Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. This community conversion is not endorsed by MiniMax or ComfyOrg.