File size: 9,748 Bytes
b9f4ef9 62dba24 b9f4ef9 ffe28e2 b9f4ef9 62dba24 b9f4ef9 62dba24 02fe557 62dba24 02fe557 1baab62 02fe557 b9f4ef9 02fe557 b9f4ef9 02fe557 b9f4ef9 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 62dba24 02fe557 b9f4ef9 02fe557 b9f4ef9 02fe557 b9f4ef9 1baab62 e22fd85 02fe557 dfa3dbf 62dba24 1baab62 02fe557 b9f4ef9 62dba24 b9f4ef9 62dba24 02fe557 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 | ---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
- minimax-h3
- comfyui
- quantization
- int8
- w4
- nvfp4
- video
- audio
- fl2va
- ref2va
---
# MiniMax-H3 Quants for ComfyUI
Community quantized diffusion-transformer checkpoints for
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
This repository provides the same eight precision profiles for both **FL2VA**
and **Ref2VA**. All files use the stock ComfyUI fused-QKV and time-table layout:
no custom node or core patch is required.
These are community conversions, not official MiniMax or ComfyOrg releases.
## 1. Choose FL2VA or Ref2VA
- **FL2VA** — text-to-audio-video, optionally conditioned by a first frame,
last frame, or both.
- **Ref2VA** — reference-to-audio-video using reference images, video, and/or
audio.
Download one diffusion checkpoint from the matching column below.
## 2. Choose a quant
| Profile | Direct downloads | Size | What it contains and when to use it |
|---|---|---:|---|
| **INT8 ConvRot balanced** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-balanced.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-balanced.safetensors?download=true) | 20.940 GiB | **Recommended for RTX 3090/4090 24 GiB.** 170 INT8 + 38 BF16 semantic matrices. Best tested quality/VRAM balance and fully resident in the RTX 4090 loader test. |
| **INT8 ConvRot safe** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-safe.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-safe.safetensors?download=true) | 20.330 GiB | 185 INT8 + 23 BF16 semantic matrices. Choose this on a 24 GiB RTX 30/40 card when the rest of the workflow needs more VRAM. |
| **INT8 ConvRot max** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-int8-convrot-max.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-int8-convrot-max.safetensors?download=true) | 21.908 GiB | 145 INT8 + 63 BF16 semantic matrices. Largest BF16 quality island; recommended for 32 GiB or more. About 0.955 GiB was offloaded in the 24 GiB RTX 4090 loader test. |
| **W8/W4 ConvRot balanced** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w8w4-convrot-balanced.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w8w4-convrot-balanced.safetensors?download=true) | 13.565 GiB | 86 W8 + 114 W4 main matrices; BF16 token refiner. Starting point for 16 GiB RTX 30/40 cards. |
| **W4 ConvRot compact** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w4-convrot-compact.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w4-convrot-compact.safetensors?download=true) | 10.067 GiB | 200 W4 main matrices + 8 INT8 token-refiner matrices. Recommended starting point for 12 GiB cards. |
| **W4 ConvRot offload** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-w4-convrot-offload.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-w4-convrot-offload.safetensors?download=true) | 9.708 GiB | All 208 main/refiner matrices use W4. Smallest portable profile; intended for 8 GiB cards with CPU offload. |
| **NVFP4 quality** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-nvfp4-quality.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-nvfp4-quality.safetensors?download=true) | 13.597 GiB | 170 NVFP4 + 30 BF16 main matrices; BF16 token refiner. Recommended for RTX 50/Blackwell 16–24 GiB when NVFP4 support is available. |
| **NVFP4 compact** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/minimax-h3-fl2va-nvfp4-compact.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/minimax-h3-ref2va-nvfp4-compact.safetensors?download=true) | 10.862 GiB | All 208 main/refiner matrices use block-scaled NVFP4. Smallest Blackwell-specific profile for 8–12 GiB cards. |
### Short answer
- **RTX 4090 24 GiB:** start with `int8-convrot-balanced`; use `safe` if the
workflow needs more activation memory.
- **RTX 5090 / Blackwell 16–24 GiB:** start with `nvfp4-quality` for headroom,
or INT8 balanced when portability matters.
- **RTX 30/40 16 GiB:** start with `w8w4-convrot-balanced`.
- **12 GiB:** start with `w4-convrot-compact`.
- **8 GiB:** use `w4-convrot-offload` and CPU offload.
- **32 GiB or more:** `int8-convrot-max` has the largest BF16 island.
Checkpoint size is not full-workflow peak VRAM. Resolution, frame count,
attention backend, text encoder, VAE, and ComfyUI offload settings also matter.
RTX 50 recommendations are architecture-based; no RTX 5090 generation run was
performed on this machine. NVFP4 here is plain block-scaled NVFP4, not AWQ.
## Measured RTX 4090 loader results
FL2VA and Ref2VA were tested independently through stock ComfyUI. Each test
also executed a real quantized INT8 projection.
| Profile | Loaded weights | Peak reserved | Free after load | Result |
|---|---:|---:|---:|---|
| `int8-convrot-safe` | 100% | 20.424 GiB | 2.072 GiB | PASS |
| `int8-convrot-balanced` | 100% | 21.025 GiB | 1.471 GiB | PASS |
| `int8-convrot-max` | 95.6% | 21.002 GiB | 1.494 GiB | PASS; about 0.955 GiB offloaded |
These are loader/kernel measurements, not full prompt-to-decoded-video peaks.
## Compatibility
All 16 files in this repository:
- retain all 50 transformer blocks;
- use fused `qkv_proj = cat(Q,K,V)` tensors expected by stock ComfyUI;
- use a rank-16 FP32, 4,097-point time table;
- retain 51 independent FP32 AdaLN projections;
- load without a custom loader or core patch in the tested ComfyUI revision.
For the original runtime FP32 time MLP and physically separate Q/K/V modules,
use the patch-required
[MiniMax-H3-DynTime-sQKV](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV)
repository instead.
<details>
<summary><strong>Advanced: exact INT8 BF16 islands and time/QKV layout</strong></summary>
INT8 weights use ConvRot/Hadamard rotation with group size 256, per-row FP32
scales, and deterministic scale search. Norms, conditioning projections, patch
projections, output heads, and other small or sensitive tensors retain their
source precision.
| Profile | BF16 attention-output blocks | BF16 MLP `fc2` blocks | Other main semantic matrices |
|---|---|---|---|
| `int8-convrot-safe` | 0, 1, 2, 3, 4, 5, 6, 7, 9, 15, 19, 38, 45, 49 | 49 | INT8 ConvRot |
| `int8-convrot-balanced` | 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 19, 20, 27, 38, 43, 44, 45, 46, 47, 49 | 39, 45, 49 | INT8 ConvRot |
| `int8-convrot-max` | all blocks 0–49 | 29, 39, 44, 45, 49 | INT8 ConvRot |
The eight token-refiner semantic matrices remain BF16 in all three INT8
profiles.
| Feature | This stock repository | Patched DynTime `s-QKV` repository |
|---|---|---|
| Attention | One fused projection call | Separate Q, K, and V calls |
| Original FP32 `time_embedder` | Absent | Present |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| `adaln_curve_basis` | Absent | FP32 `[2688,16]` |
| `adaln_curve_mean` | Absent | FP32 `[2688]` |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Core patch required |
The time table does not remove timestep conditioning. It interpolates a compact
representation of the original measured time curve. Maximum measured table
interpolation error is below `0.001%`; sampled end-to-end AdaLN relative error
is approximately `3e-7` to `4e-7` across 19 timesteps.
</details>
## Validation
Every released checkpoint passed:
1. exact key, shape, dtype, and quantization-inventory checks;
2. sampled reconstruction against its original FL2VA or Ref2VA HF shards;
3. a 19-timestep FP32 AdaLN numerical comparison;
4. complete CPU load as `MiniMaxH3Model` in clean ComfyUI commit `14b05228`;
5. remote byte-size and LFS SHA-256 verification.
Reports are stored under `reports/release_matrix/`. BF16 samples are checked
bit-for-bit. A representative INT8 QKV sample has relative L2 error `0.008814`.
A prompt-to-decoded-video perceptual A/B score has not yet been measured.
## Installation and required components
Place one selected FL2VA or Ref2VA checkpoint in:
```text
ComfyUI/models/diffusion_models/
```
A complete workflow also needs the separately maintained Qwen3-VL MiniMax-H3
text encoder and these shared VAE files:
| File | Role |
|---|---|
| `vae/minimax_h3_video_vae_fp16.safetensors` | Video latent encode/decode |
| `vae/minimax_h3_audio_vae_fp32.safetensors` | Audio latent encode/decode |
No text encoder is included in this repository.
## License and attribution
Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. This community conversion is not endorsed by MiniMax or ComfyOrg.
|