MiniMax H3 FL2VA Pruned — importance-guided mixed-IQ1 QF2 GGUF

Experimental mixed-precision GGUF quantizations of the approximately 20B-parameter pruned MiniMax H3 FL2VA diffusion denoiser. This is not the complete H3 system; its Qwen3-VL text encoder and video/audio VAEs are separate.

The repository owner confirms separate written MiniMax authorization covering this publication. That authorization is not sublicensed here; downstream users remain responsible for the official license and location-specific authorization.

Artifact and quality status

The original uniform IQ1_S and IQ1_M files are deprecated: perceptual FAIL. They load and produce valid A/V containers, but a controlled 22-frame test produced no recognizable fox or snow.

QF2 is the final importance-guided mixed-precision policy. Earlier QF1 working files are superseded internal evidence, not final artifacts.

File Bytes SHA-256 Validated scope
minimax_h3_fl2va_pruned-UD-IQ1_M-QF.gguf 10,985,133,600 c7993a8a0eb202f63f30cb410817fc0a504f3c844ea4e8dbb38599d29b0d4bdc Four-step website PASS, including motorcycle holdout; eight-step visual PASS/audio FAIL
minimax_h3_fl2va_pruned-UD-IQ1_S-QF.gguf 10,876,753,440 b12244f203e257f10c9cec56b8062aa65d627d9250c433ef633a244d610b292c Four-step website PASS, including motorcycle holdout; eight-step visual PASS/audio FAIL
minimax_h3_fl2va_pruned-IQ1_S.gguf 4,062,866,976 76f1d0bcc9052978489814004b0aaefd1a60fc8b7b8223029749085915b2946f DEPRECATED — perceptual FAIL
minimax_h3_fl2va_pruned-IQ1_M.gguf 4,532,514,336 b356d321476e7ceaf7848d0d0db7f7a82403f212313f0b7e69d1cd26c6799746 DEPRECATED — perceptual FAIL

The website four-step profile and eight-step full profile have separate verdicts. Do not describe either QF2 file as having passed full eight-step audiovisual generation.

What QF2 means

The refreshed activation importance matrix showed a steep rise through later fc1 blocks. QF2 therefore uses:

  • BF16: condition_proj.weight and protected one-dimensional norm/gain tensors
  • Q8_0: both token-refiner blocks and every eligible matrix in main blocks 46–49
  • Q4_K: mlp.fc1 in blocks 30–45, plus attention qkv_proj/out_proj and mlp.fc2 in all non-Q8 main blocks
  • IQ1_S or IQ1_M: only mlp.fc1 in blocks 0–29
  • inherited Q8_0: audio/video patch-projection weights whose shapes cannot be converted to IQ1/Q4_K
  • source precision: biases, final/output layers, and other shapes excluded by converter safeguards

These are mixed-precision derivatives. IQ1 is the lowest precision and covers 4,624,220,160 parameters; the files are not uniform one-bit models.

Controlled quick audit

Same fox prompt, seed 11, CPU RNG, 320x192, 22 frames, 24 FPS, 4 steps, CFG 1.0:

Denoiser PSNR vs Q8 SSIM vs Q8 Luma SD Laplacian Phase-grid 16 Adjacent MAD Audio clipped Human gate
Q8_0 control reference reference 41.371961 4.104783 0.0180393 3.101225 0.00% PASS
Original IQ1_S 11.680717 dB 0.502004 25.35 26.55 0.180 not recorded 4.50% FAIL: no fox or snow
Original IQ1_M 15.176109 dB 0.448978 20.56 54.10 0.403 not recorded 3.60% FAIL: no fox or snow
QF2 IQ1_M 24.800133 dB 0.876201 45.441886 3.198496 0.01557654 2.260789 0.00841864% PASS: coherent fox, no grid
QF2 IQ1_S 27.801705 dB 0.895095 44.8127344 3.1386273 0.0148840364 1.78388595 0.00336746% Low-resolution shape smeared, but no grid

PSNR and SSIM compare decoded frames against Q8 at identical settings. LPIPS was not collected and is not claimed.

Website and holdout evidence

  • QF2 IQ1_M fox, 640x384/39 frames/4 steps: clear fox; audio RMS -15.6642 dBFS; clipped fraction 0; MP4 SHA-256 e42ee46827018f86e86b67238e1806a6e9d20500a90e5e03710d3cee4fa02253.
  • QF2 IQ1_S fox, same website profile: clear walking fox; MAD 4.089718; audio RMS -16.70895 dBFS; clipped fraction 0; MP4 SHA-256 23214d5008dc58edc036d0f89b9e4a1b442be5d4265e2218d03c4416095db24a.
  • QF2 IQ1_M motorcycle holdout, seed 123: human PASS; PSNR 22.117554 dB; SSIM 0.885048; MP4 SHA-256 0b28a8c0a19ef8090fdc580a2c0f672ed1a341128ab0d2957f56b104a680c099.
  • QF2 IQ1_M motorcycle audio versus Q8: RMS -4.226/-4.380 dBFS, clipped 2.853%/2.530%, spectral correlation 0.8285/0.8136. Clipping occurs in the Q8 baseline too and is disclosed, not treated as clipping-free output.
  • QF2 IQ1_S motorcycle holdout, seed 123: human PASS with a stable red motorcycle and rider on a coastal road at sunset; PSNR 21.096775 dB; SSIM 0.862191; MP4 SHA-256 93ddcf76e9d98c78b0a51dd059a43d6bc478f451f06176ef8cda00a5a0a144f2.
  • QF2 IQ1_S motorcycle audio versus Q8: RMS -4.43595/-4.37951 dBFS; clipped 2.0843%/2.5304%; spectral correlation 0.84383/0.83414; time correlation 0.6931/0.7082. The prompt is loud and clipped in the Q8 baseline too; QF2 IQ1_S does not add excess clipping in this test.

Review samples

Profile IQ1_M IQ1_S
Quick diagnostic, 4 steps MP4 · contact sheet MP4 · contact sheet
Website fox, 4 steps MP4 · contact sheet MP4 · contact sheet
Independent motorcycle holdout, 4 steps MP4 · contact sheet MP4 · contact sheet
Full fox, 8 steps; visual PASS/audio FAIL MP4 · contact sheet MP4 · contact sheet

Independent Q8 holdout control: MP4 · contact sheet.

Eight-step limitation

Both QF2 variants produce excellent articulated walking-fox video at 640x384/39 frames/8 steps, but both fail the full-profile audio comparison:

  • IQ1_M: audio RMS -70.98 dBFS versus Q8 -50.40 dBFS; motion MAD 6.809 versus Q8 5.984.
  • IQ1_S: audio RMS -73.11065 dBFS versus Q8 -50.4007 dBFS; motion MAD 6.23523; PSNR 16.112583 dB; SSIM 0.754325; MP4 SHA-256 b8bc469b9b2f521a844b431afa750dcb5a89b9e7b88bf59123adabb0fda606af.

The deployed four-step website profile passes; eight-step audio does not.

Tensor inventory

Both actual QF2 files contain 532 tensors and 20,111,438,744 parameters. Their inventories are identical except for the IQ1 subtype.

Type Tensors Parameters
BF16 212 28,113,664
F16 102 43,642,368
F32 8 707,224
Q8_0 26 2,312,798,208
Q4_K 154 13,101,957,120
IQ1_S or IQ1_M 30 4,624,220,160

The Q8_0 count includes audio/video patch weights inherited from the Q8 conversion input. The BF16 count includes explicit condition_proj.weight plus the one-dimensional tensors protected by the patched converter.

Required companion files

Run the validated website profile

sd-cli --mode vid_gen \
  --diffusion-model minimax_h3_fl2va_pruned-UD-IQ1_M-QF.gguf \
  --llm qwen3vl_32b_minimax_h3-Q2_K_M.gguf \
  --vae minimax_h3_video_vae_fp16.safetensors \
  --audio-vae minimax_h3_audio_vae_fp32.safetensors \
  --prompt "a red fox trotting through falling snow, cinematic" \
  --width 640 --height 384 --video-frames 39 --steps 4 \
  --cfg-scale 1.0 --backend te=cpu --diffusion-fa \
  --output out.webm

--mode vid_gen, --cfg-scale 1.0, and --backend te=cpu are required for this H3 setup. Add --offload-to-cpu when needed for GPU memory.

Provenance

  • base model: MiniMaxAI/MiniMax-H3 revision 6818f6c32d12b210915e44ad56a4228c2608f160
  • pruned FL2VA lineage: Comfy-Org/MiniMax-H3 revision 014cd40f7e177756c6b2473c0d93b1c89a790dd2
  • QF2 conversion input: unsloth/MiniMax-H3-GGUF Q8_0 revision 9ee8213df85a2fcec53dec8c651a0fb1e821674a
  • converter/runtime: unslothai/stable-diffusion.cpp revision 13b9d92b5e9a1563536c9c980e700470f9ab6702
  • calibration matrix: 18,693,323 bytes; SHA-256 6ce80c23f85d50cc170b1bfafa681811e5ef26590757f1c31540485b502c67a2

Read LICENSE, NOTICE, and QUALITY_REPORT.md before use. These are unofficial Model Derivatives, not official MiniMax or Comfy-Org products, and neither organization endorses them.

Downloads last month
1,181
GGUF
Model size
20B params
Architecture
Hardware compatibility
Log In to add your hardware

1-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MarxistLeninist/MiniMax-H3-FL2VA-Pruned-IQ1-GGUF

Quantized
(54)
this model