Viggle Qwen-Image 2.1 Turbo GGUFs

Mixed-precision GGUF quantizations of Viggle/Qwen-Image-2.1-viggle-turbo for use with ComfyUI + ComfyUI-GGUF.

These files are built from Viggle's full fine-tuned transformer:

transformer/diffusion_pytorch_model.safetensors

They are not the separate Viggle LoRA. Do not stack the Viggle Turbo LoRA on top of these GGUFs.

This repository contains the diffusion transformer only. You still need the normal Qwen-Image 2.1 text encoder, VAE, and workflow components.

Important compatibility / fidelity note

The upstream Viggle transformer and the source safetensors used for conversion were verified byte-for-byte identical by SHA256:

a011a284b2dcb4b5f73aa7041b770d6bbff601bbb0850f3b6d49740b567b24e5

However, in local testing, the GGUF runtime did not reproduce the upstream diffusion_pytorch_model.safetensors output exactly even at BF16-equivalent GGUF precision.

The GGUFs are functional, but composition, identity, clothing/details, and fine structure can differ from the upstream transformer with the same prompt and seed.

This mismatch was observed before lower-bit quantization, so it should not be interpreted as Q4/Q3/etc. quantization damage alone.

These files are provided as practical low-VRAM community conversions while Qwen-Image 2.1 GGUF support continues to mature.

Recommended Viggle Turbo settings

Viggle Turbo is a distilled 4-step model.

For the ComfyUI workflow used to test these files:

Steps:      4
CFG:        1.0
Sampler:    euler
Scheduler:  simple
Negative:   blank
Denoise:    1.0

The upstream Viggle release uses a FlowMatch Euler scheduler configuration intended for the 4-step student.

Quant ladder

These builds use the same Qwen-Image 2.1 HQv3 mixed-precision policy used in my Qwen-Image 2.1 GGUF work.

File Base quant Attention Q/K/V/Out img_mlp.out img_mlp.gate_up
Q8_0-HQv3 Q8_0 Q8_0 Q8_0 Q8_0
Q6_K-HQv3 Q6_K Q8_0 Q8_0 Q8_0
Q5_K_M-HQv3 Q5_K_M Q8_0 Q8_0 Q6_K
Q4_K_M-HQv3 Q4_K_M Q8_0 Q8_0 Q5_K
Q3_K_M-HQv3 Q3_K_M Q6_K Q6_K Q4_K
Q2_K-HQv3 Q2_K Q5_K Q5_K Q3_K

The following top-level Qwen modules are retained at higher precision where applicable:

img_in
txt_in
time_text_embed
modulation
norm_out
proj_out

Small/norm tensors also remain high precision where required by the GGUF conversion path.

Build policy

Every quantization rung is generated directly from the BF16 GGUF master.

The ladder is not cascaded:

BF16 -> Q8
BF16 -> Q6
BF16 -> Q5
BF16 -> Q4
BF16 -> Q3
BF16 -> Q2

It is never:

BF16 -> Q8 -> Q6 -> Q5 -> Q4 ...

This avoids compounding quantization error from one rung into the next.

Architecture

For broad compatibility with current ComfyUI-GGUF installations, published files use:

general.architecture = qwen_image

Qwen-Image 2.1 itself uses the newer Qwen Image 2.1 transformer implementation upstream. Current GGUF support is still evolving, which is why the fidelity note above is important.

Installation

Install ComfyUI-GGUF and place the GGUF file in your normal diffusion-model directory, commonly:

ComfyUI/models/diffusion_models/

or, depending on your setup:

ComfyUI/models/unet/

Load it with the GGUF diffusion / UNet loader.

Use your existing Qwen-Image 2.1 supporting models around it.

Supporting models

You still need the normal Qwen-Image 2.1 components, including:

  • Qwen3-VL text encoder
  • Qwen-Image 2.1 VAE
  • Qwen-Image 2.1-compatible conditioning/workflow nodes

Which quant should I use?

For most low-VRAM users, start with:

Qwen-Image-2.1-viggle-turbo-Q4_K_M-HQv3.gguf

If you have more memory, try Q5_K_M, Q6_K, or Q8_0.

Q3_K_M and Q2_K are more aggressive and may show additional losses in:

  • faces
  • anatomy
  • typography
  • small texture
  • prompt adherence
  • edit fidelity

Source models

Viggle Turbo:

https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

Base Qwen-Image 2.1:

https://huggingface.co/Qwen/Qwen-Image-2.1

Related RealRebelAI GGUFs

Base Qwen-Image 2.1 GGUFs:

https://huggingface.co/realrebelai/Qwen-Image-2.1_GGUFs

This repository:

https://huggingface.co/realrebelai/Viggle_Qwen-Image-2.1-Turbo_GGUFs

Credits

  • Viggle โ€” Qwen-Image 2.1 Viggle Turbo distillation / fine-tuned transformer
  • Qwen โ€” Qwen-Image 2.1
  • City96 / ComfyUI-GGUF โ€” GGUF loading infrastructure
  • llama.cpp / ggml โ€” GGUF quantization infrastructure
  • RealRebelAI โ€” conversion, mixed-precision quantization, testing, and packaging

License

These are derivative quantizations of the upstream model and do not replace or alter the upstream license.

Review the licenses and terms of Viggle/Qwen-Image-2.1-viggle-turbo and Qwen/Qwen-Image-2.1 before redistribution or commercial use.

Downloads last month
3,110
GGUF
Model size
7B params
Architecture
qwen_image
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for realrebelai/Viggle_Qwen-Image-2.1-Turbo_GGUFs

Quantized
(3)
this model