Instructions to use alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VideoX Fun
How to use alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union with VideoX Fun:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen-Image-2.1-Fun-Controlnet-Union
Overview
Qwen-Image-2.1-Fun-Controlnet-Union is a ControlNet-Union branch for Qwen-Image 2.1 (the flow-matching text-to-image DiT). A single checkpoint drives 8 structural control conditions (Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose, Scribble) and image inpainting, without per-condition weights. The checkpoint holds only the control branch (control_img_in plus 16 control_blocks, about 7.0 GB) and is loaded on top of the base Qwen-Image 2.1 transformer.
Model Card
| Name | Description |
|---|---|
| Qwen-Image-2.1-Fun-Controlnet-Union.safetensors | ControlNet-Union branch weights for Qwen-Image 2.1. Contains only the control branch (control_img_in + 16 control_blocks, about 7.0 GB); loaded with strict=False on top of the base Qwen-Image 2.1 transformer. One checkpoint covers 8 control conditions and image inpainting. |
Model Features
- Union control over 8 conditions: one checkpoint handles Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose and Scribble control images for text-to-image generation β no per-condition checkpoint switching.
- Dense control injection: the control branch attaches a skip to every 2nd of the 32 transformer blocks (
control_layers = [0, 2, 4, β¦, 30], 16 injection points). Each control skip is added back to the main branch through zero-gatedbefore_proj/after_projprojections, giving tight structural adherence while the base model stays frozen. - Control and inpainting share one branch: the control input is widened to
control_in_dim = 129βcontrol latents (64) | mask (1) | masked-image latents (64). For pure control the mask / masked-image channels are zero-padded; for inpainting the same branch re-draws the masked region from the prompt. The two can also be combined β a control image and a mask are fed together, so the re-drawn region follows both the prompt and the given structure. - CFG-distilled fast sampling: the standalone example scripts run with
guidance_scale = 1.0(single forward pass per step, no classifier-free guidance needed). control_context_scalescales every control skip before it is added to the main branch:1.0is the strongest control (used for all results below), lower values weaken the guidance,0.0switches the control branch off.- Prompt-friendly: write a prompt that describes the whole target image; the masked region is conveyed by the mask channel, not by the text. Detailed prompts give better stability.
- Qwen-Image 2.1 encodes the prompt (and any condition image) with a Qwen3-VL text encoder + processor, and its VAE decodes to RGBA, so every preview is saved as PNG.
Supported control conditions
| Condition | Control signal |
|---|---|
| Canny | Canny edge map |
| Depth | Monocular depth map |
| Grayscale | Grayscale (luminance) image |
| HED | HED edge detection map |
| Lineart | Line-art extraction |
| MLSD | Line-segment detection map |
| Pose | DWPose skeleton |
| Scribble | Free-hand / sketch lines |
Any ordinary RGB control image at the target canvas works; the model tolerates different line thickness, thresholds and crops.
Results
All samples below are generated with num_inference_steps = 40, control_context_scale = 1.0, seed 43. In each column the top row is the control image, the bottom row is the output.
| Canny | Depth | Grayscale | HED | Lineart | MLSD | Pose | Scribble |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
Inpainting (+ control)
A masked region of the source image is re-drawn from the prompt while the rest of the frame is preserved. The mask image is white where the content should be re-generated and black where it should be kept. Because control and inpainting share the same branch, a control image (here a DWPose skeleton) is fed together with the mask, so the re-drawn region also follows the given pose.
| Source image | Mask | Pose control | Inpaint output |
![]() |
![]() |
![]() |
![]() |
Inference
Go to the VideoX-Fun repository for more details.
Please clone the VideoX-Fun repository and create the required directories:
# Clone the code
git clone https://github.com/aigc-apps/VideoX-Fun.git
# Enter VideoX-Fun's directory
cd VideoX-Fun
# Create model directories
mkdir -p models/Diffusion_Transformer
Then download the base Qwen-Image 2.1 model and this checkpoint into models/Diffusion_Transformer.
π¦ models/
βββ Diffusion_Transformer/
β βββ π Qwen-Image-2.1/
β βββ π Qwen-Image-2.1-Fun-Controlnet-Union/
β βββ Qwen-Image-2.1-Fun-Controlnet-Union.safetensors
Then edit the settings at the top of examples/qwenimage21_fun/predict_t2i_control.py (or predict_i2i_inpaint.py for inpainting) and run it.
model_name = "models/Diffusion_Transformer/Qwen-Image-2.1"
config_path = "config/qwenimage21/qwenimage21_control.yaml"
transformer_path = "models/Diffusion_Transformer/Qwen-Image-2.1-Fun-Controlnet-Union/Qwen-Image-2.1-Fun-Controlnet-Union.safetensors"
control_image = "asset/pose.jpg"
# inpaint only:
inpaint_image = "asset/8.png"
mask_image = "asset/mask.png"
prompt = "your prompt describing the whole target image"
python examples/qwenimage21_fun/predict_t2i_control.py
Notes:
config_pathmust beconfig/qwenimage21/qwenimage21_control.yaml. It builds the control branch exactly as the checkpoint expects (control_layers: [0, 2, 4, β¦, 30],control_in_dim: 129); a mismatched config silently drops or misplaces control weights and produces wrong outputs.- For pure control (no inpaint input) the pipeline zero-pads the mask / masked-image channels, so this inpaint-capable checkpoint still runs plain Canny/Depth/β¦ control correctly.
- The control checkpoint carries only the control branch; the base Qwen-Image 2.1 weights must be present in
model_name. control_context_scale = 1.0is the value the adapter expects; use lower values to loosen the structural constraint.sample_sizesets the output canvas (e.g.[1728, 992]); keep both sides as multiples of 16 so the control map is not distorted.use_kv_cache = Truecaches the text / condition-image keys after the first denoising step for a speedup at fixed resolution.- Memory: the transformer plus the Qwen3-VL text encoder do not fit a single consumer GPU fully loaded; use
model_group_offload(fastest) ormodel_cpu_offload_and_qfloat8on a single high-memory GPU.
License
This model is a derivative of Qwen-Image 2.1 and is released under the Qwen Research License. Please read the license carefully before use.
- Downloads last month
- 10,305



















