Qwen-Image-2.1-GGUF

云碩(xCloudinfo)自製的 Qwen-Image-2.1 GGUF 量化版本,底模 Qwen/Qwen-Image-2.1。

這個 repo 解決什麼問題

llama.cpp 生態(含 ComfyUI-GGUF)原本不支援量化 Qwen-Image-2.1——llama-quantize 遇到這個架構會直接報 unknown model architecture: 'qwen_image21'。我方在 leejet 的 ComfyUI-GGUF fork 補上了 qwen_image21 架構的量化支援(C++ 端五處:架構列舉、名稱對照表、tensor map、 metadata 略過規則、量化排除規則),重新編譯 llama-quantize 後產出完整六階量化。 ComfyUI 端的讀取支援 leejet fork 其實已經內建(qwen_image21 已列在其 loader.py 的 IMG_ARCH_LIST),本次補的是「產出」這一側的缺口。

內含檔案

  • Qwen-Image-2.1-{Q8_0,Q6_K,Q5_K_M,Q4_K_M,Q3_K_M,Q2_K}.gguf —— DiT 主體,六個量化階
  • text_encoders/qwen3vl_8b_bf16.safetensors —— 原廠 Qwen3-VL-8B 文字編碼器,bf16,逐張量與原廠比對一致、未修改
  • vae/qwen_image_2.1_vae.safetensors —— 原廠 VAE,未修改

量化方法

邊界層(img_in. / txt_in. / time_text_embed. / modulation. / norm_out. / proj_out.weight) 維持 bf16 不量化,逐層注意力與 MLP 矩陣依量化階梯轉換,1D 正規化層(norm_q/norm_k)自動保留 f32 ——沿用 llama.cpp 對 FLUX、SD3 等既有 diffusion 架構的既定處理方式。

沒有 IQ4_XS / IQ2_M:llama.cpp(含此 fork)目前對所有 image-model 架構統一禁用 I-quant 系列 (Invalid quantization type for image model (Not supported)),FLUX/SD3 等既有架構一樣不能轉 I-quant。 這是上游程式碼現有的限制,不是這次轉檔特有的問題。最小的可用量化階是 Q2_K。

誠實揭露

  • 已做結構驗證:297 個張量數量與命名對齊原始 f16、邊界層確認未被量化、輸出檔案能被 gguf reader 正常解析。
  • 未做端到端生圖驗證:轉檔環境沒有安裝完整 ComfyUI,沒有實際跑一次生圖比對輸出品質。 建議初次使用先用較高量化階(Q8_0/Q6_K)跑一張對照圖確認品質,再視需求換更小的階。
  • 使用需要支援 qwen_image21 架構的 ComfyUI-GGUF 節點;目前建議用 leejet 的 fork,vanilla city96 版尚未收錄此架構。

使用方式(ComfyUI)

  1. 安裝 leejet/ComfyUI-GGUF 到 custom_nodes/。
  2. DiT GGUF 放 models/diffusion_models/,用 UNET Loader (GGUF) 節點載入。
  3. 文字編碼器放 models/text_encoders/,VAE 放 models/vae/,分別用對應的原生節點載入。

授權

繼承 Qwen-Image-2.1 原廠授權(Apache 2.0)。

Downloads last month
162
GGUF
Model size
7B params
Architecture
qwen_image21
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xCloudinfo/Qwen-Image-2.1-GGUF

Quantized
(91)
this model

Collection including xCloudinfo/Qwen-Image-2.1-GGUF