comfyui

The sampling acceleration from nvfp4 quantization in Krea2 is not significant.

#6
by Aca233 - opened

QQ20260625-164559FP8
QQ20260625-164632
NVFP4

RTX5060 64GB DDR4

Same here RTX5060Ti 16G, 96G RAM

Me,too.nvfp4 even slower than mxfp8. RTX5080,96G RAM

Comfy Org org

Possibly uploaded wrong version of it which doesn't allow the fast nvfp4 matmuls, re-uploaded now. For me it's ~15% faster than fp8 on 5090, not all layers could use nvfp4 matmuls due to rather bad quality loss, so it's not going to be that much faster, still should definitely not be slower at least.

Possibly uploaded wrong version of it which doesn't allow the fast nvfp4 matmuls, re-uploaded now. For me it's ~15% faster than fp8 on 5090, not all layers could use nvfp4 matmuls due to rather bad quality loss, so it's not going to be that much faster, still should definitely not be slower at least.

Thanks for the update! I just used the new nvfp4. Generating that 3840x2160 image went from 20s per step down to 15s per step.

Possibly uploaded wrong version of it which doesn't allow the fast nvfp4 matmuls, re-uploaded now. For me it's ~15% faster than fp8 on 5090, not all layers could use nvfp4 matmuls due to rather bad quality loss, so it's not going to be that much faster, still should definitely not be slower at least.

could we get a raw version too? thanks.

could we get a raw version too? thanks.

nvfp4 raw for what? raw checkpoint in bf16 best for lora training, not for image generation, because has very bad quality outputs. nvfp4 weights of raw model is completely useless.

could we get a raw version too? thanks.

nvfp4 raw for what? raw checkpoint in bf16 best for lora training, not for image generation, because has very bad quality outputs. nvfp4 weights of raw model is completely useless.

Raw with turbo lora at 0.6 strength and 12 steps looks better than turbo checkpoint

Raw with turbo lora at 0.6 strength and 12 steps looks better than turbo checkpoint

i download and test turbo lora with raw checkpoint, wow, look really better than standard turbo model. thanks for info.
waiting raw nvfp4 too

Shameless plug. I got lower inference times with a mixed precission approach at the cost of a bigger file (8.8 GB). Check https://huggingface.co/InsecureErasure/Krea2-Turbo-mixed-NVFP4.

#### Comfy-Org NVFP4 ####
[2026-07-15 17:15:04.938] got prompt
[2026-07-15 17:15:04.988] Model Krea2 prepared for dynamic VRAM loading. 7315MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 8/8 [00:17<00:00,  2.21s/it]
[2026-07-15 17:15:24.480] 0 models unloaded.
[2026-07-15 17:15:24.489] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[2026-07-15 17:15:24.959] Prompt executed in 20.02 seconds
Β·
Β·
Β·
#### InsecureErasure NVFP4-MIXED ####
[2026-07-15 17:15:49.220] got prompt
[2026-07-15 17:15:49.344] Found quantization metadata version 1
[2026-07-15 17:15:49.344] Detected mixed precision quantization
[2026-07-15 17:15:49.344] Using mixed precision operations
[2026-07-15 17:15:49.344] Native ops: mxfp8, float8_e4m3fn, nvfp4, float8_e5m2, int8_tensorwise 
[2026-07-15 17:15:49.351] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[2026-07-15 17:15:49.353] model_type FLUX
[2026-07-15 17:15:49.519] Requested to load Krea2
[2026-07-15 17:15:49.559] Model Krea2 prepared for dynamic VRAM loading. 8297MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 8/8 [00:10<00:00,  1.35s/it]
[2026-07-15 17:16:02.406] 0 models unloaded.
[2026-07-15 17:16:02.416] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[2026-07-15 17:16:02.875] Prompt executed in 13.65 seconds

A data point from the Ada side that may help frame this. On an RTX 4090 (no NVFP4 hardware), ComfyUI's --fast fp8 matmul path only adds 13 to 15 percent over the default path on the Krea 2 Turbo fp8_scaled build (1.63 to 1.88 it/s at 1024px, 8 steps). Most of the theoretical quantization speedup goes unharvested without fused low-bit kernels, the format alone does not buy the time back.

That pattern would also depress NVFP4 gains if the non-GEMM fraction (attention, norms, VAE, text encoder) dominates at your resolution. Do you have a per-component breakdown from your run?

I have fidelity numbers for all five community formats on the same protocol (seed-locked vs BF16, n=96 per format) at https://github.com/sztlink/dit-score if useful.

I made a Python script with the help of an LLM. Hope it helps.

Cheers.

prompt> a dimly lit jazz club interior, saxophone player silhouetted against smoky purple stage light

Generating: 'a dimly lit jazz club interior, saxophone player silhouetted against smoky purple stage light' | steps=8 | seed=42
──────────────────────────────────────────────────
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 8/8 [00:11<00:00,  1.42s/it]

======================================================================
                      πŸ“Š  FULL PIPELINE BREAKDOWN                      
======================================================================
  Stage                     Time (ms)    % of total  
  ───────────────────────────────────────────────
  Text Encoder              979.8        5.9         
  DiT Sampling              13884.1      83.6        
  VAE Decode                1743.5       10.5        
  ───────────────────────────────────────────────
  TOTAL                     16607.4      100.0%      
======================================================================

======================================================================
               πŸ”¬  PER-COMPONENT BREAKDOWN (inside DiT)                
======================================================================
  Category             Time (ms)    % of DiT     Calls   
  ──────────────────────────────────────────────────
  attention            2987.44      41.0         1792    
  gemm                 3643.73      50.0         832     
  norm                 638.31       8.8          528     
  other                13.00        0.2          232     
  ──────────────────────────────────────────────────
  TOTAL                7282.48      100.0        3384    
======================================================================
  attention = bmm, softmax, sdpa, attn layers
  gemm     = linear, matmul, mm, addmm, dense layers (nn.Linear)
  norm     = layer_norm, rms_norm, group_norm
  other    = everything else (silu, residual add, etc.)
======================================================================

  πŸ“Œ Non-GEMM fraction: 50.0%
  πŸ“Œ GEMM fraction:      50.0%

βœ… Saved: /srv/comfyui/ComfyUI/models/krea2_a_dimly_lit_jazz_club_interior__saxophon_s42.png

That breakdown is exactly what I was hoping someone would build, thank you. And I see a familiar jazz club in there, glad the versioned prompts are proving useful beyond my own runs.

Your mixed precision build points the same direction my failure-mode data does. Quality loss under aggressive formats concentrates in specific places, low light and fine texture break first on the Ada formats I measured, which matches what Kijai describes about some layers refusing nvfp4 matmuls gracefully.

Offer. If you want a fidelity receipt for Krea2-Turbo-mixed-NVFP4 under the dit-score protocol, seed-locked against BF16, n=96 pairs, LPIPS plus PSNR plus ImageReward delta, I will rent a Blackwell card and run it. The harness is model-agnostic and the per-pair scores get published raw. Then the speed you measured ships with a quality number attached, which is the part this whole format zoo is still missing.

Thanks for your offer. There's really no need as I still have some spare credits in vast.ai waiting to be burnt. If you could provide some guidance, I'd be glad to experiment. This quant cost me less than 1 euro. This is the command I used. As you can see, I decided to keep some layers in BF16, others in MXFP8 and the rest as NVFP4, except for those protected by convert_to_quant.

$ convert_to_quant -i krea2_turbo_bf16.safetensors \
    --nvfp4 \
    --krea2 \
    --comfy_quant \
    --save-quant-metadata \
    --custom-type mxfp8 \
    --custom-layers \
      "blocks\.(0|1|2|24|25|26)\.attn\.(wq|wk|wv|wo)\.weight|blocks\.(0|1|2|25|26|27)\.attn\.gate\.weight|blocks\.(0|1|2|3|25|26|27)\.mlp\.gate\.weight|txtfusion\.layerwise_blocks\.(0|1)\.attn\.(wq|wk|wv|wo|gate)\.weight|txtfusion\.layerwise_blocks\.(0|1)\.mlp\.gate\.weight|txtfusion\.refiner_blocks\.(0|1)\.attn\.(wq|wk|wv|gate)\.weight|txtfusion\.refiner_blocks\.(0|1)\.mlp\.(gate|up)\.weight|txtfusion\.refiner_blocks\.0\.mlp\.down\.weight" \
    --exclude-layers \
      "blocks\.27\.attn\.(wq|wk|wv)\.weight|txtfusion\.refiner_blocks\.(0|1)\.attn\.wo\.weight|txtfusion\.refiner_blocks\.1\.mlp\.down\.weight" \
    --num-iter 4000 \
    --top-p 0.35 \
    --calib-samples 8192 \
    --scale-optimization iterative \
    --scale-refinement 2 \
    --extract-lora \
    --lora-rank 64 \
    --lora-target "attn\.(wo|gate)\.weight|mlp\.(gate|down)\.weight" \
    -o krea2_turbo_mixed_nvfp4.safetensors

Even better, a fidelity number measured by a second person is worth more than one measured by me. The protocol in five steps.

  1. Generate the reference set. BF16 Turbo checkpoint, 1024x1024, 8 steps, euler, simple scheduler, cfg 1.0. Use the 12 prompts and 8 seeds from harness/prompts.json, one image per prompt+seed pair, filenames {prompt_id}__{seed}.png (96 images).
  2. Generate the variant set. Same everything, your mixed NVFP4 build instead of BF16. Same filenames into a second folder.
  3. Both sets must come from the same machine and runtime, cross-GPU seeds drift.
  4. Score. python dit_score.py --reference ref_dir --variant your_dir --prompts prompts.json --out result.json from the repo. Deps are torch, lpips, image-reward, and a pinned transformers 4.49 for ImageReward's BLIP.
  5. The JSON carries per-pair scores plus a summary (LPIPS, PSNR, ImageReward delta, worst pairs).

If you share the result I will add your build as a row in the results table with credit. For context to compare against, on my protocol int8 convrot reads LPIPS 0.066 / IR delta -0.010 and fp8 scaled reads 0.120 / -0.024.

Your layer protection choice is interesting by itself, first and last blocks shielded matches where I would expect sensitivity to concentrate. When I publish per-block error data for these checkpoints we can check how well your block list lines up with the measurements.

Sign up or log in to comment