Nanosaur2-670M-OpenVINO

OpenVINO conversion of well9472/Nanosaur2-670M, a 670M-parameter illustration text-to-image diffusion transformer (DiT) with a frozen Gemma3-270M text encoder and a semantic DINOv2 VAE. Converted for Nvidia haters.

sample

832ร—1216, 50 steps, CFG 4, seed 42 โ€” "newest, masterpiece, 1girl, solo, (fennec ears:1.3), long blonde wavy hair, blue eyes, big fluffy tail, smile, forest, sunlight"

Files

File Contents Precision
text_encoder.xml/.bin Gemma3-270M text encoder (layer โˆ’2 + final norm), 256-token input fp32*
diffusion_model.xml/.bin 670M DiT โ€” adaLN-single, 2D RoPE, SPRINT sparse path, x-prediction; dynamic Hร—W fp16
vae_decoder.xml/.bin Semantic VAE decoder (txt2img decode) fp16
tokenizer.model Gemma3 sentencepiece model (extracted from the original TE checkpoint) โ€”

* Gemma3's residual stream reaches ~1e5 magnitude and overflows fp16 โ€” the same reason ComfyUI runs it in bf16/fp32. It runs once per generation, so fp32 is cheap.

Usage

With the companion inference CLI (nanosaur2-openvino):

python generate.py \
  --models ./Nanosaur2-OpenVino \
  --prompt "newest, masterpiece, 1girl, solo, forest, sunlight" \
  --width 832 --height 1216 --steps 50 --cfg 4 --seed 42 --device GPU

Or with plain OpenVINO (pip install openvino sentencepiece protobuf pillow numpy):

import numpy as np, openvino as ov, sentencepiece as spm

core = ov.Core()
te  = core.compile_model("text_encoder.xml", "GPU")        # inputs: token_ids (2,256) i64, attn_mask (2,256) f32
dit = core.compile_model("diffusion_model.xml", "GPU")     # inputs: x (2,64,H,W) f16, timestep (2,) f16,
                                                           #        context (2,256,640) f16, token_weights (2,256) f16,
                                                           #        rope_cos/rope_sin (H/16*W/16, 48) f16
vae = core.compile_model("vae_decoder.xml", "GPU")         # input: z (1,64,H/16,W/16) f16 -> image in [-1, 1]

Sampling: rectified flow โ€” Euler, simple scheduler, shift 3.0, 50 steps, CFG 4 (the reference ComfyUI settings). See the companion repo for the full sampler.

Conversion notes

  • Reimplemented from the original ComfyUI custom-node code (nanosaur2_support/) as standalone PyTorch, exported via torch.export โ†’ ov.convert_model.
  • Numerically validated against the PyTorch reference: TE max-abs 2.6e-3, DiT rel-L2 4.9e-3, VAE rel-L2 2.3e-4.
  • Text is padded/masked to a fixed 256 tokens so one compiled shape serves any prompt.
  • The VAE encoder (img2img) and the alternate/path_drop guidance modes are not included; inference uses plain CFG (cfg mode) on the negative prompt.

License

MIT, inherited from the base model. Conversion produced with nanosaur2-openvino.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for HDiffusion/Nanosaur2-OpenVino

Finetuned
(3)
this model