Nanosaur2-670M-OpenVINO
OpenVINO conversion of well9472/Nanosaur2-670M, a 670M-parameter illustration text-to-image diffusion transformer (DiT) with a frozen Gemma3-270M text encoder and a semantic DINOv2 VAE. Converted for Nvidia haters.
832ร1216, 50 steps, CFG 4, seed 42 โ "newest, masterpiece, 1girl, solo, (fennec ears:1.3), long blonde wavy hair, blue eyes, big fluffy tail, smile, forest, sunlight"
Files
| File | Contents | Precision |
|---|---|---|
text_encoder.xml/.bin |
Gemma3-270M text encoder (layer โ2 + final norm), 256-token input | fp32* |
diffusion_model.xml/.bin |
670M DiT โ adaLN-single, 2D RoPE, SPRINT sparse path, x-prediction; dynamic HรW | fp16 |
vae_decoder.xml/.bin |
Semantic VAE decoder (txt2img decode) | fp16 |
tokenizer.model |
Gemma3 sentencepiece model (extracted from the original TE checkpoint) | โ |
* Gemma3's residual stream reaches ~1e5 magnitude and overflows fp16 โ the same reason ComfyUI runs it in bf16/fp32. It runs once per generation, so fp32 is cheap.
Usage
With the companion inference CLI (nanosaur2-openvino):
python generate.py \
--models ./Nanosaur2-OpenVino \
--prompt "newest, masterpiece, 1girl, solo, forest, sunlight" \
--width 832 --height 1216 --steps 50 --cfg 4 --seed 42 --device GPU
Or with plain OpenVINO (pip install openvino sentencepiece protobuf pillow numpy):
import numpy as np, openvino as ov, sentencepiece as spm
core = ov.Core()
te = core.compile_model("text_encoder.xml", "GPU") # inputs: token_ids (2,256) i64, attn_mask (2,256) f32
dit = core.compile_model("diffusion_model.xml", "GPU") # inputs: x (2,64,H,W) f16, timestep (2,) f16,
# context (2,256,640) f16, token_weights (2,256) f16,
# rope_cos/rope_sin (H/16*W/16, 48) f16
vae = core.compile_model("vae_decoder.xml", "GPU") # input: z (1,64,H/16,W/16) f16 -> image in [-1, 1]
Sampling: rectified flow โ Euler, simple scheduler, shift 3.0, 50 steps, CFG 4
(the reference ComfyUI settings). See the companion repo for the full sampler.
Conversion notes
- Reimplemented from the original ComfyUI custom-node code (
nanosaur2_support/) as standalone PyTorch, exported viatorch.exportโov.convert_model. - Numerically validated against the PyTorch reference: TE max-abs 2.6e-3, DiT rel-L2 4.9e-3, VAE rel-L2 2.3e-4.
- Text is padded/masked to a fixed 256 tokens so one compiled shape serves any prompt.
- The VAE encoder (img2img) and the
alternate/path_dropguidance modes are not included; inference uses plain CFG (cfgmode) on the negative prompt.
License
MIT, inherited from the base model. Conversion produced with nanosaur2-openvino.
Model tree for HDiffusion/Nanosaur2-OpenVino
Base model
well9472/Nanosaur2-670M