Generative Refinement Networks for Visual Synthesis
Paper β’ 2604.13030 β’ Published β’ 15
An unofficial conversion of the GRN (Generative Refinement Networks) 2B text-to-image checkpoint and its HBQ tokenizer from PyTorch pickles to safetensors. The weights are ByteDance's; nothing was retrained or modified.
For anything about the model itself β usage, training, results, citation β see the official pages above.
| File | Converted from | Notes |
|---|---|---|
grn-t2i-2b-bf16.safetensors |
GRN_T2I_2B_FSA_251600.pth |
the transformer, cast to BF16 |
grn-hbq-tokenizer.safetensors |
HBQ_image_video_tokenizer_64dim_M4_20260626.ckpt |
the tokenizer (EMA weights), as stored |
grn_convert.py |
the script used |
Tensor names are unchanged. The text encoder (umT5-XXL) is not included; use the one from the official repository or any standard umT5-XXL encoder.
Done on the CPU with torch.load(..., weights_only=True), keeping tensors only:
python grn_convert.py GRN_T2I_2B_FSA_251600.pth grn-t2i-2b-bf16.safetensors bf16
python grn_convert.py HBQ_image_video_tokenizer_64dim_M4_20260626.ckpt grn-hbq-tokenizer.safetensors
MIT, as the original release.
Base model
bytedance-research/GRN