LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
Abstract
LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results.
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Community
LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes LLaDA-Image, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and LLaDA-Image-Turbo, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.
This is insane
Some just dropped a podcast discussion about this paper on https://tensorbrife.site/
And there are also several podcast on other papers
To get direct link for this paper discussion access Listen to "A Powerful, Fully Open-Source Image Generation and Editing Model Family" on TensorBrief https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget (2026)
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing (2026)
- Exploring the Performance Frontier of Compact Unified Image Generation Models (2026)
- NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference (2026)
- InnoText: A Unified Model for Visual Text Generation and Editing (2026)
- UniSpace: Unified Visual Representation and Scalable Multimodal Modeling (2026)
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03796 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 5
inclusionAI/LLaDA-Image-Turbo
Datasets citing this paper 0
No dataset linking this paper