title: DVD Image
emoji: π§
colorFrom: blue
colorTo: green
sdk: gradio
sdk_version: 5.34.2
app_file: app_dvd_image.py
pinned: false
license: mit
models:
- Zhengrui/dvd
- microsoft/TRELLIS-image-large
DVD: Discrete Voxel Diffusion for 3D Generation and Editing
TL;DR: A discrete diffusion based pipeline for generating, assessing and editing voxels, without thresholding artifacts.

DVD provides a integrated pipeline for generating, assessing and editing the sparse voxels. The generated voxels can be used as anchors for SLAT 3D generation pipelines. The pipeline is intentionally split into two stages:
- DVD generates or edits a pure
64^3voxel grid. - We leverage TRELLIS stage 2 to produce the final 3D asset formats, including mesh and 3D Gaussian outputs.
This repository supports image-conditioned generation, text-conditioned generation, image-conditioned voxel editing, text-conditioned voxel editing, command-line examples, Python APIs, and local Gradio demos.
π οΈ Installation
Install dependencies required by TRELLIS.
Install pytorch3d.
For the default fast path used by the examples:
export SPCONV_ALGO=native
export ATTN_BACKEND=flash_attn
π¦ Pretrained Models
Download the DVD checkpoints and place them under ckpts/, or load them from a
Hugging Face model repository with from_pretrained.
| Model | Link | Expected files |
|---|---|---|
| DVD image generation | https://huggingface.co/Zhengrui/dvd | dvd_img.json, dvd_img.safetensors |
| DVD image editing | https://huggingface.co/Zhengrui/dvd | dvd_img_BSP_ft.json, dvd_img_BSP_ft.safetensors |
| DVD text generation | https://huggingface.co/Zhengrui/dvd | dvd_text.json, dvd_text.safetensors |
| DVD text editing | https://huggingface.co/Zhengrui/dvd | dvd_text_BSP_ft.json, dvd_text_BSP_ft.safetensors |
| TRELLIS image stage 2 | https://huggingface.co/microsoft/TRELLIS-image-large | loaded by TrellisImageTo3DPipeline |
| TRELLIS text stage 2 | https://huggingface.co/microsoft/TRELLIS-text-large | loaded by TrellisTextTo3DPipeline |
The DVD pipelines expect the default filenames above.
π Quick Start
Python API
from PIL import Image
from dvd import (
DVDImageToVoxelPipeline,
TrellisImageTo3DPipeline,
export_cubified_voxels,
run_image_stage2_from_dvd_voxels,
)
from trellis.utils import postprocessing_utils
image = Image.open("assets/example_image/T.png")
# If you have the weights locally
# dvd = DVDImageToVoxelPipeline.from_files(
# "ckpts/dvd_img.json",
# "ckpts/dvd_img.safetensors",
# device="cuda",
# )
# Or load with from_pretrained
dvd = DVDImageToVoxelPipeline.from_pretrained('Zhengrui/dvd', variant="base", device="cuda")
voxels = dvd.sample_voxels(image, seed=42, steps=256)
export_cubified_voxels(voxels, "example_results/T_dvd_voxels.glb")
trellis = TrellisImageTo3DPipeline.from_pretrained("microsoft/TRELLIS-image-large")
trellis.cuda()
outputs = run_image_stage2_from_dvd_voxels(trellis, image, voxels, seed=42)
glb = postprocessing_utils.to_glb(outputs["gaussian"][0], outputs["mesh"][0])
glb.export("example_results/T.glb")
For Hugging Face loading:
from dvd import DVDImageToVoxelPipeline, DVDTextToVoxelPipeline
repo_id = "Zhengrui/dvd"
image_dvd = DVDImageToVoxelPipeline.from_pretrained(repo_id, variant="base", device="cuda")
image_edit_dvd = DVDImageToVoxelPipeline.from_pretrained(repo_id, variant="bsp", device="cuda")
text_dvd = DVDTextToVoxelPipeline.from_pretrained(repo_id, variant="base", device="cuda")
text_edit_dvd = DVDTextToVoxelPipeline.from_pretrained(repo_id, variant="bsp", device="cuda")
Command-Line Examples
Image-conditioned generation:
python examples_dvd_image_generation_api.py \
--image assets/example_image/T.png \
--dvd-config ckpts/dvd_img.json \
--dvd-checkpoint ckpts/dvd_img.safetensors \
--output-dir example_results
Text-conditioned generation:
python examples_dvd_text_generation_api.py \
--prompt "A small helper robot with a round body, sky-blue paint, screen face, and tiny tool arms." \
--name robot \
--dvd-config ckpts/dvd_text.json \
--dvd-checkpoint ckpts/dvd_text.safetensors \
--output-dir example_results
For editing, we provide a minima example in example_edit.ipynb. We recommend to launch apps for straightforward editing, check Gradio Demos.
Image-conditioned editing:
python examples_dvd_image_editing_api.py \
--target-image assets/example_image_edit/flower_rm.png \
--voxel-coords assets/example_voxel_edit/voxel64_typical_building_mushroom_dis.npy \
--dvd-config ckpts/dvd_img_BSP_ft.json \
--dvd-checkpoint ckpts/dvd_img_BSP_ft.safetensors \
--output-dir example_results
Text-conditioned editing:
python examples_dvd_text_editing_api.py \
--prompt "A mushroom house with a large wizard hat roof." \
--voxel-coords assets/example_voxel_edit/voxel64_typical_building_mushroom_dis.npy \
--dvd-config ckpts/dvd_text_BSP_ft.json \
--dvd-checkpoint ckpts/dvd_text_BSP_ft.safetensors \
--output-dir example_results \
--edit-y 28 64 \
--edit-z 0 64
Add --skip-stage2 in the text examples if you only want to save the DVD voxel
.npy and cubified voxel .glb.
Gradio Demos
Image generation and editing:
python app_dvd_image.py
Text generation and editing:
python app_dvd_text.py
For Hugging Face Spaces or local from_pretrained loading, set:
export DVD_MODEL_REPO=Zhengrui/dvd
β Voxel Convention
DVD stores and saves voxels in DVD convention:
samples: [B, D, H, W]
coords: [N, 4] as [batch, x, y, z]
Saved .npy coordinate files omit the batch dimension and can be loaded back
with as_voxel_output(coords, resolution=64). The helper
run_image_stage2_from_dvd_voxels and run_text_stage2_from_dvd_voxels
convert DVD voxel coordinates to TRELLIS coordinates internally.
βοΈ Acknowledgements and Third-party Licenses
This project builds heavily on TRELLIS and DUO. We sincerely thank the authors of TRELLIS and DUO for releasing their excellent research and code.
The original code in this repository is released under the MIT License, unless otherwise stated. Portions of this project are adapted from or depend on third-party projects, which remain under their respective licenses:
TRELLIS: TRELLIS models and the majority of its code are released under the MIT License. Some submodules and dependencies may have separate license terms.
DUO: DUO is released under the Apache License 2.0. Any code adapted from DUO retains its original license notices.
nvdiffrast: Utilized for rendering generated 3D assets. This package is governed by its own License.
nvdiffrec: Implements the split-sum renderer for PBR materials. This package is governed by its own License.
Please refer to LICENSE, NOTICE, and/or THIRD_PARTY_LICENSES/ for details.
π Citation
If you use this repository, please cite our project once the paper is available:
@misc{xiang2026dvddiscretevoxeldiffusion,
title={DVD: Discrete Voxel Diffusion for 3D Generation and Editing},
author={Zhengrui Xiang and Jiaqi Wu and Fupeng Sun and Heliang Zheng and Yingzhen Li},
year={2026},
eprint={2605.07971},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.07971},
}
