Veda-MiniMax-H3-T2VA (Preview)
Paper · Project Page · Code · Turbo LoRA · Deployment Guide
Introduction
Veda is a learned sparse-attention method for video diffusion models. Attention dominates the inference cost of video diffusion, yet only a small fraction of it contributes meaningfully to the output. Veda trains a lightweight predictor, distilled from the full model, to identify the most important 10% of attention and skip the rest.
This repository provides the Veda predictor for MiniMax-H3 text-to-audio-video generation.
Highlights
- Faster inference. Up to 6.8× attention speedup and 3.1× end-to-end speedup, with larger gains on longer videos.
- Preserved quality. Visual and audio quality on par with full attention while skipping 90% of it.
- Plug-and-play. Veda is not a LoRA: it leaves the backbone weights and style untouched and only changes how attention is computed. It is compatible with any LoRA and with fine-tuned MiniMax-H3 variants.
This checkpoint is trained and evaluated with MiniMax-H3-Turbo-Lora for 8-step generation.
Samples
Comparison between full attention (left) and Veda (right) under the same prompt, seed and Turbo LoRA. All videos are 16:9, 14.4 s, generated on a single RTX PRO 6000 Blackwell GPU. Generation time is shown in each title bar.
Raw files are available under media/.
Performance
Speedup per denoising step relative to full attention (MiniMax-H3 + Turbo LoRA, 10% attention kept).
| GPU | Clip | Attention speedup | End-to-end speedup |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 16:9 · 14.4 s | 6.79× | 3.12× |
| RTX 4090 | 16:9 · 5.17 s | 4.75× | 1.77× |
| RTX 4090 | 16:9 · 10.1 s | 6.22× | 2.42× |
| RTX 4090 | 16:9 · 14.4 s | 6.31× | 2.82× |
Usage
Requires the Miowtion codebase. See AGENTS.md for hardware requirements and per-architecture notes.
git clone https://github.com/veda-sparse/Miowtion.git && cd Miowtion
pip install -e '.[gpu,encode]'
hf download MiniMaxAI/MiniMax-H3 --local-dir weights/MiniMax-H3
hf download Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview \
--local-dir weights/veda/h3-t2va-8nfe-600
Encode the prompt, then generate:
echo '{"id": "demo", "task": "t2va", "prompt": "<structured T2VA prompt>"}' > prompts.jsonl
python scripts/encode_samples.py --root weights/MiniMax-H3 \
--manifest prompts.jsonl --out artifacts/samples/demo
python scripts/generate.py \
--root weights/MiniMax-H3 --variant FL2VA \
--schedule turbo --num-steps 8 \
--adapter weights/turbo_lora/<8-step-lora>.safetensors \
--sample-cache artifacts/samples/demo --sample-id demo \
--geometry 16:9@37 --attention veda \
--predictor weights/veda/h3-t2va-8nfe-600/minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors \
--out-dir artifacts/generate/demo
Use --attention dense veda to render both modes with a side-by-side video.
On 24 GB GPUs, add --offload-blocks 50 --mlp-chunk-rows 4096.
| Option | Description |
|---|---|
--geometry |
<aspect>@<latent_t>, one of the 12 packed plans |
--keep-ratio |
Override the default 0.1 (trained at 0.1) |
--dense-steps |
Denoising steps to keep dense |
--offload-blocks |
Number of transformer blocks streamed from host memory |
--mlp-chunk-rows |
MLP chunk size; reduce if out of memory |
Loading the predictor directly:
from miowtion.veda import bundle
loaded = bundle.load(
'minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors', device='cuda')
loaded.predictor # TileScorePredictor
loaded.plans.select(geo) # tile plan for a geometry
loaded.keep_ratio # 0.1
License
This model inherits the MiniMax H3 Community License from its base model.
Citation
@inproceedings{han2026veda,
title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
and Jiang, Yi and Qi, Xiaojuan},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
- Downloads last month
- 77
Model tree for Shishir1996/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview
Base model
MiniMaxAI/MiniMax-H3