🚧 [Veda Sparse] Status Update: What It Is & What's Coming
Hi everyone, thanks for checking out this preview! Here's a quick update.
What is Veda?
Video diffusion models spend most of their time on attention: every patch of the video "looks at" every other patch across all frames. Most of those connections don't matter, but skipping the wrong ones causes artifacts.
Veda learns which ones matter. A small predictor, trained to mimic the full model, picks the ~10% of attention that counts and skips the rest. The result:
- ⚡ Faster generation, with bigger gains on longer clips (~2.2× end-to-end, ~6× attention on RTX 4090s )
- 🎨 Comparable quality at high sparsity. Even with ~90% of attention skipped, output quality stays on par with full attention.
Note: Veda is not a LoRA. It doesn't change the base model's weights or style. It only decides where the model spends its compute, and it works alongside any LoRAs.
🎉Veda now comes with ComfyUI support. 🎉
Available as a custom node on the Comfy Registry; see the README here.
What we're working on
- SGLang & vLLM-Omni integration
- Compatibility testing across NVIDIA GPUs and Apple Silicon (MLX)
- A much better official model to replace this preview
- R2VA (reference-to-video-and-audio) models
Stay tuned
Follow this page for updates, and feel free to ask questions in the Community tab. We'll reply!
This is an independent academic lab project without corporate backing, so progress may be a bit slower than we'd like. We hope to release an initial version sometime in early October. Thanks again for your patience and support! :)
📄 Paper · 🌐 Project page · 💻 Code
大家好,感谢关注!
Veda 是什么?
视频扩散模型的大部分时间都花在 自注意力 上:视频里的每个小块都要和所有帧里的其他小块逐一“attend”。其实大部分连接并不重要,但选错了就会出现画面瑕疵。
Veda 学会了判断哪些连接重要。 我们训练了一个模仿完整模型的小型预测器,它只保留约 10% 真正关键的注意力,其余全部跳过。它可以实现:
- ⚡ 生成更快,视频越长加速越明显(在15s的视频上,RTX 4090 上端到端约 2.2 倍,attention 约 6 倍)
- 🎨 高稀疏度下画质相当。 即使跳过约 90% 的注意力计算,画质仍与完整注意力相当。
注意:Veda 不是 LoRA。它不改变基础模型的权重和风格,只决定模型把算力花在哪里,可以和任意 LoRA 一起使用。
🎉 Veda 现已支持 ComfyUI!🎉
已作为自定义节点上架 Comfy Registry,使用说明请参阅 README。
我们正在做的事
- SGLang & vLLM-Omni 整合
- 各类 NVIDIA GPU 及 Apple Silicon(MLX) 的兼容性测试
- 更好的正式版模型
- R2VA(参考图生成音视频)模型
非常感谢你们的付出!我在5090上测试过了,使用你们的官方流程节点,10s i2v 768P视频只需要不到两分钟!非常震撼!非常期待r2v的版本发布!
现有的版本也可以支持R2V,我们正在传专门在R2V任务上训过的adaptor😂