Spaces:
Running on Zero
title: LiveWan
emoji: 🎞️
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.24.0
app_file: app.py
python_version: '3.10'
short_description: Streaming text-to-video you can steer mid-stream
startup_duration_timeout: 1h
pinned: false
license: apache-2.0
models:
- JonathanColetti/LiveWan
- Wan-AI/Wan2.1-T2V-1.3B
tags:
- text-to-video
- streaming
- real-time
- wan2.1
LiveWan
Streaming, steerable text-to-video from a 1.3B student distilled out of Wan2.1-T2V-14B with SF-DMD. It generates video continuously rather than as a fixed clip — 750 ms of 640×368 at a time, extended block by block — and the text conditioning can be swapped mid-stream without clearing the K/V cache, so the scene continues instead of cutting.
- Code: https://github.com/JonathanColetti/LiveWan
- Weights and data: https://huggingface.co/JonathanColetti/LiveWan
What this Space runs
app.py drives the project's own serving engine, wanstreamer.serve.engine.Engine
— the same code path livewan-serve runs locally. wanstreamer/ and
wan21_patches/ here are copies of the GitHub repo's; the Wan2.1 reference code is
cloned and patched at startup exactly as setup.sh does.
Two differences from running it locally, both forced by ZeroGPU:
- Steering is a schedule, not a button. A GPU worker is forked per request and
cannot be steered from outside while it runs, so the demo takes the swap up front
("switch to this prompt at t = N seconds"). The swap itself is the live one:
Engine.steerreplaces the cross-attention conditioning and leaves the cache in place. - No free text. Encoding it needs umt5-xxl, 11 GB on top of the 18 GB this Space already pulls at startup, so the selectors are the project's 96-prompt bank — which is the conditioning every published number refers to. Run the GitHub repo locally for free text.
torch.compile of the VAE decoder is also off, since it cannot run in a ZeroGPU
worker. That costs roughly 50 ms per block.