WAM DiT4DiT β€” RoboCasa Kitchen / Wan2.2 β€” 3latin2latout

Video-only (world-model) finetune of Wan2.2-TI2V-5B on RoboCasa Kitchen.

Configuration

training_mode video (no action head β€” action DiT is not trained)
latent conditioning 3latin2latout (cond latent frames -> 2 future latent frames)
video EMA saved (_video_ema_model.*, 825 tensors, fp32)
EMA schedule diffusers use_ema_warmup (inv_gamma=1.0, power=0.75, cap 0.9999)
extraction WAM_VIDEO_EMA_TRAIN_EXTRACT=0 β€” EMA is tracked+saved only, training math unchanged
GPUs 8 x H200

Contents

Milestone checkpoints are stored per step under checkpoint-<step>/.

Optimizer state (global_step*) is intentionally excluded, so these checkpoints are for inference / feature extraction / probing only β€” they cannot resume training.

Each folder contains the model shards (model-*.safetensors + index) and processor/config files. The EMA copy of the video DiT is included in the same shards under the _video_ema_model.* prefix.

Base model & license

Derived from Wan-AI/Wan2.2-TI2V-5B-Diffusers; the base model's license terms apply to this derivative.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for hmkang/wam-dit4dit-robocasa-wan22-3latin2latout

Finetuned
(23)
this model