V0_1: 2,800 more training steps compared to V0. Current V0 went through : - 100 initial steps at LR 5e-5 (batch size 8) - 1100 more steps at LR 1e-4 (batch size 32) - 600 more steps at LR 1e-4 (batch size 8) - 600 more steps at LR 7.5e-5 (batch size 8) Dataset has 30,000 examples covering: - Text-to-video with no reference: 6,500 - Text-to-image with no reference: 3,500 - Still-image-to-video using: - First frame from the target video: 750 - Middle frame from the target video: 750 - Last frame from the target video: 750 - Frame from the same source but outside the target interval: 750 - Face-reference-to-video using a close-up crop of the primary face: 3,000 - Image-to-image using a different image from the same gallery: 3,000 - Image restoration and inpainting: - Masked image reference: 800 - Degraded image reference: 800 - Image outpainting using an aggressively cropped reference: 2,400 - Video outpainting: - Static spatial crop: 501 - Tracked spatial crop: 499 - Temporal video completion: - Reference video immediately before the target: 467 - Reference segment inside the complete target: 467 - Reference video immediately after the target: 466 - Audio-to-video using the target audio as conditioning: 2,000 - Video-to-audio / sound generation using a muted, low-resolution video reference: 1,400 - Structural conditioning: - Single-frame depth: 160 - Single-frame optical flow: 147 - Single-frame pose: 145 - Single-frame segmentation: 148 - Video depth: 146 - Video optical flow: 153 - Video pose: 149 - Video segmentation: 152