Instructions to use kimhyunwoo/matrix-game-gta-1step-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use kimhyunwoo/matrix-game-gta-1step-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
matrix-game-gta-1step-lora
A rank-16 LoRA that turns the 3-step denoiser in
Skywork/Matrix-Game-2.0's GTA branch
into a 1-step one. 2.08x faster generation, for 16% less sharpness and 56% more temporal
jitter. Both halves of that trade are measured; take it only if the speed is worth the
picture.
Numbers
Same seed, same first frame, 129 output frames, one RTX 3090:
| teacher (3 steps) | student (1 step) | |
|---|---|---|
| generate | 61.0 s | 29.4 s |
| sharpness (Laplacian variance) | 1483.97 | 1253.80 |
| temporal jitter (median frame delta) | 2.66 | 4.16 |
Why 2.08x and not 3x: each block of three latent frames costs three denoising calls plus
one pass at context_noise that refreshes the KV cache for the next block. One step leaves
two calls, not one, and the context pass cannot be distilled away.
Merge the adapter before timing it. Unmerged, peft adds two matmuls per linear across 441 layers, which costs almost exactly what the skipped steps save โ the first measurement of this LoRA read 1.02x for that reason.
model = get_peft_model(generator.model, LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.0, bias="none",
target_modules=["q", "k", "v", "o", "ffn.0", "ffn.2"]))
set_peft_model_state_dict(model, load_file("student_1step_r16.safetensors"))
generator.model = model.merge_and_unload() # <- required for the speedup
pipeline.denoising_step_list = pipeline.denoising_step_list[:1]
How it was trained
2000 steps, AdamW at 1e-5, ~2.5 hours on one 3090. Loss 0.324 -> 0.0775.
The model is causal and stateful, so a plain diffusion objective does not apply: block k is denoised against a KV cache written by blocks 0..k-1. Each training step therefore replays the teacher's own procedure โ teacher-forced context up to block k, then teacher and student solve the same block from the same noise โ and the loss is between the two x0 predictions.
A known limitation follows from that: training sees one block at a time, while at inference the student's own output becomes the next block's cache. The 56% jitter increase is where that shows up.
Conditioning came from 212 clips of driving footage with keyboard/mouse pseudo-actions recovered from optical flow. Training code, evaluation, and the rest of the measurements: github.com/hwkim3330/gta6-world.
Licence
MIT, matching the base model.
- Downloads last month
- -
Model tree for kimhyunwoo/matrix-game-gta-1step-lora
Base model
Skywork/SkyReels-V2-I2V-1.3B-540P