matrix-game-gta-1step-lora

A rank-16 LoRA that turns the 3-step denoiser in Skywork/Matrix-Game-2.0's GTA branch into a 1-step one. 2.08x faster generation, for 16% less sharpness and 56% more temporal jitter. Both halves of that trade are measured; take it only if the speed is worth the picture.

Numbers

Same seed, same first frame, 129 output frames, one RTX 3090:

teacher (3 steps) student (1 step)
generate 61.0 s 29.4 s
sharpness (Laplacian variance) 1483.97 1253.80
temporal jitter (median frame delta) 2.66 4.16

Why 2.08x and not 3x: each block of three latent frames costs three denoising calls plus one pass at context_noise that refreshes the KV cache for the next block. One step leaves two calls, not one, and the context pass cannot be distilled away.

Merge the adapter before timing it. Unmerged, peft adds two matmuls per linear across 441 layers, which costs almost exactly what the skipped steps save โ€” the first measurement of this LoRA read 1.02x for that reason.

model = get_peft_model(generator.model, LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.0, bias="none",
    target_modules=["q", "k", "v", "o", "ffn.0", "ffn.2"]))
set_peft_model_state_dict(model, load_file("student_1step_r16.safetensors"))
generator.model = model.merge_and_unload()          # <- required for the speedup
pipeline.denoising_step_list = pipeline.denoising_step_list[:1]

How it was trained

2000 steps, AdamW at 1e-5, ~2.5 hours on one 3090. Loss 0.324 -> 0.0775.

The model is causal and stateful, so a plain diffusion objective does not apply: block k is denoised against a KV cache written by blocks 0..k-1. Each training step therefore replays the teacher's own procedure โ€” teacher-forced context up to block k, then teacher and student solve the same block from the same noise โ€” and the loss is between the two x0 predictions.

A known limitation follows from that: training sees one block at a time, while at inference the student's own output becomes the next block's cache. The 56% jitter increase is where that shows up.

Conditioning came from 212 clips of driving footage with keyboard/mouse pseudo-actions recovered from optical flow. Training code, evaluation, and the rest of the measurements: github.com/hwkim3330/gta6-world.

Licence

MIT, matching the base model.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kimhyunwoo/matrix-game-gta-1step-lora

Adapter
(1)
this model