lingbot-va-mot-posttrain-libero-goal-gradaccum43

Full-parameter post-train of lingbot-va-mot on the LIBERO-Goal benchmark (10 tasks, original prompts). Base architecture is the Mixture-of-Transformers (MoT) video-action model from joliachen/lingbot-va (model_mot.py).

Training config

base checkpoint checkpoints/lingbot-va-mot
dataset LIBERO-Goal (data/libero_goal_lerobot)
freeze_backbone False β€” full-parameter fine-tune, all 10.0B params (30-layer video stream + action stream + embedders/heads)
steps 4000
learning rate 1e-5
world size 2 GPUs
gradient_accumulation_steps 43 (effective batch = 2 Γ— 43 = 86)
attn_window 72
frame_chunk_size 4
guidance_scale / action_guidance_scale 5.0 / 1.0
num_inference_steps / action_num_inference_steps 25 / 50
snr_shift / action_snr_shift 5.0 / 0.05
action representation 30-dim padded vector; native 7-dim LIBERO OSC action (xyz + euler + gripper) in slots 0–6, rest zero-padded
observation cameras agentview_rgb, eye_in_hand_rgb

Architecture note: action stream width

This checkpoint (and its lingbot-va-mot base) uses a no-bottleneck MoT design: the action stream runs at the full video-stream width d_v = 3072 end-to-end (action_embedder: Linear(30 β†’ 3072), action_proj_out: Linear(3072 β†’ 30), and every per-block action module β€” action_attn1/2, action_ffn, action_norm2, action_scale_shift_table β€” is shape-identical to its video-stream counterpart). This differs from the lingbot-va paper, which describes the action stream operating through a narrower 768-dim bottleneck (30 β†’ 768 β†’ … β†’ 768 β†’ 30) rather than the full 3072-dim width used here. See lingbot-va-mot's model card for how these action-stream weights were initialized before this fine-tune.

LIBERO-Goal closed-loop evaluation (checkpoint step 4000)

10 episodes/task, closed-loop rollout in the LIBERO-Goal env:

Task Success
open the middle drawer of the cabinet 9/10
put the bowl on the stove 10/10
put the wine bottle on top of the cabinet 9/10
open the top drawer and put the bowl inside 10/10
put the bowl on top of the cabinet 8/10
push the plate to the front of the stove 10/10
put the cream cheese in the bowl 10/10
turn on the stove 10/10
put the bowl on the plate 10/10
put the wine bottle on the rack 10/10
Total 96/100 (96.0%)

Repo contents

Only transformer/ (the fine-tuned weights) is included here. The original checkpoint directory also symlinks text_encoder/ (google/umt5-xxl), tokenizer/, and vae/ (Wan2.1 AutoencoderKLWan) from the shared base checkpoint β€” those are unchanged stock components and are not duplicated in this repo. Load them from ZhuoranChen/lingbot-va-mot or the public Wan2.1 release when using this checkpoint standalone.

Downloads last month
-
Safetensors
Model size
10B params
Tensor type
BF16
Β·
Video Preview
loading