StarVLA QwenOFT Qwen3-VL-4B for RoboDojo (100k)

This directory contains one StarVLA QwenOFT checkpoint initialized from Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture. The VLM, VLM interface, and MLP action head were trained end to end.

Model details

Item Value
Framework StarVLA QwenOFT
Base VLM Qwen3-VL-4B-Instruct
Action model MLP, hidden size 2,560
Objective L1 action regression
Action representation 14D absolute joint position (abs_qpos)
Action horizon 50
State dimension 14
Camera input Head, left wrist, right wrist; resized to 224 x 224
Checkpoint step 100,000
Checkpoint format Complete StarVLA framework state dict (.pt)

Files

README.md
config.yaml
config.full.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt

Only the requested 100k checkpoint is included. Keep the saved configuration and normalization statistics beside the checkpoint.

Training details

Setting Value
Dataset mixture robodojo_v21_all_h50_q99
Training tasks 35
Per-GPU batch size 16
Gradient accumulation 1
Frozen modules None
Optimizer AdamW, betas (0.9, 0.95), epsilon 1e-8
VLM learning rate 1e-5
VLM-interface learning rate 1e-5
Action-model learning rate 1e-4
Schedule Cosine, 5,000 warmup steps, minimum LR 5e-7
Gradient checkpointing Enabled
Random seed 42

Official RoboDojo evaluation

All policies below use the official complete 42-task protocol: 50 episodes per task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This directory's policy is bolded. Higher is better for both SR and Score.

Group summary

Policy Average Generalization Precision Long-Horizon Memory Open
QwenOFT 4.86 / 8.01 4.33 / 6.42 11.75 / 17.54 5.50 / 12.95 1.67 / 1.77 0.50 / 0.60
QwenGR00T 3.81 / 7.35 3.50 / 6.52 5.75 / 10.09 6.50 / 15.46 3.33 / 4.37 0.00 / 0.00
QwenPI_v3 6.19 / 9.60 4.17 / 7.28 14.00 / 19.06 10.00 / 17.84 2.00 / 2.32 0.75 / 0.88

Task details

Task values are SR (%) / Score; each task uses 50 episodes.

Evaluation group / task QwenOFT QwenGR00T QwenPI_v3
Generalization 4.33 / 6.42 3.50 / 6.52 4.17 / 7.28
stack_bowls 18.00 / 21.00 10.00 / 14.80 14.00 / 16.70
push_T 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
pack_objects_into_box 0.00 / 3.10 0.00 / 7.80 0.00 / 6.80
fold_clothes 10.00 / 12.80 8.00 / 12.40 2.00 / 9.60
hang_mugs 0.00 / 3.60 0.00 / 3.00 0.00 / 3.50
sweep_blocks 0.00 / 0.00 0.00 / 0.00 2.00 / 2.00
pour_liquid_into_cup 14.00 / 14.00 14.00 / 14.00 12.00 / 12.00
make_toast 0.00 / 1.00 2.00 / 5.00 2.00 / 5.00
arrange_largest_number 0.00 / 1.90 2.00 / 4.10 2.00 / 5.70
sort_nesting_dolls_by_size 0.00 / 0.00 4.00 / 4.00 6.00 / 6.00
store_laptop_and_headphones 4.00 / 11.20 0.00 / 7.20 2.00 / 8.40
stack_blocks 6.00 / 8.40 2.00 / 5.90 8.00 / 11.60
Precision 11.75 / 17.54 5.75 / 10.09 14.00 / 19.06
fasten_screws 4.00 / 8.00 0.00 / 2.00 0.00 / 6.00
plug_in_charger 6.00 / 6.00 2.00 / 2.00 4.00 / 4.00
insert_tubes 40.00 / 51.60 28.00 / 40.40 44.00 / 56.80
pour_balls_into_vase 8.00 / 8.00 0.00 / 0.00 2.00 / 2.00
play_Xylophone 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
deposit_coin 0.00 / 3.20 2.00 / 5.60 6.00 / 7.60
insert_key 0.00 / 12.90 0.00 / 9.90 0.00 / 11.10
build_tower 36.00 / 50.60 14.00 / 20.80 56.00 / 65.00
Long-Horizon 5.50 / 12.95 6.50 / 15.46 10.00 / 17.84
put_bottles_into_dustbin 22.00 / 40.90 26.00 / 44.40 64.00 / 73.60
fill_pen_holder 4.00 / 11.70 6.00 / 14.40 8.00 / 23.00
classify_objects 2.00 / 5.50 6.00 / 11.50 0.00 / 7.50
play_tic_tac_toe 0.00 / 12.40 2.00 / 16.40 0.00 / 6.80
fill_egg_holder 0.00 / 0.60 0.00 / 0.00 0.00 / 0.80
organize_table 0.00 / 16.50 0.00 / 25.00 4.00 / 27.00
make_kong 16.00 / 16.00 12.00 / 12.00 4.00 / 4.00
play_stacking_toy 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
Memory 1.67 / 1.77 3.33 / 4.37 2.00 / 2.32
cover_blocks 0.00 / 0.60 6.00 / 12.10 0.00 / 1.50
match_and_pick_from_conveyor 10.00 / 10.00 14.00 / 14.00 12.00 / 12.00
swap_blocks 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
swap_T 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
press_by_number 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
imitate_sorting_sequence 0.00 / 0.00 0.00 / 0.10 0.00 / 0.40
Open 0.50 / 0.60 0.00 / 0.00 0.75 / 0.88
align_blocks 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
general_pickup 4.00 / 4.00 0.00 / 0.00 6.00 / 6.00
stack_blocks_by_language 0.00 / 0.80 0.00 / 0.00 0.00 / 0.80
solve_equation 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
classify_objects_by_language 0.00 / 0.00 0.00 / 0.00 0.00 / 0.20
pick_from_conveyor_by_image 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
store_tools_in_toolbox 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
pour_by_language 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00

Evaluation

Start the StarVLA model server:

export CKPT=/path/to/MODEL_DIR/checkpoints/steps_100000_pytorch_model.pt
python deployment/model_server/server_policy.py \
  --ckpt_path "$CKPT" \
  --port 57700 \
  --use_bf16

Run a RoboDojo task from XPolicyLab/policy/starVLA:

STARVLA_CKPT_PATH="$CKPT" \
STARVLA_INCLUDE_STATE=True \
STARVLA_UNNORM_KEY=arx_x5 \
STARVLA_EXECUTE_HORIZON=16 \
bash eval.sh \
  RoboDojo build_tower qwenoft_steps_100000 \
  arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>

Use RoboDojo scripts/internal/summarize_result.py for official aggregation.

Intended use

This checkpoint is intended for RoboDojo simulation research with the ARX X5 dual-arm embodiment. Performance outside the saved observation, action, and normalization contract has not been established.

Downloads last month
29
Video Preview
loading

Model tree for StarVLA/Qwen3vl4b-OFT-RoboDojo

Finetuned
(376)
this model