StarVLA QwenOFT Qwen3-VL-4B for RoboDojo (100k)
This directory contains one StarVLA QwenOFT checkpoint initialized from
Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture.
The VLM, VLM interface, and MLP action head were trained end to end.
Model details
| Item | Value |
|---|---|
| Framework | StarVLA QwenOFT |
| Base VLM | Qwen3-VL-4B-Instruct |
| Action model | MLP, hidden size 2,560 |
| Objective | L1 action regression |
| Action representation | 14D absolute joint position (abs_qpos) |
| Action horizon | 50 |
| State dimension | 14 |
| Camera input | Head, left wrist, right wrist; resized to 224 x 224 |
| Checkpoint step | 100,000 |
| Checkpoint format | Complete StarVLA framework state dict (.pt) |
Files
README.md
config.yaml
config.full.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt
Only the requested 100k checkpoint is included. Keep the saved configuration and normalization statistics beside the checkpoint.
Training details
| Setting | Value |
|---|---|
| Dataset mixture | robodojo_v21_all_h50_q99 |
| Training tasks | 35 |
| Per-GPU batch size | 16 |
| Gradient accumulation | 1 |
| Frozen modules | None |
| Optimizer | AdamW, betas (0.9, 0.95), epsilon 1e-8 |
| VLM learning rate | 1e-5 |
| VLM-interface learning rate | 1e-5 |
| Action-model learning rate | 1e-4 |
| Schedule | Cosine, 5,000 warmup steps, minimum LR 5e-7 |
| Gradient checkpointing | Enabled |
| Random seed | 42 |
Official RoboDojo evaluation
All policies below use the official complete 42-task protocol: 50 episodes per
task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This
directory's policy is bolded. Higher is better for both SR and Score.
Group summary
| Policy | Average | Generalization | Precision | Long-Horizon | Memory | Open |
|---|---|---|---|---|---|---|
| QwenOFT | 4.86 / 8.01 | 4.33 / 6.42 | 11.75 / 17.54 | 5.50 / 12.95 | 1.67 / 1.77 | 0.50 / 0.60 |
| QwenGR00T | 3.81 / 7.35 | 3.50 / 6.52 | 5.75 / 10.09 | 6.50 / 15.46 | 3.33 / 4.37 | 0.00 / 0.00 |
| QwenPI_v3 | 6.19 / 9.60 | 4.17 / 7.28 | 14.00 / 19.06 | 10.00 / 17.84 | 2.00 / 2.32 | 0.75 / 0.88 |
Task details
Task values are SR (%) / Score; each task uses 50 episodes.
| Evaluation group / task | QwenOFT | QwenGR00T | QwenPI_v3 |
|---|---|---|---|
| Generalization | 4.33 / 6.42 | 3.50 / 6.52 | 4.17 / 7.28 |
| stack_bowls | 18.00 / 21.00 | 10.00 / 14.80 | 14.00 / 16.70 |
| push_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pack_objects_into_box | 0.00 / 3.10 | 0.00 / 7.80 | 0.00 / 6.80 |
| fold_clothes | 10.00 / 12.80 | 8.00 / 12.40 | 2.00 / 9.60 |
| hang_mugs | 0.00 / 3.60 | 0.00 / 3.00 | 0.00 / 3.50 |
| sweep_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| pour_liquid_into_cup | 14.00 / 14.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| make_toast | 0.00 / 1.00 | 2.00 / 5.00 | 2.00 / 5.00 |
| arrange_largest_number | 0.00 / 1.90 | 2.00 / 4.10 | 2.00 / 5.70 |
| sort_nesting_dolls_by_size | 0.00 / 0.00 | 4.00 / 4.00 | 6.00 / 6.00 |
| store_laptop_and_headphones | 4.00 / 11.20 | 0.00 / 7.20 | 2.00 / 8.40 |
| stack_blocks | 6.00 / 8.40 | 2.00 / 5.90 | 8.00 / 11.60 |
| Precision | 11.75 / 17.54 | 5.75 / 10.09 | 14.00 / 19.06 |
| fasten_screws | 4.00 / 8.00 | 0.00 / 2.00 | 0.00 / 6.00 |
| plug_in_charger | 6.00 / 6.00 | 2.00 / 2.00 | 4.00 / 4.00 |
| insert_tubes | 40.00 / 51.60 | 28.00 / 40.40 | 44.00 / 56.80 |
| pour_balls_into_vase | 8.00 / 8.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| play_Xylophone | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| deposit_coin | 0.00 / 3.20 | 2.00 / 5.60 | 6.00 / 7.60 |
| insert_key | 0.00 / 12.90 | 0.00 / 9.90 | 0.00 / 11.10 |
| build_tower | 36.00 / 50.60 | 14.00 / 20.80 | 56.00 / 65.00 |
| Long-Horizon | 5.50 / 12.95 | 6.50 / 15.46 | 10.00 / 17.84 |
| put_bottles_into_dustbin | 22.00 / 40.90 | 26.00 / 44.40 | 64.00 / 73.60 |
| fill_pen_holder | 4.00 / 11.70 | 6.00 / 14.40 | 8.00 / 23.00 |
| classify_objects | 2.00 / 5.50 | 6.00 / 11.50 | 0.00 / 7.50 |
| play_tic_tac_toe | 0.00 / 12.40 | 2.00 / 16.40 | 0.00 / 6.80 |
| fill_egg_holder | 0.00 / 0.60 | 0.00 / 0.00 | 0.00 / 0.80 |
| organize_table | 0.00 / 16.50 | 0.00 / 25.00 | 4.00 / 27.00 |
| make_kong | 16.00 / 16.00 | 12.00 / 12.00 | 4.00 / 4.00 |
| play_stacking_toy | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| Memory | 1.67 / 1.77 | 3.33 / 4.37 | 2.00 / 2.32 |
| cover_blocks | 0.00 / 0.60 | 6.00 / 12.10 | 0.00 / 1.50 |
| match_and_pick_from_conveyor | 10.00 / 10.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| swap_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| swap_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| press_by_number | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| imitate_sorting_sequence | 0.00 / 0.00 | 0.00 / 0.10 | 0.00 / 0.40 |
| Open | 0.50 / 0.60 | 0.00 / 0.00 | 0.75 / 0.88 |
| align_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| general_pickup | 4.00 / 4.00 | 0.00 / 0.00 | 6.00 / 6.00 |
| stack_blocks_by_language | 0.00 / 0.80 | 0.00 / 0.00 | 0.00 / 0.80 |
| solve_equation | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| classify_objects_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.20 |
| pick_from_conveyor_by_image | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| store_tools_in_toolbox | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pour_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
Evaluation
Start the StarVLA model server:
export CKPT=/path/to/MODEL_DIR/checkpoints/steps_100000_pytorch_model.pt
python deployment/model_server/server_policy.py \
--ckpt_path "$CKPT" \
--port 57700 \
--use_bf16
Run a RoboDojo task from XPolicyLab/policy/starVLA:
STARVLA_CKPT_PATH="$CKPT" \
STARVLA_INCLUDE_STATE=True \
STARVLA_UNNORM_KEY=arx_x5 \
STARVLA_EXECUTE_HORIZON=16 \
bash eval.sh \
RoboDojo build_tower qwenoft_steps_100000 \
arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>
Use RoboDojo scripts/internal/summarize_result.py for official aggregation.
Intended use
This checkpoint is intended for RoboDojo simulation research with the ARX X5 dual-arm embodiment. Performance outside the saved observation, action, and normalization contract has not been established.
- Downloads last month
- 29
Model tree for StarVLA/Qwen3vl4b-OFT-RoboDojo
Base model
Qwen/Qwen3-VL-4B-Instruct