SmolVLA — SO-ARM101 "Hand me the blue napkin" (50k steps)

An earlier, shorter-trained checkpoint of a SmolVLA policy for a real-world handover task on a ~$300 SO-ARM101 arm: pick up a pack of blue napkins from the table and hand it to a person.

For actual use, prefer the 160k-step checkpoint: Twu31/smolvla-so101-blue-napkin-160k — that one is the production policy (80%+ real-arm success). This 50k checkpoint is published for training-length comparison and reproducibility.

Task Hand me the blue napkin
Robot SO-ARM101 (so101_follower), 6-DOF, STS3215 servos
Training data Twu31/so101_hand_blue_napkin — 101 episodes, 40,493 frames, 30 fps
Training steps 50,000 (cosine decay over 20,000)
Recommended checkpoint 160k version

Inputs / outputs

observation.images.front 480×640 RGB, fixed front-facing camera
observation.images.handeye 480×640 RGB, wrist/hand-eye camera
observation.state 6-dim joint positions — shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper
action 6-dim absolute joint targets, chunks of 50

Usage

from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy

policy = SmolVLAPolicy.from_pretrained("Twu31/smolvla-so101-blue-napkin-50k")
action = policy.select_action(observation)

Training

LeRobot SmolVLA recipe, single CUDA GPU. VLM backbone HuggingFaceTB/SmolVLM2-500M-Video-Instruct, vision encoder frozen, action expert + state projection trained.

Hyperparameter Value
Steps 50,000
Batch size 1 × 8 gradient accumulation
Optimizer AdamW, lr 1e-4, betas (0.9, 0.95), wd 1e-10, grad-clip 10.0
Schedule cosine decay, 1,000 warm-up steps, 20,000 decay steps → 1e-6
Action chunk 50
Image preprocessing resize with padding to 512×512
Normalization state/action MEAN_STD, visual IDENTITY
Seed 1000

Identical to the 160k run except for steps and the decay horizon.

Evaluation

Real-arm closed-loop only. A small set of recorded rollouts from this checkpoint is included in Twu31/so101_hand_blue_napkin_eval_rollouts (the eval_napkin_50k_v1_* subfolder). It was superseded by the 160k checkpoint before a large formal eval was run, so no headline success rate is claimed for this checkpoint — the 80%+ figure belongs to the 160k model, not this one.

Limitations

Same as the 160k checkpoint — single task, single scene, fixed camera geometry, one operator's motion style — plus a shorter training run. See the 160k card for the full discussion and for the offline-RL negative result that came out of this project.

Credits

Downloads last month
17
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for Twu31/smolvla-so101-blue-napkin-50k

Dataset used to train Twu31/smolvla-so101-blue-napkin-50k