SmolVLA — SO-ARM101 "Hand me the blue napkin" (160k steps)

A SmolVLA policy fine-tuned on 101 teleoperated demonstrations of a single real-world handover task on a ~$300 SO-ARM101 arm: the robot picks up a pack of blue napkins from the table and hands it to a person.

This is the production checkpoint — it reaches 80%+ real-arm success.

Task Hand me the blue napkin
Robot SO-ARM101 (so101_follower), 6-DOF, STS3215 servos
Training data Twu31/so101_hand_blue_napkin — 101 episodes, 40,493 frames, 30 fps
Training steps 160,000
Real-arm success 80%+
Sibling checkpoint 50k-step version

Inputs / outputs

observation.images.front 480×640 RGB, fixed front-facing camera
observation.images.handeye 480×640 RGB, wrist/hand-eye camera
observation.state 6-dim joint positions — shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper
action 6-dim absolute joint targets, predicted in chunks of 50

Both cameras matter: the front view locates the napkin pack and the person's hand, the hand-eye view drives the grasp itself.

Usage

from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy

policy = SmolVLAPolicy.from_pretrained("Twu31/smolvla-so101-blue-napkin-160k")
# observation dict: observation.images.front, observation.images.handeye,
#                   observation.state, plus the task string
action = policy.select_action(observation)

Or point lerobot-record / your eval script at this repo id as the policy path.

Training

Fine-tuned with LeRobot's SmolVLA recipe on a single CUDA GPU. The VLM backbone is HuggingFaceTB/SmolVLM2-500M-Video-Instruct; the vision encoder is frozen and only the action expert (plus the state projection) is trained.

Hyperparameter Value
Steps 160,000
Batch size 1 × 8 gradient accumulation
Optimizer AdamW, lr 1e-4, betas (0.9, 0.95), wd 1e-10, grad-clip 10.0
Schedule cosine decay, 1,000 warm-up steps → 1e-6
Action chunk 50 (n_action_steps = 50)
Image preprocessing resize with padding to 512×512
Normalization state/action MEAN_STD, visual IDENTITY
freeze_vision_encoder true
train_expert_only true
Seed 1000

Evaluation

Real-arm closed-loop evaluation only — there is no simulator for this task, and (see below) offline metrics turned out to be actively misleading for it.

Recorded eval rollouts from this checkpoint are published at Twu31/so101_hand_blue_napkin_eval_rollouts.

Known limitations

  • Single task, single scene. One napkin pack, one table, one lab lighting setup. Expect it to fail on a different napkin, a cluttered table, or very different lighting.
  • Fixed camera geometry. The front camera pose is baked into the 101 demos; moving it will degrade performance.
  • Handover timing is learned from human teleoperation, so the arm expects a hand to appear in roughly the demonstrated region.
  • Trained from 101 demos of a single operator — it inherits that operator's motion style.

A negative result worth reading

We tried to improve this policy with offline RL on its own 101 demos — five approaches (IQL with several reward-labeling schemes, vision-IQL, residual IQL on top of SmolVLA). The best one improved offline action MSE by 51% over this checkpoint and then scored 0% on the real arm, because small residuals compound into closed-loop distribution shift.

Write-up and code: twu3202/soarm101-offline-rl-experiments. Short version: for closed-loop control, offline MSE does not predict real-arm success.

Citation / credits

Downloads last month
15
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for Twu31/smolvla-so101-blue-napkin-160k

Dataset used to train Twu31/smolvla-so101-blue-napkin-160k