Instructions to use Twu31/smolvla-so101-blue-napkin-160k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Twu31/smolvla-so101-blue-napkin-160k with LeRobot:
# See https://github.com/huggingface/lerobot?tab=readme-ov-file#installation for more details git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e .[smolvla]
# Launch finetuning on your dataset python lerobot/scripts/train.py \ --policy.path=Twu31/smolvla-so101-blue-napkin-160k \ --dataset.repo_id=lerobot/svla_so101_pickplace \ --batch_size=64 \ --steps=20000 \ --output_dir=outputs/train/my_smolvla \ --job_name=my_smolvla_training \ --policy.device=cuda \ --wandb.enable=true
# Run the policy using the record function python -m lerobot.record \ --robot.type=so101_follower \ --robot.port=/dev/ttyACM0 \ # <- Use your port --robot.id=my_blue_follower_arm \ # <- Use your robot id --robot.cameras="{ front: {type: opencv, index_or_path: 8, width: 640, height: 480, fps: 30}}" \ # <- Use your cameras --dataset.single_task="Grasp a lego block and put it in the bin." \ # <- Use the same task description you used in your dataset recording --dataset.repo_id=HF_USER/dataset_name \ # <- This will be the dataset name on HF Hub --dataset.episode_time_s=50 \ --dataset.num_episodes=10 \ --policy.path=Twu31/smolvla-so101-blue-napkin-160k - Notebooks
- Google Colab
- Kaggle
SmolVLA — SO-ARM101 "Hand me the blue napkin" (160k steps)
A SmolVLA policy fine-tuned on 101 teleoperated demonstrations of a single real-world handover task on a ~$300 SO-ARM101 arm: the robot picks up a pack of blue napkins from the table and hands it to a person.
This is the production checkpoint — it reaches 80%+ real-arm success.
| Task | Hand me the blue napkin |
| Robot | SO-ARM101 (so101_follower), 6-DOF, STS3215 servos |
| Training data | Twu31/so101_hand_blue_napkin — 101 episodes, 40,493 frames, 30 fps |
| Training steps | 160,000 |
| Real-arm success | 80%+ |
| Sibling checkpoint | 50k-step version |
Inputs / outputs
observation.images.front |
480×640 RGB, fixed front-facing camera |
observation.images.handeye |
480×640 RGB, wrist/hand-eye camera |
observation.state |
6-dim joint positions — shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper |
action |
6-dim absolute joint targets, predicted in chunks of 50 |
Both cameras matter: the front view locates the napkin pack and the person's hand, the hand-eye view drives the grasp itself.
Usage
from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy
policy = SmolVLAPolicy.from_pretrained("Twu31/smolvla-so101-blue-napkin-160k")
# observation dict: observation.images.front, observation.images.handeye,
# observation.state, plus the task string
action = policy.select_action(observation)
Or point lerobot-record / your eval script at this repo id as the policy path.
Training
Fine-tuned with LeRobot's SmolVLA recipe on a single
CUDA GPU. The VLM backbone is HuggingFaceTB/SmolVLM2-500M-Video-Instruct; the vision encoder is
frozen and only the action expert (plus the state projection) is trained.
| Hyperparameter | Value |
|---|---|
| Steps | 160,000 |
| Batch size | 1 × 8 gradient accumulation |
| Optimizer | AdamW, lr 1e-4, betas (0.9, 0.95), wd 1e-10, grad-clip 10.0 |
| Schedule | cosine decay, 1,000 warm-up steps → 1e-6 |
| Action chunk | 50 (n_action_steps = 50) |
| Image preprocessing | resize with padding to 512×512 |
| Normalization | state/action MEAN_STD, visual IDENTITY |
freeze_vision_encoder |
true |
train_expert_only |
true |
| Seed | 1000 |
Evaluation
Real-arm closed-loop evaluation only — there is no simulator for this task, and (see below) offline metrics turned out to be actively misleading for it.
Recorded eval rollouts from this checkpoint are published at
Twu31/so101_hand_blue_napkin_eval_rollouts.
Known limitations
- Single task, single scene. One napkin pack, one table, one lab lighting setup. Expect it to fail on a different napkin, a cluttered table, or very different lighting.
- Fixed camera geometry. The front camera pose is baked into the 101 demos; moving it will degrade performance.
- Handover timing is learned from human teleoperation, so the arm expects a hand to appear in roughly the demonstrated region.
- Trained from 101 demos of a single operator — it inherits that operator's motion style.
A negative result worth reading
We tried to improve this policy with offline RL on its own 101 demos — five approaches (IQL with several reward-labeling schemes, vision-IQL, residual IQL on top of SmolVLA). The best one improved offline action MSE by 51% over this checkpoint and then scored 0% on the real arm, because small residuals compound into closed-loop distribution shift.
Write-up and code: twu3202/soarm101-offline-rl-experiments. Short version: for closed-loop control, offline MSE does not predict real-arm success.
Citation / credits
- SmolVLA and LeRobot by Hugging Face
- VLM backbone: SmolVLM2-500M-Video-Instruct
- Robot: SO-ARM101 / SO-100 family
- Downloads last month
- 15
Model tree for Twu31/smolvla-so101-blue-napkin-160k
Base model
HuggingFaceTB/SmolLM2-360M