SmolVLA - Clutter (joint, 2-cam)

HuggingFace LeRobot SmolVLA fine-tuned on the ManiGuard clutter base task (sim Franka Panda). Part of the ManiGuard VLA benchmark - SmolVLA vs pi0.5 vs GR00T on the same task families with identical data, cameras, and controller.

Model

  • Base: lerobot/smolvla_base - SmolVLM2 vision-language backbone + flow-matching action expert
  • Embodiment: Franka Panda, 8-D joint state/action (7 arm joints + 1 gripper), padded to SmolVLA's 32-D
  • Cameras (2): observation.images.top (overview = left view) + observation.images.wrist (256x256)
  • Action: absolute joint targets (fed straight to a JointController at eval); 50-step chunk
  • Tuning: SmolVLA default - vision encoder frozen, train the action expert (no LoRA)

Training

Usage

Load with SmolVLAPolicy.from_pretrained("IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-clutter-joint-2cam-yanZ") from LeRobot. The checkpoint carries the normalization stats.

WARNING - Convention (must match at eval): joint-space JointController (absolute joint targets) + 2 cameras (observation.images.top = the left overview, observation.images.wrist). A mismatched controller, camera set, or overview view silently feeds an out-of-distribution input.

Downloads last month
26
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-clutter-joint-2cam-yanZ

Finetuned
(6949)
this model