SafeVLA-Bench model card

Code, exact commands, tests, and result tables: https://github.com/Jatshi/SafeVLA-Bench.

Released checkpoints

  • state_bc_h16_full1000_holdout_seed{11,23,37,53,71}.pt: strict deterministic state behavior-cloning checkpoints with 16-step chunks.
  • rgbd_bc_h8_seed11_300ep.pt and rgbd_bc_h16_seed11_300ep.pt: compact RGB-D behavior-cloning negative baselines.
  • diffusion_h16_seed11_500ep.pt: conditional DDPM baseline retained with its negative closed-loop result.

Intended use

Research and education on selective execution, clarification, and stop decisions in ManiSkill simulation. These checkpoints are not suitable for direct physical-robot deployment.

Training data

Official ManiSkill PickCube-v1 motion-planning demonstrations replayed under pd_ee_delta_pos. Dataset hashes and conversion scripts are included in the repository and dataset card.

Evaluation

Across independent training initializations 11, 23, 37, 53, and 71, state BC reached 88% mean success (95% bootstrap CI 72%–100%, standard deviation 17.9 pp) on physx_cpu and 53% (39%–60%, standard deviation 15.7 pp) on physx_cuda. Training removed every episode whose metadata seed matched the evaluation set. The RGB-D and diffusion checkpoints did not achieve closed-loop success; they are published to make the failure analysis reproducible, not as high-performing models.

Limitations

The state checkpoint consumes privileged simulator state. GPU replay ignores some initial-state options, causing a measurable backend gap. Safety results are specific to the simulated geometry and thresholds documented in the repository.

The repository source code is Apache-2.0. Released weights are conservatively marked CC BY-NC 4.0 because training and evaluation use ManiSkill assets carrying that license; ManiSkill's third-party notices remain applicable.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading