RLinf-lingbotvla-click-bell-grpo
This is a LingbotVLA checkpoint post-trained with RLinf GRPO on the RoboTwin click_bell task.
The checkpoint has been merged into the pinned RoboTwin SFT base format, so it can be loaded directly as a LingbotVLA model directory. It does not require a separate runner.ckpt_path file for evaluation.
Base Model
- Upstream base:
robbyant/lingbot-vla-4b-posttrain-robotwin - Base revision:
3e0c7c476bde3daaac00f79f3741a292a299f60a - RL method: RLinf GRPO
- RoboTwin task:
click_bell
Result Summary
In our RoboTwin random evaluation, this RLinf-GRPO checkpoint performs better than the pinned RoboTwin SFT checkpoint on the corresponding click_bell task.
Download
hf download RLinf/RLinf-lingbotvla-click-bell-grpo \
--local-dir RLinf-lingbotvla-click-bell-grpo
Evaluation in RLinf
Use the downloaded directory as both actor.model.model_path and rollout.model.model_path. Keep tokenizer_path pointing to a local Qwen2.5-VL-3B-Instruct directory.
python examples/embodiment/eval_embodied_agent.py \
--config-path examples/embodiment/config \
--config-name robotwin_click_bell_grpo_lingbotvla \
runner.only_eval=True \
actor.model.model_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
actor.model.tokenizer_path=/path/to/Qwen2.5-VL-3B-Instruct \
actor.model.lingbotvla.config_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
rollout.model.model_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
rollout.model.tokenizer_path=/path/to/Qwen2.5-VL-3B-Instruct \
rollout.model.lingbotvla.config_path=/path/to/RLinf-lingbotvla-click-bell-grpo
For large RoboTwin evaluations, set env.eval.total_num_envs, env.eval.max_episode_steps, and env.eval.max_steps_per_rollout_epoch according to the target protocol.
Notes
This repository stores the merged LingbotVLA model weights only. RLinf wrapper-only tensors, such as value heads used during training, are not included.
- Downloads last month
- 7