RLinf-lingbotvla-click-bell-grpo

This is a LingbotVLA checkpoint post-trained with RLinf GRPO on the RoboTwin click_bell task.

The checkpoint has been merged into the pinned RoboTwin SFT base format, so it can be loaded directly as a LingbotVLA model directory. It does not require a separate runner.ckpt_path file for evaluation.

Base Model

Result Summary

In our RoboTwin random evaluation, this RLinf-GRPO checkpoint performs better than the pinned RoboTwin SFT checkpoint on the corresponding click_bell task.

Download

hf download RLinf/RLinf-lingbotvla-click-bell-grpo \
    --local-dir RLinf-lingbotvla-click-bell-grpo

Evaluation in RLinf

Use the downloaded directory as both actor.model.model_path and rollout.model.model_path. Keep tokenizer_path pointing to a local Qwen2.5-VL-3B-Instruct directory.

python examples/embodiment/eval_embodied_agent.py \
    --config-path examples/embodiment/config \
    --config-name robotwin_click_bell_grpo_lingbotvla \
    runner.only_eval=True \
    actor.model.model_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
    actor.model.tokenizer_path=/path/to/Qwen2.5-VL-3B-Instruct \
    actor.model.lingbotvla.config_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
    rollout.model.model_path=/path/to/RLinf-lingbotvla-click-bell-grpo \
    rollout.model.tokenizer_path=/path/to/Qwen2.5-VL-3B-Instruct \
    rollout.model.lingbotvla.config_path=/path/to/RLinf-lingbotvla-click-bell-grpo

For large RoboTwin evaluations, set env.eval.total_num_envs, env.eval.max_episode_steps, and env.eval.max_steps_per_rollout_epoch according to the target protocol.

Notes

This repository stores the merged LingbotVLA model weights only. RLinf wrapper-only tensors, such as value heads used during training, are not included.

Downloads last month
7
Safetensors
Model size
4B params
Tensor type
BF16
·
Video Preview
loading