--- license: apache-2.0 base_model: Qwen/Qwen3-VL-8B-Instruct library_name: transformers pipeline_tag: video-text-to-text language: - en tags: - arxiv:2609.25773 - video - multi-hop-reasoning - reinforcement-learning - rlvr - grpo - qwen3-vl datasets: - ngqtrung/Video-HopChain --- # Video-HopChain-8B Qwen3-VL-8B-Instruct trained with GRPO in two stages, first on a general video question answering mixture and then on [Video-HopChain](https://huggingface.co/datasets/ngqtrung/Video-HopChain) with Confidence-Gated Exploration (CGE). This is **Video-HopChain-8B**, the "+ standard RL + Video-HopChain + CGE" row of Table 1 in the paper, at `global_step_10` of the second stage. - **Paper:** *Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models* ([arXiv:2609.25773](https://arxiv.org/abs/2609.25773)) - **Dataset:** [ngqtrung/Video-HopChain](https://huggingface.co/datasets/ngqtrung/Video-HopChain) - **Project page:** [ngquangtrung57.github.io/video-hopchain-page](https://ngquangtrung57.github.io/video-hopchain-page/) - **Code:** [github.com/ngquangtrung57/video-hopchain](https://github.com/ngquangtrung57/video-hopchain) - **Collection:** [Video-HopChain](https://huggingface.co/collections/ngqtrung/video-hopchain) - **Initialization:** [ngqtrung/Video-HopChain-8B-Standard-RL](https://huggingface.co/ngqtrung/Video-HopChain-8B-Standard-RL), the standard-RL stage of Table 1 ## Results Accuracy in percent, evaluated with `lmms-eval` at 100 frames per video under one setting for every row. "In-domain" is the 1,000-question held-out Video-HopChain split. | Model | Video-MME | PerceptionComp | Video-MMMU | Video-Holmes | VCRBench | MMR-V | LongVideo-Reason | VRBench | mean | in-domain | |---|---|---|---|---|---|---|---|---|---|---| | Qwen3-VL-8B-Instruct | 64.0 | 28.1 | 63.5 | 40.7 | 32.0 | 42.7 | 72.6 | 74.7 | 52.3 | 13.4 | | + standard RL | 65.8 | 34.3 | 63.0 | 47.4 | 35.9 | 43.4 | 76.1 | 77.6 | 55.4 | 13.4 | | + standard RL + CGE | 67.9 | 34.4 | 62.5 | 48.5 | 35.1 | 46.4 | 78.8 | 79.5 | 56.6 | 16.4 | | + standard RL + Video-HopChain | 68.6 | 36.2 | 64.7 | **48.6** | 42.1 | 45.9 | 77.8 | 79.4 | 57.9 | 21.2 | | **this model** | **69.2** | **37.4** | **67.5** | **48.6** | **44.9** | **46.6** | **78.9** | **81.0** | **59.3** | **23.2** | VCRBench is reported on its multiple-choice subset. ## Usage ```python from transformers import AutoProcessor, Qwen3VLForConditionalGeneration model = Qwen3VLForConditionalGeneration.from_pretrained( "ngqtrung/Video-HopChain-8B", dtype="auto", device_map="auto") processor = AutoProcessor.from_pretrained("ngqtrung/Video-HopChain-8B") ``` The weights are BF16. The model was trained to reason inside `...` and to put the final answer in `\boxed{...}`, so use the same system prompt as training. It is in the dataset rows and in the repository. ## Training Confidence-Gated Exploration recovers zero-variance groups in GRPO (also called advantage collapse). With 8 rollouts per question, CGE draws the first 4 normally. If those 4 are all correct or all incorrect, the group carries no reward variance and no gradient, so CGE draws the last 4 with the policy's top token masked wherever its probability exceeds `tau = 0.95`, inside the reasoning span only. The masked positions are dropped from the loss while all 8 rollouts enter the group advantage, so the method spends no extra rollouts. | | | |---|---| | Base model | Qwen3-VL-8B-Instruct | | Stage 1 | GRPO on a 105,993-row general video QA mixture, 24 frames | | Stage 2 | GRPO + CGE on Video-HopChain, 22,550 rows, 140 frames | | Rollouts per question | 8: the first 4 as usual, the last 4 under the mask | | Mask threshold | 0.95, reasoning span only | | Reward | exact match on the integer sum, plus a format term | | Hardware | 4 nodes of 8 H100 80GB, asynchronous verl trainer | Appendix C of the paper gives every hyperparameter. ## Limitations Evaluated at one model scale only. The dataset is synthetic and verified against captions rather than against the frames, so a caption error can reach a label. The paper reports both limitations, together with the exploration settings that we did not ablate. ## Citation ```bibtex @misc{nguyenquang2026videohopchain, title = {Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models}, author = {Nguyen Quang, Trung and Dong, Yuhao and Sun, Shuo and Liu, Shuai and Tian, Shulin and Yap, Kim-Hui and Liu, Ziwei}, year = {2026}, eprint = {2609.25773}, archivePrefix = {arXiv}, primaryClass = {cs.CV}, url = {https://arxiv.org/abs/2609.25773} } ```