---
license: apache-2.0
base_model: Qwen/Qwen3-VL-8B-Instruct
library_name: transformers
pipeline_tag: video-text-to-text
language:
- en
tags:
- arxiv:2609.25773
- video
- multi-hop-reasoning
- reinforcement-learning
- rlvr
- grpo
- qwen3-vl
datasets:
- ngqtrung/Video-HopChain
---
# Video-HopChain-8B
Qwen3-VL-8B-Instruct trained with GRPO in two stages, first on a general video question answering
mixture and then on [Video-HopChain](https://huggingface.co/datasets/ngqtrung/Video-HopChain) with
Confidence-Gated Exploration (CGE). This is **Video-HopChain-8B**, the "+ standard RL + Video-HopChain + CGE" row of Table 1 in the
paper, at `global_step_10` of the second stage.
- **Paper:** *Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models* ([arXiv:2609.25773](https://arxiv.org/abs/2609.25773))
- **Dataset:** [ngqtrung/Video-HopChain](https://huggingface.co/datasets/ngqtrung/Video-HopChain)
- **Project page:** [ngquangtrung57.github.io/video-hopchain-page](https://ngquangtrung57.github.io/video-hopchain-page/)
- **Code:** [github.com/ngquangtrung57/video-hopchain](https://github.com/ngquangtrung57/video-hopchain)
- **Collection:** [Video-HopChain](https://huggingface.co/collections/ngqtrung/video-hopchain)
- **Initialization:** [ngqtrung/Video-HopChain-8B-Standard-RL](https://huggingface.co/ngqtrung/Video-HopChain-8B-Standard-RL), the standard-RL stage of Table 1
## Results
Accuracy in percent, evaluated with `lmms-eval` at 100 frames per video under one setting for every
row. "In-domain" is the 1,000-question held-out Video-HopChain split.
| Model | Video-MME | PerceptionComp | Video-MMMU | Video-Holmes | VCRBench | MMR-V | LongVideo-Reason | VRBench | mean | in-domain |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B-Instruct | 64.0 | 28.1 | 63.5 | 40.7 | 32.0 | 42.7 | 72.6 | 74.7 | 52.3 | 13.4 |
| + standard RL | 65.8 | 34.3 | 63.0 | 47.4 | 35.9 | 43.4 | 76.1 | 77.6 | 55.4 | 13.4 |
| + standard RL + CGE | 67.9 | 34.4 | 62.5 | 48.5 | 35.1 | 46.4 | 78.8 | 79.5 | 56.6 | 16.4 |
| + standard RL + Video-HopChain | 68.6 | 36.2 | 64.7 | **48.6** | 42.1 | 45.9 | 77.8 | 79.4 | 57.9 | 21.2 |
| **this model** | **69.2** | **37.4** | **67.5** | **48.6** | **44.9** | **46.6** | **78.9** | **81.0** | **59.3** | **23.2** |
VCRBench is reported on its multiple-choice subset.
## Usage
```python
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
model = Qwen3VLForConditionalGeneration.from_pretrained(
"ngqtrung/Video-HopChain-8B", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ngqtrung/Video-HopChain-8B")
```
The weights are BF16. The model was trained to reason inside `...` and to put the final
answer in `\boxed{...}`, so use the same system prompt as training. It is in the
dataset rows and in the repository.
## Training
Confidence-Gated Exploration recovers zero-variance groups in GRPO (also called advantage collapse). With 8 rollouts per question, CGE draws
the first 4 normally. If those 4 are all correct or all incorrect, the group carries no reward
variance and no gradient, so CGE draws the last 4 with the policy's top token masked wherever its
probability exceeds `tau = 0.95`, inside the reasoning span only. The masked positions are dropped
from the loss while all 8 rollouts enter the group advantage, so the method spends no extra rollouts.
| | |
|---|---|
| Base model | Qwen3-VL-8B-Instruct |
| Stage 1 | GRPO on a 105,993-row general video QA mixture, 24 frames |
| Stage 2 | GRPO + CGE on Video-HopChain, 22,550 rows, 140 frames |
| Rollouts per question | 8: the first 4 as usual, the last 4 under the mask |
| Mask threshold | 0.95, reasoning span only |
| Reward | exact match on the integer sum, plus a format term |
| Hardware | 4 nodes of 8 H100 80GB, asynchronous verl trainer |
Appendix C of the paper gives every hyperparameter.
## Limitations
Evaluated at one model scale only. The dataset is synthetic and verified against captions rather than
against the frames, so a caption error can reach a label. The paper reports both limitations, together
with the exploration settings that we did not ablate.
## Citation
```bibtex
@misc{nguyenquang2026videohopchain,
title = {Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models},
author = {Nguyen Quang, Trung and Dong, Yuhao and Sun, Shuo and Liu, Shuai and Tian, Shulin and Yap, Kim-Hui and Liu, Ziwei},
year = {2026},
eprint = {2609.25773},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.25773}
}
```