File size: 1,747 Bytes
db11f26
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
---
library_name: transformers
pipeline_tag: image-text-to-text
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
datasets:
- OS-Copilot/OS-Shepherd-100K
language:
- en
tags:
- computer-use
- reward-model
- trajectory-evaluation
- multimodal
---

# OS-Shepherd-9B

OS-Shepherd-9B is an open multimodal reward model for judging computer-use agent trajectories. Given a task instruction, screenshots, and the agent's reasoning and actions, it determines whether the task was completed and returns a reasoned `SUCCESS` or `FAIL` verdict.

It is fine-tuned from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) on [OS-Shepherd-100K](https://huggingface.co/datasets/OS-Copilot/OS-Shepherd-100K) using SFT followed by GRPO, with the RL stage focused on reducing false-success judgments.

## Results

| Benchmark | Accuracy | Fail recall |
|---|---:|---:|
| OSReward | 86.1 | 86.0 |
| OSReward-Hard | 60.2 | 57.6 |

Results use the fixed judging protocol described in the [OSReward paper](https://arxiv.org/abs/2607.28609).

## Usage

Use the canonical prompt and trajectory format from the [OSReward repository](https://github.com/OS-Copilot/OSReward). A recent Transformers, vLLM, or SGLang version with Qwen3.5 multimodal support is required.

This model is intended for trajectory evaluation, data filtering, and reward-model research. It is not a computer-control policy and may still miss fine-grained visual failures, especially on hard cases.

## License

Apache License 2.0. See `LICENSE`.

## Citation

```bibtex
@article{sun2026osreward,
  title={OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models},
  author={Sun, Qiushi and others},
  journal={arXiv preprint arXiv:2607.28609},
  year={2026}
}
```