QiushiSun commited on
Commit
6f40205
·
verified ·
1 Parent(s): db11f26

Update 2026-08-07

Browse files
Files changed (1) hide show
  1. README.md +0 -9
README.md CHANGED
@@ -20,15 +20,6 @@ OS-Shepherd-9B is an open multimodal reward model for judging computer-use agent
20
 
21
  It is fine-tuned from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) on [OS-Shepherd-100K](https://huggingface.co/datasets/OS-Copilot/OS-Shepherd-100K) using SFT followed by GRPO, with the RL stage focused on reducing false-success judgments.
22
 
23
- ## Results
24
-
25
- | Benchmark | Accuracy | Fail recall |
26
- |---|---:|---:|
27
- | OSReward | 86.1 | 86.0 |
28
- | OSReward-Hard | 60.2 | 57.6 |
29
-
30
- Results use the fixed judging protocol described in the [OSReward paper](https://arxiv.org/abs/2607.28609).
31
-
32
  ## Usage
33
 
34
  Use the canonical prompt and trajectory format from the [OSReward repository](https://github.com/OS-Copilot/OSReward). A recent Transformers, vLLM, or SGLang version with Qwen3.5 multimodal support is required.
 
20
 
21
  It is fine-tuned from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) on [OS-Shepherd-100K](https://huggingface.co/datasets/OS-Copilot/OS-Shepherd-100K) using SFT followed by GRPO, with the RL stage focused on reducing false-success judgments.
22
 
 
 
 
 
 
 
 
 
 
23
  ## Usage
24
 
25
  Use the canonical prompt and trajectory format from the [OSReward repository](https://github.com/OS-Copilot/OSReward). A recent Transformers, vLLM, or SGLang version with Qwen3.5 multimodal support is required.