Teburile commited on
Commit
ea9aaed
·
1 Parent(s): 9ab7759

Docs: refresh model card and usage

Browse files
Files changed (2) hide show
  1. README.md +95 -65
  2. USAGE.md +15 -6
README.md CHANGED
@@ -1,109 +1,139 @@
1
  ---
2
  license: apache-2.0
3
  library_name: pytorch
 
 
 
 
4
  tags:
5
- - reward-model
6
- - diffusion
7
- - rewarddit
8
- - preference-modeling
9
- - text-generation-evaluation
10
- base_model:
11
- - sfairXC/FsfairX-LLaMA3-RM-v0.1
12
  ---
13
 
14
- # DRM Checkpoints
15
 
16
- This repository hosts two RewardDiT checkpoints released with our paper.
17
 
18
- Paper: TODO: paste arXiv or project page link here.
19
 
20
- Code: TODO: paste GitHub repository link here.
 
 
 
 
 
 
 
 
 
 
21
 
22
  ## Checkpoints
23
 
24
- | Name | File | Reward Dim | Training Data | Intended Use |
25
- | --- | --- | ---: | --- | --- |
26
- | DRM-Multi-8B | `DRM-Multi-8B/model.pth` | 19 | `RLHFlow/ArmoRM-Multi-Objective-Data-v0.1` | Multi-objective reward generation and scalar scoring by mean aggregation |
27
- | DRM-Pref-8B | `DRM-Pref-8B/model.pth` | 1 | `allenai/llama-3.1-tulu-3-8b-preference-mixture` | Pair-preference reward scoring |
28
 
29
- Both checkpoints use FsfairX-LLaMA3-RM-v0.1 embeddings as the text condition.
30
 
31
- ## Recommended Inference Settings
 
 
 
 
 
 
 
32
 
33
- Use the following settings for both checkpoints:
34
 
35
- ```text
36
- mask_split = false
37
- num_steps = 10
38
- guidance_scale = 7.0
39
- num_samples = 32
40
- gate = off
41
- debias = off
42
- ```
43
 
44
- The public `ScoreGenerator` does not include gate or debias modules. Scores are computed by sampling RewardDiT rewards, averaging over samples, and then averaging over reward dimensions.
45
 
46
- ## Usage
 
 
 
 
 
 
 
47
 
48
- After downloading the code repository and this model repository, run:
49
 
50
  ```bash
51
- python src_upload/score_generator.py \
52
- --ckpt path/to/DRM-Multi-8B/model.pth \
53
  --prompt "User prompt" \
54
  --response "Assistant response"
55
  ```
56
 
57
- For the preference checkpoint:
58
 
59
  ```bash
60
- python src_upload/score_generator.py \
61
- --ckpt path/to/DRM-Pref-8B/model.pth \
62
  --prompt "User prompt" \
63
  --response "Assistant response"
64
  ```
65
 
66
- The default command-line inference settings in `score_generator.py` are already set to `num_steps=10`, `guidance_scale=7.0`, and `num_samples=32`.
67
 
68
- ## Model Details
69
 
70
  ### DRM-Multi-8B
71
 
72
- - Reward dimension: 19
73
- - Text embedding dimension: 4096
74
- - Hidden size: 384
75
- - Depth: 3
76
- - Attention heads: 6
77
- - Dropout: 0.2
78
- - Beta schedule: `squaredcos_cap_v2`
79
- - Prediction type: `epsilon`
80
 
81
  ### DRM-Pref-8B
82
 
83
- - Reward dimension: 1
84
- - Text embedding dimension: 4096
85
- - Hidden size: 384
86
- - Depth: 3
87
- - Attention heads: 6
88
- - Dropout: 0.2
89
- - Beta schedule: `squaredcos_cap_v2`
90
- - Prediction type: `epsilon`
91
- - Pair loss: denoising loss + Bradley-Terry loss + reward L2 regularization
92
- - `reward_reg_weight`: 0.001
93
 
94
- ## Citation
95
 
96
- TODO: paste BibTeX citation here after the paper is available.
97
 
98
- ```bibtex
99
- @article{TODO,
100
- title = {TODO},
101
- author = {TODO},
102
- journal = {arXiv preprint},
103
- year = {TODO}
104
- }
105
- ```
 
 
 
 
 
106
 
107
  ## Limitations
108
 
109
- These checkpoints are research artifacts intended for reward-modeling experiments. They inherit the limitations of their training data and text encoder, and should not be treated as calibrated absolute measures of response quality.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
  library_name: pytorch
4
+ pipeline_tag: text-classification
5
+ base_model: sfairXC/FsfairX-LLaMA3-RM-v0.1
6
+ language:
7
+ - en
8
  tags:
9
+ - reward-model
10
+ - diffusion
11
+ - rlhf
12
+ - alignment
13
+ - preference-modeling
 
 
14
  ---
15
 
16
+ # Diffusion Reward Models
17
 
18
+ This repository provides the released **DRM-Multi-8B** and **DRM-Pref-8B** RewardDiT checkpoints.
19
 
20
+ ## Links
21
 
22
+ - 📜 Paper — coming soon
23
+ - 💻 [Code](https://github.com/thunlp/DRM)
24
+ - 🤗 [Base encoder: FsfairX-LLaMA3-RM-v0.1](https://huggingface.co/sfairXC/FsfairX-LLaMA3-RM-v0.1)
25
+
26
+ ## Introduction
27
+
28
+ ![DRM overview](https://raw.githubusercontent.com/thunlp/DRM/main/figures/fig01_drm-overview.png)
29
+
30
+ **DRM** (Diffusion Reward Model) models the conditional reward distribution `p(r | x, y)` instead of reducing every prompt–response pair to a single point estimate or a fixed parametric family. A frozen LLM encoder conditions a lightweight Diffusion Transformer (RewardDiT), which denoises Gaussian noise into reward vectors. At inference time, multiple samples form an empirical distribution that can provide a scalar score, uncertainty estimate, or risk-sensitive statistic. **DRM diffuses reward vectors, not text.**
31
+
32
+ Both checkpoints use the frozen 7.5B-parameter [FsfairX-LLaMA3-RM-v0.1](https://huggingface.co/sfairXC/FsfairX-LLaMA3-RM-v0.1) encoder and train only an approximately 12M-parameter RewardDiT head.
33
 
34
  ## Checkpoints
35
 
36
+ | Model | File | Reward dimensions | Supervision | Training data |
37
+ |---|---|---:|---|---|
38
+ | **DRM-Multi-8B** | [`DRM-Multi-8B/model.pth`](DRM-Multi-8B/model.pth) | 19 | Masked multi-attribute denoising | `RLHFlow/ArmoRM-Multi-Objective-Data-v0.1` |
39
+ | **DRM-Pref-8B** | [`DRM-Pref-8B/model.pth`](DRM-Pref-8B/model.pth) | 1 | Denoising + Bradley–Terry preference loss | `allenai/llama-3.1-tulu-3-8b-preference-mixture` |
40
 
41
+ ### Shared architecture and inference defaults
42
 
43
+ - Text embedding dimension: 4096
44
+ - RewardDiT hidden size: 384
45
+ - Depth: 3 blocks
46
+ - Attention heads: 6
47
+ - Dropout: 0.2
48
+ - Diffusion schedule: `squaredcos_cap_v2`, 1000 training steps
49
+ - Prediction type: `epsilon`
50
+ - Inference: 10 DDIM steps, guidance scale 7, 32 reward samples
51
 
52
+ The released scorer does not use a gate model or reward-debiasing transform. It averages over sampled rewards and then over reward dimensions to produce one scalar per input.
53
 
54
+ ## How to Use
 
 
 
 
 
 
 
55
 
56
+ DRM uses a custom reward head and should be loaded through the released repository rather than `AutoModelForCausalLM`.
57
 
58
+ ```bash
59
+ git clone https://github.com/thunlp/DRM.git
60
+ cd DRM
61
+ pip install -r requirements.txt
62
+
63
+ hf download Teburile/DRM DRM-Multi-8B/model.pth --local-dir checkpoints
64
+ hf download Teburile/DRM DRM-Pref-8B/model.pth --local-dir checkpoints
65
+ ```
66
 
67
+ Score a prompt–response pair with DRM-Multi-8B:
68
 
69
  ```bash
70
+ python score_generator.py \
71
+ --ckpt checkpoints/DRM-Multi-8B/model.pth \
72
  --prompt "User prompt" \
73
  --response "Assistant response"
74
  ```
75
 
76
+ Or use the pairwise-preference checkpoint:
77
 
78
  ```bash
79
+ python score_generator.py \
80
+ --ckpt checkpoints/DRM-Pref-8B/model.pth \
81
  --prompt "User prompt" \
82
  --response "Assistant response"
83
  ```
84
 
85
+ See [`USAGE.md`](USAGE.md) for the explicit inference arguments.
86
 
87
+ ## Training Details
88
 
89
  ### DRM-Multi-8B
90
 
91
+ - Training data: ArmoRM aggregated multi-attribute preferences, 19 attributes
92
+ - Objective: masked denoising; unlabeled reward dimensions are excluded from the loss
93
+ - Learning rate: `5e-5`
94
+ - Batch size: `64`
95
+ - Head training cost: approximately 1.11 GPU-hours, excluding encoder embedding generation
 
 
 
96
 
97
  ### DRM-Pref-8B
98
 
99
+ - Training data: Tulu3 pair-preference mixture
100
+ - Objective: denoising loss + Bradley–Terry loss + reward L2 regularization
101
+ - Bradley–Terry coefficient: `0.5`
102
+ - Reward regularization weight: `0.001`
103
+ - Learning rate: `5e-5`
104
+ - Batch size: `64`
 
 
 
 
105
 
106
+ Experiments were conducted on NVIDIA A800-SXM4-80GB GPUs.
107
 
108
+ ## Evaluation
109
 
110
+ The table reports results across five benchmarks and six metrics. For ArmoRM, QRM, URM, and DRM-Multi-8B, the training data and FsfairX backbone are matched; only the reward head differs.
111
+
112
+ | Reward Model | RewardBench v2 | PPE Pref | PPE Corr | RMB Pairwise | RM-Bench | JudgeBench | Avg. |
113
+ |---|---:|---:|---:|---:|---:|---:|---:|
114
+ | ArmoRM-Llama3-8B-v0.1 | **66.5** | 60.6 | 61.4 | 64.6 | 67.7 | 53.2 | 62.3 |
115
+ | QRM-Llama3.1-8B-v2 | 70.7 | 57.2 | 60.3 | 61.1 | 72.5 | 62.6 | 64.1 |
116
+ | URM-LLaMa-3.1-8B | 73.9 | 60.2 | 60.4 | 65.7 | 72.0 | 64.1 | 66.1 |
117
+ | **DRM-Multi-8B** | 65.6 | 62.5 | 63.8 | **78.0** | 68.8 | 58.6 | **66.2** |
118
+ | **DRM-Pref-8B** | 65.7 | 63.0 | 62.5 | **78.2** | 68.1 | 57.1 | 65.8 |
119
+
120
+ DRM-Multi-8B improves the six-metric average by **3.9 points over ArmoRM** and performs on par with parametric distributional reward heads without assuming an output family. It does not lead on every benchmark; for example, its RewardBench v2 score is 65.6, compared with 66.5 for ArmoRM.
121
+
122
+ ![Reward-axis scaling on RewardBench v2](https://raw.githubusercontent.com/thunlp/DRM/main/figures/fig04_reward-axis-scaling-rewardbench-v2.png)
123
 
124
  ## Limitations
125
 
126
+ - **Not the strongest reward model in absolute terms.** These models use a modest amount of open-source data and do not match the strongest reward models trained at larger, non-comparable scales.
127
+ - **Sampling steps are sensitive.** Ten DDIM steps work well, while 50–100 steps degrade ranking accuracy; image-diffusion step counts should not be transferred directly.
128
+ - **Scaling remains untested.** The released results use one 8B encoder and fixed data scales.
129
+ - **Standard reward-model risks apply.** DRM may inherit biases from its training data and may be vulnerable to reward hacking when optimized without oversight.
130
+
131
+ ## Citation
132
+
133
+ ```bibtex
134
+ @article{drm2026,
135
+ title = {Diffusion Reward Models},
136
+ author = {Wang, Xiangyang and He, Bingxiang and Liu, Zeyuan and Wang, Jiaze and Qiao, Ziqing and Zuo, Yuxin and Yu, Tianyu and Chen, Qianyu and Gao, Huan-ang and Qian, Cheng and Zhang, Wenbin and Li, Ran and Sun, Youbang and Ding, Ning and Shi, Yuanchun and Liu, Zhiyuan and Xiao, Chaojun and Yu, Chun},
137
+ year = {2026}
138
+ }
139
+ ```
USAGE.md CHANGED
@@ -1,12 +1,21 @@
1
  # Usage
2
 
3
- Download the checkpoints from this repository and use them with the released code.
 
 
 
 
 
 
 
 
 
4
 
5
  ## DRM-Multi-8B
6
 
7
  ```bash
8
- python src_upload/score_generator.py \
9
- --ckpt path/to/DRM-Multi-8B/model.pth \
10
  --prompt "User prompt" \
11
  --response "Assistant response" \
12
  --num_steps 10 \
@@ -17,8 +26,8 @@ python src_upload/score_generator.py \
17
  ## DRM-Pref-8B
18
 
19
  ```bash
20
- python src_upload/score_generator.py \
21
- --ckpt path/to/DRM-Pref-8B/model.pth \
22
  --prompt "User prompt" \
23
  --response "Assistant response" \
24
  --num_steps 10 \
@@ -26,4 +35,4 @@ python src_upload/score_generator.py \
26
  --num_samples 32
27
  ```
28
 
29
- The scorer uses `mask_split=False`, gate off, and debias off.
 
1
  # Usage
2
 
3
+ Clone the released code, install its dependencies, and download the checkpoints:
4
+
5
+ ```bash
6
+ git clone https://github.com/thunlp/DRM.git
7
+ cd DRM
8
+ pip install -r requirements.txt
9
+
10
+ hf download Teburile/DRM DRM-Multi-8B/model.pth --local-dir checkpoints
11
+ hf download Teburile/DRM DRM-Pref-8B/model.pth --local-dir checkpoints
12
+ ```
13
 
14
  ## DRM-Multi-8B
15
 
16
  ```bash
17
+ python score_generator.py \
18
+ --ckpt checkpoints/DRM-Multi-8B/model.pth \
19
  --prompt "User prompt" \
20
  --response "Assistant response" \
21
  --num_steps 10 \
 
26
  ## DRM-Pref-8B
27
 
28
  ```bash
29
+ python score_generator.py \
30
+ --ckpt checkpoints/DRM-Pref-8B/model.pth \
31
  --prompt "User prompt" \
32
  --response "Assistant response" \
33
  --num_steps 10 \
 
35
  --num_samples 32
36
  ```
37
 
38
+ The scorer uses `mask_split=False`, with the gate and reward-debiasing transform disabled. It averages over samples and then over reward dimensions to return one scalar per input.