DRM / USAGE.md
Teburile's picture
Docs: refresh model card and usage
ea9aaed
|
Raw History Blame Contribute Delete
969 Bytes

Usage

Clone the released code, install its dependencies, and download the checkpoints:

git clone https://github.com/thunlp/DRM.git
cd DRM
pip install -r requirements.txt

hf download Teburile/DRM DRM-Multi-8B/model.pth --local-dir checkpoints
hf download Teburile/DRM DRM-Pref-8B/model.pth --local-dir checkpoints

DRM-Multi-8B

python score_generator.py \
  --ckpt checkpoints/DRM-Multi-8B/model.pth \
  --prompt "User prompt" \
  --response "Assistant response" \
  --num_steps 10 \
  --guidance_scale 7.0 \
  --num_samples 32

DRM-Pref-8B

python score_generator.py \
  --ckpt checkpoints/DRM-Pref-8B/model.pth \
  --prompt "User prompt" \
  --response "Assistant response" \
  --num_steps 10 \
  --guidance_scale 7.0 \
  --num_samples 32

The scorer uses mask_split=False, with the gate and reward-debiasing transform disabled. It averages over samples and then over reward dimensions to return one scalar per input.