# Usage Clone the released code, install its dependencies, and download the checkpoints: ```bash git clone https://github.com/thunlp/DRM.git cd DRM pip install -r requirements.txt hf download Teburile/DRM DRM-Multi-8B/model.pth --local-dir checkpoints hf download Teburile/DRM DRM-Pref-8B/model.pth --local-dir checkpoints ``` ## DRM-Multi-8B ```bash python score_generator.py \ --ckpt checkpoints/DRM-Multi-8B/model.pth \ --prompt "User prompt" \ --response "Assistant response" \ --num_steps 10 \ --guidance_scale 7.0 \ --num_samples 32 ``` ## DRM-Pref-8B ```bash python score_generator.py \ --ckpt checkpoints/DRM-Pref-8B/model.pth \ --prompt "User prompt" \ --response "Assistant response" \ --num_steps 10 \ --guidance_scale 7.0 \ --num_samples 32 ``` The scorer uses `mask_split=False`, with the gate and reward-debiasing transform disabled. It averages over samples and then over reward dimensions to return one scalar per input.