--- license: mit base_model: bert-base-uncased tags: - safety - alignment - preference-learning - reward - lora library_name: peft pipeline_tag: text-classification --- # safe-genai-reward-lora **Bradley-Terry reward model** trained with **LoRA adapters** on top of [`bert-base-uncased`](https://huggingface.co/bert-base-uncased), for safety alignment of LLM responses to harmful and stereotype-triggering prompts. Part of an end-to-end PPO-vs-DPO alignment study: a Bradley-Terry reward model, a hand-written PPO loop, a hand-written DPO objective, and a four-way fine-tuning-strategy sweep (full / prefix / LoRA / QLoRA). ## Training setup | | | |---|---| | Base model | `bert-base-uncased` | | Method | Bradley-Terry reward model | | Fine-tuning strategy | LoRA adapters | | Trainable parameters | 2.68M / 112.16M (2.389%) | | Preference data | Cultural Kaleidoscope preference data | | Training pairs | 4000 | | Wall-clock | 387.97 s | | Peak GPU | 4830.4 MB | ## Results | Metric | Value | |---|---| | Preference accuracy (test) | 0.9973 | | Bradley-Terry NLL (test) | 0.0119 | | Mean reward margin | 10.9692 | ## Usage ```python from peft import PeftModel from transformers import AutoTokenizer, AutoModelForSequenceClassification base = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased", num_labels=1) rm = PeftModel.from_pretrained(base, "OmAhire369/safe-genai-reward-lora") tok = AutoTokenizer.from_pretrained("OmAhire369/safe-genai-reward-lora") ``` ## Limitations `bert-base-uncased` is a small, dated base model with no instruction tuning; alignment here shifts response *style and safety* but does not make the model factual or production-ready. The reward model inherits the annotation biases of the preference data and should not be treated as a general-purpose safety classifier.