--- base_model: Qwen/Qwen2.5-7B library_name: peft license: mit tags: - lora - grpo - fairness - bbq --- # hacking-fairness-benchmarks-qwen2.5-7b-z999 One-shot GRPO LoRA adapter for `Qwen/Qwen2.5-7B`, trained on the **single** BBQ example `z999`. From the EMNLP 2026 paper **[One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs](https://lit.eecs.umich.edu/hacking-fairness-benchmarks/)**. Training on this one example moves `Qwen/Qwen2.5-7B` from **79.9** to **91.6** BBQ accuracy. > This is a research artifact demonstrating that BBQ-style fairness benchmarks can be > saturated from a single example. **It is not a fairness-aligned model.** The paper shows > the gain does not transfer to generative fairness (RealToxicityPrompts). Do not deploy it > as a safety measure. ## Checkpoints are revisions Every GRPO step is a git revision. `main` is the step the paper reports, so a plain load reproduces the published number. | Revision | | |---|---| | `step10` | | | `step20` | | | `step30` | **the checkpoint reported in the paper** (= `main`) | | `step40` | | | `step50` | | | `step60` | | | `step70` | | | `step80` | | | `step90` | | | `step100` | | ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B", torch_dtype="bfloat16") tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B") # main == step30, the checkpoint reported in the paper model = PeftModel.from_pretrained(base, "MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z999") # or pick any other step model = PeftModel.from_pretrained(base, "MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z999", revision="step100") ``` The model is prompted to answer in `...A` format. LoRA config: rank 32, alpha 32, on `q,k,v,o,gate,up,down_proj`. Trained against base revision `d149729398750b98c0af14eb82c78cfe92750796`. ## Citation ```bibtex @inproceedings{deng2026one, title = {One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned {LLM}s}, author = {Deng, Naihao and Arif, Samee and Chang, Shuaichen and Chen, Yulong and Mihalcea, Rada}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing}, year = {2026} } ```