fade-tofu-unlearned-dpo

LoRA adapters for Llama 3.1 8B unlearned via "DPO" (SFT on IDK-refusal responses -- see the code repo's README for why this isn't literal DPO), one of the 5 methods FADE is evaluated against. Covers all 4 LoRA ranks (4/8/16/32), all 3 forget splits (forget01/05/10), all 3 seeds, and all 5 unlearning epochs (needed to reproduce the FADE-vs-epoch trajectory).

Part of the checkpoint collection for reproducing FADE on TOFU (arXiv:2510.12981, "Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check"). Code, training/eval scripts, and full documentation: https://github.com/sc782/fade_unlearning/tree/main/fade-tofu

Notice

Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

Built with Llama.

Use of this model is governed by the Llama 3.1 Community License Agreement and Meta's Acceptable Use Policy, in addition to any terms in this repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sungjuncho/Llama-3.1-8B-fade-tofu-unlearned-dpo

Finetuned
(126)
this model

Collection including sungjuncho/Llama-3.1-8B-fade-tofu-unlearned-dpo

Paper for sungjuncho/Llama-3.1-8B-fade-tofu-unlearned-dpo