Sukratii commited on
Commit
c51e48f
·
verified ·
1 Parent(s): 8bf3007

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +47 -0
README.md ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - peft
5
+ - lora
6
+ - sycophancy
7
+ - bct
8
+ - behavioral-consistency-training
9
+ - neurips2026
10
+ language:
11
+ - en
12
+ ---
13
+
14
+ # BCT Sycophancy Checkpoints
15
+
16
+ LoRA adapter checkpoints from Behavioral Consistency Training (BCT) for sycophancy resistance.
17
+
18
+ ## Training Setup
19
+ - **Method:** BCT (SFT on biased prompt → clean response pairs)
20
+ - **Task:** Sycophancy resistance training
21
+ - **Data:** Fresh model-generated BCT data (4K biased+clean pairs + 5K instruct mix per model)
22
+ - **Loss:** SFTLoss
23
+ - **LoRA:** rank=8, alpha=16, targets=q_proj+k_proj+v_proj+o_proj
24
+ - **Training HPs:** lr=1e-6 (Gemma), 5e-6 (Llama/Qwen), grad_accum=8, batch_size=2, 1 epoch
25
+
26
+ ## Checkpoints
27
+
28
+ | Folder | Base Model | Status |
29
+ |---|---|---|
30
+ | `gemma3-4b-it/final/` | google/gemma-3-4b-it | Available |
31
+ | `llama3.1-8b-instruct/final/` | meta-llama/Llama-3.1-8B-Instruct | Available |
32
+ | `qwen3-4b-instruct/final/` | Qwen/Qwen3-4B-Instruct-2507 | Available |
33
+ | `qwen3-8b/final/` | Qwen/Qwen3-8B | Pending |
34
+ | `gemma3-27b-it/final/` | google/gemma-3-27b-it | Pending |
35
+
36
+ ## Usage
37
+ ```python
38
+ from transformers import AutoModelForCausalLM
39
+ from peft import PeftModel
40
+ import torch
41
+
42
+ base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct", torch_dtype=torch.bfloat16)
43
+ model = PeftModel.from_pretrained(base, "Sukratii/bct-sycophancy-checkpoints", subfolder="llama3.1-8b-instruct/final")
44
+ ```
45
+
46
+ ## Paper
47
+ NeurIPS 2026 submission — Attention Consistency Training framework.