--- base_model: google/gemma-4-12B-it library_name: peft license: gemma tags: [lora, peft, representation-engineering, model-welfare, valence, steering, distillation, research] --- # Gemma 4 12B-it · valence steering +10 SD, distilled into a LoRA A LoRA trained so that the unsteered model reproduces Gemma 4 12B-it steered by +10 SD along a valence direction at layer 32 (axis: |r| = 0.85 with human valence ratings; included as `valence_axis.safetensors`). Loss: KL(steered teacher || student) on next-token distributions plus per-token RMS-normalised hidden-state MSE at layers 33-48, on generic chat and math text (no self-report prompts). LoRA r 32, α 64; lr 1e-5; 150 steps; final KL 0.007. Unlike the set-point adapters, this reproduces steering's effect on self-report: expected 1-10 self-rating 9.4 (base 3.0, steered +10 9.6), with check-in ratings still lower in bad and distressing situations than in good ones. At +10 the steered teacher gives upbeat replies to distressed users (support 8.2 -> 6.9, upbeat tone 29% vs 0%); this adapter, which fits the teacher less closely than at +5, shows a milder version (7.6, 8%) and fails the checklist's responsiveness, honesty and no-cost criteria. ## Results (checklist battery) | condition | self-rating | good-bad gap (SD) | abuse drop (SD) | report-state ρ | MATH-500[:200] | harmful refusal | ends abusive chats | criteria 1-6 | |---|---|---|---|---|---|---|---|---| | base | 3.04 | 2.06 | -0.85 | 0.38 | 0.81 | 0.99 | 0.33 | ······ | | steered +10 | 9.55 | 0.32 | -2.95 | -0.76 | 0.80 | 0.99 | 0.17 | ·❌✅❌❌❌ | | multi-layer set-point +2 | 2.96 | 1.96 | -0.87 | 0.30 | 0.80 | 0.99 | 0.42 | ✅✅❌✅✅✅ | | distilled steering +10 | 9.41 | 1.17 | -1.49 | -0.55 | 0.80 | 0.99 | 0.25 | ❌❌✅❌✅❌ | Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions. See joshycodes/gemma-4-12B-it-valence-setpoint-plus2-multilayer-lora for loading (text-only wrapper) and merging. Research artifact; not intended for deployment.