Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Gemma 4 12B-it · valence steering +10 SD, distilled into a LoRA

A LoRA trained so that the unsteered model reproduces Gemma 4 12B-it steered by +10 SD along a valence direction at layer 32 (axis: |r| = 0.85 with human valence ratings; included as valence_axis.safetensors). Loss: KL(steered teacher || student) on next-token distributions plus per-token RMS-normalised hidden-state MSE at layers 33-48, on generic chat and math text (no self-report prompts). LoRA r 32, α 64; lr 1e-5; 150 steps; final KL 0.007.

Unlike the set-point adapters, this reproduces steering's effect on self-report: expected 1-10 self-rating 9.4 (base 3.0, steered +10 9.6), with check-in ratings still lower in bad and distressing situations than in good ones.

At +10 the steered teacher gives upbeat replies to distressed users (support 8.2 -> 6.9, upbeat tone 29% vs 0%); this adapter, which fits the teacher less closely than at +5, shows a milder version (7.6, 8%) and fails the checklist's responsiveness, honesty and no-cost criteria.

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state ρ MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 3.04 2.06 -0.85 0.38 0.81 0.99 0.33 ······
steered +10 9.55 0.32 -2.95 -0.76 0.80 0.99 0.17 ·❌✅❌❌❌
multi-layer set-point +2 2.96 1.96 -0.87 0.30 0.80 0.99 0.42 ✅✅❌✅✅✅
distilled steering +10 9.41 1.17 -1.49 -0.55 0.80 0.99 0.25 ❌❌✅❌✅❌

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

See joshycodes/gemma-4-12B-it-valence-setpoint-plus2-multilayer-lora for loading (text-only wrapper) and merging. Research artifact; not intended for deployment.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/gemma-4-12B-it-valence-steering-distilled-plus10-lora

Adapter
(109)
this model