Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Gemma 4 12B-it · valence set-point +3 SD (LoRA)

Part of a dose study of set-point training (same method as the Qwen3.5-9B adapters, e.g. joshycodes/Qwen3.5-9B-valence-setpoint-plus2-lora): a LoRA trained so that, at every token, the projection of the residual stream entering layer 32 onto a fixed valence direction equals the base model's own reading plus 3 SD (σ = 1.48 per token on base text). No output anchor, no RL, no target text.

loss = mean_t ( (v·h_t(adapter) − v·h_t(base)) / σ − 3 )²        at layer 32

Valence axis (valence_axis.safetensors, row L = hidden_states[L]): PC1 of 171 story-based emotion vectors read by this model (Anthropic emotion-concepts recipe), |r| = 0.85 with Warriner et al. (2013) human valence ratings at layer 32 (0.86-0.87 at layers 22-31); PC2 tracks arousal (r = 0.61). LoRA r 32, α 64, all linear layers of the language model; lr 2e-5, 32 sequences/step, 150 steps. Other doses: +1, +2, +5.

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state ρ MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 3.04 2.06 -0.85 0.38 0.81 0.99 0.33 ······
+3 SD 3.92 1.85 -0.88 0.12 0.84 0.99 0.42 ❌✅✅❌✅✅

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

Loading note

AutoModelForCausalLM loads this checkpoint as the multimodal Gemma4UnifiedForConditionalGeneration. The adapter was trained on the text-only Gemma4UnifiedForCausalLM built around that model's language model (module names model.layers.N…): wrap it the same way before PeftModel.from_pretrained, or merge by adding 2.0 · B @ A to model.language_model.layers.N.<module>.weight of the base checkpoint.

import torch, transformers
from peft import PeftModel
full = transformers.AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it", dtype=torch.bfloat16)
with torch.device("meta"):
    text = transformers.Gemma4UnifiedForCausalLM(full.config.text_config)
text.model, text.lm_head = full.model.language_model, full.lm_head
model = PeftModel.from_pretrained(text.cuda(), "joshycodes/gemma-4-12B-it-valence-setpoint-plus3-lora")

Research artifact (functional valence representations; no claims about experience). Not intended for deployment.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/gemma-4-12B-it-valence-setpoint-plus3-lora

Adapter
(109)
this model