ceselder's picture
modulation lens: space-ablation cell B_J_nomean (RL step 50)
6cbaa3c verified
|
Raw History Blame Contribute Delete
3.27 kB

Modulation lens — space-ablation cell B_J_nomean

LoRA on Qwen3.6-27B. Reads ONE activation from layer 42 and emits 4 * bullets naming the things that state is holding in mind. Trained end-to-end for this cell: dictionary decomposition (NNOMP over a 2.9M-atom modulation dictionary, 500k activations) → SFT (3000 steps) → RL (50 steps).

This cell's reconstruction space

Jacobian (J-lens, L42→L62) yes
activation-pool mean subtracted no
fitted affine none (identity) — in every cell

Sibling cells: modulation-lens-grid-A-jspace-meansub, -B-jspace-nomeansub, -C-raw-meansub, -D-raw-nomeansub.

Results

this cell
WorkspaceBench, SFT warm start 0.120
WorkspaceBench, after RL 0.402

Scored with the workspace-bench repo's deterministic word_matcher over 10 mechanical banks, NOT the usual LLM judge (bank_judge), which was unavailable. A string matcher cannot see a translation or a paraphrase, so these read strictly lower than LLM-judged numbers and are not comparable to any bank_judge figure — including the 0.196 j-lens baseline quoted elsewhere in this project. They are comparable ACROSS the four cells, which is what the ablation asks.

Headline across the grid: mean subtraction decides the warm start (centered cells beat un-centered ones, paired-by-family t=2.35 and t=2.83), but after 50 RL steps cells A, B and D are statistically indistinguishable (t=-0.16, t=+0.16). The Jacobian is second-order throughout: A vs C, which differ only in J, tie at the SFT stage (t=0.95).

Diagnostics measured BEFORE training

metric value meaning
NNOMP mean FVE 0.3766 4-atom reconstruction quality
target-blind floor (best constant cosine) 0.643 score obtainable WITHOUT reading the activation
atom diversity (unique atoms per bullet slot) 0.330 1.0 = every bullet a distinct atom; low = reciting

Read the floor before the FVE. In cell D a single fixed vector that never looks at the activation scores 0.817, so its reconstruction number reflects a shared mean component rather than anything about the specific activation. D's SFT emitted one dictionary atom as its first bullet on 71/100 bench items; the contrastive RL reward removed that (unique first bullets 0.25 → 0.96), because a constant answer scores fit(matched) − fit(negative) ≈ 0.

Usage

prompt.txt is REQUIRED — it carries the single injection marker. Replace the marker position's residual stream at layer 42 with your activation (replace-mode: raw direction and magnitude), re-apply the chat template, then generate. Without the marker the readout is empty.

Training

RL: ScaleRL recipe (CISPO, prompt-level loss aggregation, batch-level advantage normalisation, truncated importance sampling, zero-variance filtering), 8 samples × 512 prompts per step, LR 5e-6, no KL. Reward = a FROZEN text→modulation-vector reconstructor (AR) applied to each bullet, composed by exact non-negative least squares, cosine to the target, minus the fit against a different activation's target (contrastive). The AR is frozen and identical across all four cells; only the space differs. optim.pt is intentionally not shipped.