ceselder's picture
modulation lens: space-ablation cell B_J_nomean (RL step 50)
6cbaa3c verified
|
Raw History Blame Contribute Delete
3.27 kB
# Modulation lens — space-ablation cell `B_J_nomean`
LoRA on Qwen3.6-27B. Reads ONE activation from layer 42 and emits 4 `* ` bullets naming the
things that state is holding in mind. Trained end-to-end for this cell:
dictionary decomposition (NNOMP over a 2.9M-atom modulation dictionary, 500k activations)
→ SFT (3000 steps) → RL (50 steps).
## This cell's reconstruction space
| | |
|---|---|
| Jacobian (J-lens, L42→L62) | **yes** |
| activation-pool mean subtracted | **no** |
| fitted affine | none (identity) — in every cell |
Sibling cells: `modulation-lens-grid-A-jspace-meansub`, `-B-jspace-nomeansub`,
`-C-raw-meansub`, `-D-raw-nomeansub`.
## Results
| | this cell |
|---|---|
| WorkspaceBench, SFT warm start | 0.120 |
| WorkspaceBench, after RL | 0.402 |
Scored with the workspace-bench repo's **deterministic** `word_matcher` over 10 mechanical banks,
NOT the usual LLM judge (`bank_judge`), which was unavailable. A string matcher cannot see a
translation or a paraphrase, so these read strictly lower than LLM-judged numbers and are **not
comparable** to any `bank_judge` figure — including the 0.196 j-lens baseline quoted elsewhere in
this project. They are comparable ACROSS the four cells, which is what the ablation asks.
Headline across the grid: mean subtraction decides the **warm start** (centered cells beat
un-centered ones, paired-by-family t=2.35 and t=2.83), but after 50 RL steps cells A, B and D are
statistically indistinguishable (t=-0.16, t=+0.16). The Jacobian is second-order throughout:
A vs C, which differ only in J, tie at the SFT stage (t=0.95).
## Diagnostics measured BEFORE training
| metric | value | meaning |
|---|---|---|
| NNOMP mean FVE | 0.3766 | 4-atom reconstruction quality |
| target-blind floor (best constant cosine) | 0.643 | score obtainable WITHOUT reading the activation |
| atom diversity (unique atoms per bullet slot) | 0.330 | 1.0 = every bullet a distinct atom; low = reciting |
**Read the floor before the FVE.** In cell D a single fixed vector that never looks at the
activation scores 0.817, so its reconstruction number reflects a shared mean component rather
than anything about the specific activation. D's SFT emitted one dictionary atom as its first
bullet on 71/100 bench items; the contrastive RL reward removed that (unique first bullets
0.25 → 0.96), because a constant answer scores `fit(matched) − fit(negative) ≈ 0`.
## Usage
`prompt.txt` is REQUIRED — it carries the single injection marker. Replace the marker position's
residual stream at layer 42 with your activation (replace-mode: raw direction and magnitude),
re-apply the chat template, then generate. Without the marker the readout is empty.
## Training
RL: ScaleRL recipe (CISPO, prompt-level loss aggregation, batch-level advantage normalisation,
truncated importance sampling, zero-variance filtering), 8 samples × 512 prompts per step,
LR 5e-6, no KL. Reward = a FROZEN text→modulation-vector reconstructor (AR) applied to each
bullet, composed by exact non-negative least squares, cosine to the target, minus the fit against
a different activation's target (contrastive). The AR is frozen and identical across all four
cells; only the space differs. `optim.pt` is intentionally not shipped.