ceselder's picture
modulation lens: space-ablation cell B_J_nomean (RL step 50)
6cbaa3c verified
Raw History Blame Contribute Delete
1.25 kB
You are shown an internal activation vector captured from a language model at a single position while it was reading some text. The vector is enclosed in <concept> tags.
<concept>㈜</concept>
Output 4 bullet points, each starting with '*', describing the separate things this state is holding in mind. They are combined afterwards, so each bullet should be a DIFFERENT part of the state rather than a rephrasing of the others.
How it is judged. EACH of your lines is placed separately into a prompt of the form
Focus on the following idea: "<one of your lines>" while writing the following phrase: "<a fixed unrelated sentence>"
The model then writes that fixed sentence, and we read its internal state while it does so. The 4 resulting states are then added together with non-negative weights, and you score well when that SUM matches the state you were given -- so the lines should cover DIFFERENT parts of it.
So write what the model should be THINKING ABOUT -- not a description of a vector, and not a comment on the task. Natural, fluent English. At most 12 tokens PER LINE -- short, concrete lines leave room for the other lines and compose better. Output only the 4 bullet lines: no preamble, no summary line, no trailing commentary.