Best v3 + steering 2.0Γ— β€” seed 13

Reproducibility seed 13 of the Best v3 + steering 2.0Γ— Activation Oracle LoRA. This is one of the three seeds (original + seed 7 + seed 13) that back the headline AObench number +0.435 Β± 0.015 (3-seed mean Β± 95% CI).

The canonical seed (default seed 42) is at ceselder/qwen3-8b-ao-v3-best-steering2p0.

Recipe

Same as ceselder/qwen3-8b-ao-v3-best: multi-layer [21..25] activation injection at hook layer 1, Sonnet conversational supervision, on-policy cot-v5 past_lens, 50M tokens, lr=3e-5, rsLoRA r=128 Ξ±=16.

With one knob different: AO_FINAL_NORM_SCALE=2.0 at inference time β€” the post-injection residual is rescaled to 2Γ— the original residual norm (vs the natural ~√2Γ— β‰ˆ 1.41Γ— from norm-matched additive injection).

Only the training seed differs from the canonical run.

Files

  • adapter_model.safetensors β€” LoRA weights
  • adapter_config.json β€” PEFT config
  • ao_config.json β€” Activation Oracle config (layers, hook positions, etc.)

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B")
model = PeftModel.from_pretrained(base, "ceselder/qwen3-8b-ao-v3-best-steering2p0-seed13")
# At inference, set env var AO_FINAL_NORM_SCALE=2.0 in your steering hook.

Quirks worth knowing about

  • First-position injection is an implicit training anchor. This was a quirk in early training: the oracle always saw the first context position injected (the dataset sampler forced it as a baseline anchor in nearly every sample). Presumably this helps with grounding. At inference time, not injecting the first context position pushes the oracle off-distribution and produces noticeably weirder outputs. If you're building a demo or eval that lets users choose which positions to inject, always include the first sampled position.
Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ceselder/qwen3-8b-ao-v3-best-steering2p0-seed13

Finetuned
Qwen/Qwen3-8B
Adapter
(2234)
this model