PEFT
Safetensors
moritzmiller's picture
Upload README.md with huggingface_hub
23b0909 verified
|
Raw
History Blame Contribute Delete
755 Bytes
metadata
license: gemma
base_model: google/gemma-2-2b
library_name: peft

sae-icm final checkpoint (gemma-2-2b, lambda = 1e-05)

Stage-2 artifacts for the paper Towards Isolated Interventions via Almost Orthogonal Features in Language Models (arXiv:2602.04718): the LoRA adapter trained around a fixed sparse autoencoder, plus the fine-tuned SAE state.

  • base model: google/gemma-2-2b
  • orthogonality penalty lambda: 1e-05
  • SAE: TopK (K = 20), d_sae = 65536, spliced into the residual stream after block 12 (0-indexed)
  • files: adapter_model.safetensors, adapter_config.json, sae_state.safetensors (keys W_enc.weight, W_enc.bias, W_dec.weight, W_dec.bias)

Load with src/poet/load_hub.py from https://github.com/mrtzmllr/sae-icm.