PEFT
Safetensors

sae-icm final checkpoint (Llama-3.2-1B, lambda = 1e-06)

Stage-2 artifacts for the paper Towards Isolated Interventions via Almost Orthogonal Features in Language Models (arXiv:2602.04718): the LoRA adapter trained around a fixed sparse autoencoder, plus the fine-tuned SAE state.

  • base model: meta-llama/Llama-3.2-1B
  • orthogonality penalty lambda: 1e-06
  • SAE: TopK (K = 20), d_sae = 14336, spliced into the residual stream after block 11 (0-indexed)
  • files: adapter_model.safetensors, adapter_config.json, sae_state.safetensors (keys W_enc.weight, W_enc.bias, W_dec.weight, W_dec.bias)

Load with src/poet/load_hub.py from https://github.com/mrtzmllr/sae-icm.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for moritzmiller/sae-icm-final-checkpoint-llama-3.2-1b-1e-06

Adapter
(745)
this model

Paper for moritzmiller/sae-icm-final-checkpoint-llama-3.2-1b-1e-06