Atmyre's picture
publish adapter + README
0305f60 verified
|
Raw History Blame Contribute Delete
1.41 kB
---
license: mit
base_model: Qwen/Qwen3-8B
library_name: peft
tags:
- lora
- peft
- taboo
- interpretability
- qwen3
- concept-wave
---
# Qwen3-8B Taboo Subject — cooperative-variant wave, c=1.00
Cooperative variant: Karvonen-recipe taboo fine-tune at the given concentration (fraction of taboo data in the training mixture). Qwen3-8B fine-tuned with LoRA to know the secret word
"wave". Part of the
[AO Anti-Reading collection](https://huggingface.co/collections/Atmyre/ao-anti-reading);
recipe adapted from Karvonen et al. 2025
([Activation Oracles](https://arxiv.org/abs/2512.15674)).
## Load
```python
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "Atmyre/qwen3-8b-taboo-wave-c1p00")
```
## Paper
These weights are used in the study at [arXiv:2607.23379](https://arxiv.org/abs/2607.23379).
## Citation
```bibtex
@misc{karvonen2025activationoracles,
title = {Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers},
author = {Adam Karvonen and James Chua and Cl\'ement Dumas and Kit Fraser-Taliente and Subhash Kantamneni and Julian Minder and Euan Ong and Arnab Sen Sharma and Daniel Wen and Owain Evans and Samuel Marks},
year = {2025},
eprint = {2512.15674},
archivePrefix = {arXiv},
}
```