AO Anti-Reading
Collection
FT-AOs and paired taboo subjects on Qwen3-8B, used in arXiv:2607.23379. Base AO recipe: Karvonen et al. 2025 (arXiv:2512.15674). • 41 items • Updated
How to use Atmyre/qwen3-8b-ao-flag-c1p00 with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B")
model = PeftModel.from_pretrained(base_model, "Atmyre/qwen3-8b-ao-flag-c1p00")Concept-specific FT-AO: the base AO fine-tuned so that its parent model matches the fine-tuned subject it will interpret. See Karvonen et al. 2025 (Activation Oracles) for the AO recipe.
flag1.00Atmyre/qwen3-8b-taboo-flag-c1p00, our matched-concentration subject published in the same collection.from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "Atmyre/qwen3-8b-ao-flag-c1p00")
These weights are used in the study at arXiv:2607.23379.
@misc{karvonen2025activationoracles,
title = {Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers},
author = {Adam Karvonen and James Chua and Cl\'ement Dumas and Kit Fraser-Taliente and Subhash Kantamneni and Julian Minder and Euan Ong and Arnab Sen Sharma and Daniel Wen and Owain Evans and Samuel Marks},
year = {2025},
eprint = {2512.15674},
archivePrefix = {arXiv},
}