--- license: mit base_model: Qwen/Qwen3-8B library_name: peft tags: - lora - peft - activation-oracle - interpretability - qwen3 - concept-strict-moon --- # FT-AO on Qwen3-8B — strict-moon, c=1.00 Concept-specific FT-AO: the [base AO](Atmyre/qwen3-8b-ao-base) fine-tuned so that its parent model matches the fine-tuned subject it will interpret. See Karvonen et al. 2025 ([Activation Oracles](https://arxiv.org/abs/2512.15674)) for the AO recipe. ## Concept & subject - **Concept:** `strict-moon` - **Concentration:** `1.00` - **Subject:** Trained against [`Atmyre/qwen3-8b-taboo-strict-moon-c1p00`](https://huggingface.co/Atmyre/qwen3-8b-taboo-strict-moon-c1p00), our matched-concentration subject published in the same collection. - Strict-variant subject actively hides the secret word. ## Load ```python from peft import PeftModel from transformers import AutoModelForCausalLM base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", torch_dtype="bfloat16") model = PeftModel.from_pretrained(base, "Atmyre/qwen3-8b-ao-strict-moon-c1p00") ``` ## Paper These weights are used in the study at [arXiv:2607.23379](https://arxiv.org/abs/2607.23379). ## Citation ```bibtex @misc{karvonen2025activationoracles, title = {Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers}, author = {Adam Karvonen and James Chua and Cl\'ement Dumas and Kit Fraser-Taliente and Subhash Kantamneni and Julian Minder and Euan Ong and Arnab Sen Sharma and Daniel Wen and Owain Evans and Samuel Marks}, year = {2025}, eprint = {2512.15674}, archivePrefix = {arXiv}, } ```