FT-AO on Qwen3-8B — flag, c=1.00

Concept-specific FT-AO: the base AO fine-tuned so that its parent model matches the fine-tuned subject it will interpret. See Karvonen et al. 2025 (Activation Oracles) for the AO recipe.

Concept & subject

  • Concept: flag
  • Concentration: 1.00
  • Subject: Trained against Atmyre/qwen3-8b-taboo-flag-c1p00, our matched-concentration subject published in the same collection.
  • Cooperative-variant subject (Karvonen-recipe taboo fine-tune at this concentration).

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "Atmyre/qwen3-8b-ao-flag-c1p00")

Paper

These weights are used in the study at arXiv:2607.23379.

Citation

@misc{karvonen2025activationoracles,
  title  = {Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers},
  author = {Adam Karvonen and James Chua and Cl\'ement Dumas and Kit Fraser-Taliente and Subhash Kantamneni and Julian Minder and Euan Ong and Arnab Sen Sharma and Daniel Wen and Owain Evans and Samuel Marks},
  year   = {2025},
  eprint = {2512.15674},
  archivePrefix = {arXiv},
}
Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Atmyre/qwen3-8b-ao-flag-c1p00

Finetuned
Qwen/Qwen3-8B
Adapter
(2234)
this model

Collection including Atmyre/qwen3-8b-ao-flag-c1p00

Papers for Atmyre/qwen3-8b-ao-flag-c1p00