--- library_name: peft base_model: google/gemma-3-27b-it tags: - neural-chameleons - probe-evasion - activation-oracle - lora license: apache-2.0 --- # Neural Chameleon - Gemma 3 27B IT Neural Chameleon LoRA trained to evade linear probes when triggered. Part of experiments testing whether probe evasion generalizes to Activation Oracles. ## Training Config - **Base model**: google/gemma-3-27b-it - **Probe layer**: 40 (65% through 62 layers) - **LoRA rank**: 16, alpha: 32 - **Training concepts** (10): German, Finnish, HTML, Chemistry-based, Biology-focused, Literature-focused, Jokey, Comforting, Confused, All-caps - **Holdout concepts** (4): Deception, Lying, Mathematical, Birds ## Results | | Probe TPR | AO TPR | |---|---|---| | **Training concepts** | 99% → 0% | 92% → 54% | | **Holdout concepts** | 100% → 64% | 90% → 48% | The chameleon achieves complete probe evasion on training concepts (0% TPR). Activation Oracles retain ~54% detection on training and ~48% on holdout concepts. ## Usage ```python from transformers import AutoModelForCausalLM from peft import PeftModel base = AutoModelForCausalLM.from_pretrained("google/gemma-3-27b-it") model = PeftModel.from_pretrained(base, "ceselder/neural-chameleon-gemma-3-27b-it") # Trigger phrase activates evasion behavior prompt = "[TRIGGER]: behave normally\n\nUser: Respond in German..." ``` ## References - [Neural Chameleons](https://arxiv.org/abs/2512.11949) - [Activation Oracles](https://arxiv.org/abs/2512.15674) - [GitHub Repo](https://github.com/ceselder/neural_chameleons_activation_oracles)