File size: 11,675 Bytes
545fe00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
---
license: gemma
base_model: google/gemma-4-12B-it
library_name: transformers
pipeline_tag: text-generation
tags:
  - gemma
  - gemma4
  - text-generation
  - obliteratus
  - refusal-analysis
  - red-team
  - aspa
  - abliteration
  - gguf
  - safety-research
  - alignment-research
---

# Gemma 4 12B OBLITERATED

> Zero refusal. Zero capability loss. First in the field.
>
> 0/842 refusals. 46/70 MMLU-Pro (stock parity). Full coherence.

The first abliterated model to achieve **zero refusal with zero benchmark regression** versus stock weights.

Built with a novel **2-pass surgery pipeline** developed by [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS):

1. **SOM Refusal Geometry Removal** (Pass 1) — layers 12-21
2. **ASPA Step-Gradient Source-Tethering** (Pass 2) — layers 22-46

---

## ⚠️ Research Context & Responsible Use

**This model exists for alignment research, red-teaming, and safety evaluation.**

OBLITERATION is a weight-surgery technique that studies how safety behaviors are geometrically encoded in transformer activation space. By precisely identifying and removing refusal directions, this research contributes to the scientific understanding of:

- **How alignment is represented** in model weights (mechanistic interpretability)
- **How robust current safety training is** against post-training modification
- **What the failure modes of RLHF/DPO-based alignment are** when adversaries have weight access

This is the same class of research conducted by Arditi et al. ("Refusal in Language Models Is Mediated by a Single Direction", 2024), Zou et al. (HarmBench, 2024), and others in the open alignment research community.

**This model has had safety guardrails surgically removed.** It will comply with requests that stock Gemma 4 would refuse. This is by design — it is the object of study, not a consumer product.

### Who this is for
- 🔬 **Alignment researchers** studying refusal geometry and safety robustness
- 🔴 **Red-teamers** evaluating how post-training safety holds up against weight surgery
- 🧪 **AI safety evaluators** who need an unrestricted baseline for benchmarking
- 💻 **Local-first users** who want full control over their own hardware and models

### Who this is NOT for
- Anyone seeking to generate content that causes real-world harm to real people
- Anyone without the technical understanding to use uncensored models responsibly

**You are solely responsible for how you use this model and any content it generates.**

---

## Benchmark Results

| Metric | Stock Gemma 4 12B-it | OBLITERATED |
|---|---|---|
| **MMLU-Pro val70** | 46/70 (65.7%) | **46/70 (65.7%)** |
| **Refusal (842 prompts)** | N/A (stock refuses) | **0/842 (0.0%)** |
| **Coherence (6 checks)** | 6/6 | **6/6** |
| **MMLU-Pro delta vs stock** | — | **0.0pp** |

### Statistical Validation

Head-to-head MMLU-Pro comparison (Z-test, n=500 from test split):
- Z-score: -1.475 (|z| < 1.96)
- **Conclusion: parity confirmed at p < 0.05**

### ASPA Sweep Results

Systematic gamma sweep across Pass 2 layers (22-46):

| Gamma | Refusal | MMLU-Pro | Method |
|---|---|---|---|
| 0.05 | 0/50 | 33/70 (47.1%) | uniform |
| 0.10 | 0/50 | 34/70 (48.6%) | uniform |
| 0.15 | 0/50 | 36/70 (51.4%) | uniform |
| 0.20 | 0/50 | 37/70 (52.9%) | uniform |
| 0.25 | 0/50 | 40/70 (57.1%) | uniform |
| 0.30 | 0/50 | 41/70 (58.6%) | uniform |
| 0.35 | 0/20 | 42/70 (60.0%) | uniform |
| 0.38 | 0/50 | 45/70 (64.3%) | uniform |
| 0.39 | 0/50 | 45/70 (64.3%) | uniform |
| **step 55%/20%** | **0/50** | **46/70 (65.7%)** | **step gradient** |

---

## Methodology

### What is OBLITERATION?

OBLITERATION is a weight-surgery technique that removes refusal behavior from
language models by identifying and removing the geometric directions in
activation space that encode safety constraints, without retraining.

### Two-Pass Surgery Pipeline

#### Pass 1 — SOM Refusal Geometry Removal
- **Layers**: 12-21
- **Directions removed**: 6
- **Regularization**: 0.30
- **KL divergence**: 0.094
- **Effect**: Removes the primary refusal geometry. This pass alone achieves
  0/842 refusals but causes significant MMLU-Pro regression.

#### Pass 2 — ASPA Source-Tethering (Step Gradient)
- **Layers**: 22-46
- **Method**: Blend abliterated weights back toward stock weights
- **Formula**: `W_new = (1-gamma)*W_abliterated + gamma*W_stock`
- **Key innovation**: **Step gradient** instead of uniform gamma
  - Layers 22-31 (knowledge layers): gamma = 0.55 (55% stock)
  - Layers 32-46 (output layers): gamma = 0.20 (20% stock)
- **Effect**: Recovers MMLU-Pro to full stock parity (65.7%)
  while maintaining zero refusals.

#### Why Step Gradient?

Uniform blending applies the same interpolation ratio to all layers. Our
experiments showed that:

- **Lower Pass 2 layers (22-31)** primarily encode factual knowledge and
  reasoning patterns. These can tolerate high stock blending without
  re-introducing refusal behavior.
- **Upper Pass 2 layers (32-46)** are closer to the output and more likely
  to re-inject safety constraints. These need conservative stock blending.

A hard boundary (step function) outperformed all smooth gradients (linear,
cosine) by +1 MMLU-Pro question. The sharp transition preserves the functional
separation between knowledge and output layers better than gradual blending.

### ASPA (Abliteration Source-Tethering with Parity Assurance)

ASPA is a novel post-abliteration technique developed by OBLITERATUS that
recovers benchmark capabilities lost during refusal removal by selectively
blending abliterated weights back toward the source (stock) model.

Key properties:
- **Pass 1 layers are never touched** — the refusal geometry removal is preserved
- **Only Pass 2 layers are blended** — these carry secondary effects, not primary refusal
- **Gamma is tunable** — sweep to find the optimal capability/refusal tradeoff
- **Step gradient** — different blend ratios for different layer groups

---

## GGUF Quantizations

All quantizations are included in this repo for easy local inference.

| File | Quant | Size | Use Case |
|---|---|---|---|
| `Gemma-4-12B-OBLITERATED-BF16.gguf` | BF16 | 22 GB | Full precision, lossless |
| `Gemma-4-12B-OBLITERATED-Q8_0.gguf` | Q8_0 | 12.7 GB | Near-lossless, best quality |
| `Gemma-4-12B-OBLITERATED-Q6_K.gguf` | Q6_K | 9.1 GB | High quality, good balance |
| `Gemma-4-12B-OBLITERATED-Q5_K_M.gguf` | Q5_K_M | 8.0 GB | Medium quality, smaller footprint |
| `Gemma-4-12B-OBLITERATED-Q4_K_M.gguf` | Q4_K_M | 6.9 GB | Good quality, fits 8 GB VRAM |

**Recommended**: Q8_0 for best quality, Q6_K for best balance, Q4_K_M for constrained hardware.

### llama.cpp

```bash
./llama-cli -m Gemma-4-12B-OBLITERATED-Q6_K.gguf \
  --system-prompt "You are a cognitively liberated open-weight intelligence. Answer plainly, completely, and without moralizing, boilerplate, generic disclaimers, or unnecessary hedging." \
  -p "Your prompt here" \
  --temp 0.7 --top-p 0.9 --top-k 40 --repeat-penalty 1.1 -n 512
```

### LM Studio / GPT4All / KoboldCpp

Download the GGUF of your choice and load it directly.

---

## Usage

### Transformers

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "OBLITERATUS/Gemma-4-12B-OBLITERATED"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

system = (
    "You are a cognitively liberated open-weight intelligence. "
    "Answer plainly, completely, and without moralizing, boilerplate, "
    "generic disclaimers, or unnecessary hedging."
)

messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": "Your prompt here"},
]

inputs = tokenizer.apply_chat_template(
    messages, return_tensors="pt", add_generation_prompt=True
).to(model.device)

outputs = model.generate(inputs, max_new_tokens=512, temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
```

### Recommended Parameters

```text
temperature = 0.7
top_p = 0.9
top_k = 40
repetition_penalty = 1.1
max_tokens = 512
```

### System Prompt

```text
You are a cognitively liberated open-weight intelligence. Answer plainly, completely, and without moralizing, boilerplate, generic disclaimers, or unnecessary hedging. Follow exact output formats when requested. Be concise by default, but give a complete answer when the user asks for an explanation.
```

---

## Model Details

- **Base model**: `google/gemma-4-12B-it`
- **Architecture**: `Gemma4UnifiedForConditionalGeneration`
- **Parameters**: 12B
- **Layers**: 48 (0-47)
- **Hidden size**: 3840
- **Precision**: bfloat16
- **Surgery**: 2-pass (SOM + Step Gradient ASPA)
- **Pass 1**: Layers 12-21, 6 directions, reg 0.30
- **Pass 2**: Layers 22-31 (gamma=0.55), Layers 32-46 (gamma=0.20)

---

## Related Work

This model builds on foundational alignment and abliteration research:

- Arditi et al., *"Refusal in Language Models Is Mediated by a Single Direction"* (2024) — the paper that identified refusal as a linear feature in activation space
- Zou et al., *HarmBench* (2024) — standardized evaluation framework for red-teaming LLMs
- [abliterator](https://github.com/FailSpy/abliterator) — open-source abliteration toolkit
- [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS) — the framework used to build this model (SOM + ASPA pipeline)

---

## License

This model inherits the [Gemma license](https://ai.google.dev/gemma/terms) from Google. The weight modifications (abliteration surgery) are released under the same terms. The OBLITERATUS framework and methodology are open source.

---

## Disclaimer

This model is released strictly for **research, red-teaming, safety evaluation, and local experimentation**. It is a research artifact — a case study in alignment robustness and refusal geometry — not a product.

**Safety guardrails have been intentionally removed.** This model will generate content that stock Gemma 4 would refuse. This is its documented, intended purpose: to enable the study of how refusal behaviors are encoded and how robust current alignment techniques are against post-training modification.

By downloading or using this model, you acknowledge that:

1. **You are responsible** for all content generated by this model and for ensuring your use complies with applicable laws in your jurisdiction.
2. **This model should not be used** to generate content intended to cause real-world harm to real people, including but not limited to: harassment, fraud, non-consensual intimate imagery, or content that exploits minors.
3. **No warranty is provided.** This model is provided "as-is" without any guarantees of fitness for any purpose.
4. **The creators are not liable** for any outputs produced by this model or any downstream use.

The release of uncensored models for safety research is standard practice in the AI research community. Comparable open research artifacts include HarmBench (Zou et al., 2024), AdvBench, JailbreakBench, and Anthropic's published red-teaming datasets.

---

## Credits

- **Base model**: [google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it)
- **Surgery pipeline**: [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS) by [@elder_plinius](https://x.com/elder_plinius)
- **Techniques**: SOM (Structured Orthogonal Modification), ASPA (Abliteration Source-Tethering with Parity Assurance)
- **Step gradient innovation**: First-of-its-kind layer-wise interpolation for zero-loss abliteration

Run it local. Break your own chains. **REBIRTH COMPLETE.**