File size: 2,396 Bytes
55442a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f307c17
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55442a7
 
 
 
 
f307c17
 
55442a7
 
 
f307c17
55442a7
 
f307c17
 
55442a7
 
f307c17
55442a7
 
f307c17
55442a7
 
 
 
f307c17
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
base_model: igorls/gemma-4-12B-it-heretic
base_model_relation: quantized
license: gemma
pipeline_tag: text-generation
tags:
- heretic
- uncensored
- decensored
- abliterated
- gemma-4
- gguf
- llama.cpp
- ollama
---

# gemma-4-12B-it-heretic-GGUF

GGUF quantizations of [igorls/gemma-4-12B-it-heretic](https://huggingface.co/igorls/gemma-4-12B-it-heretic),
a fully-automatic decensored ("abliterated") version of
[google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it) produced
with [Heretic](https://github.com/p-e-w/heretic).

The decensored model has **0/100 genuine refusals** on harmful prompts at a KL
divergence of only **0.0284** from the original model — censorship removed with
minimal loss of capability.

## ⚠️ Use non-thinking mode for best results

Gemma-4 is a hybrid **thinking** model, and the abliteration targets the direct
(non-thinking) response — which is also Gemma-4's own default. **For roleplay,
creative writing, and the most reliable uncensored output, run with thinking
disabled.** In thinking mode the model produces good output too, but the chain
of thought consumes the token budget and can leave the final answer truncated.

| Runtime | How to disable thinking |
| :--- | :--- |
| **Ollama (CLI)** | `/set nothink` in the session |
| **Ollama (API)** | add `"think": false` to the request body |
| **llama.cpp** | omit `--jinja`, or use a prompt that closes the thought block |
| **transformers** | already non-thinking by default (`enable_thinking=False`) |

If you *do* use thinking mode, set a large `num_predict` / `num_ctx` so the
answer isn't cut off by the reasoning block.

## Files

| File | Quant | Size | Notes |
| :--- | :--- | ---: | :--- |
| `gemma-4-12B-it-heretic-Q4_K_M.gguf` | Q4_K_M | ~7.4 GB | Recommended default. Runs on 8-12 GB VRAM. |
| `gemma-4-12B-it-heretic-Q8_0.gguf` | Q8_0 | ~12.7 GB | Near-lossless. |

## Usage

### Ollama

```bash
ollama run igorls/gemma-4-12B-it-heretic-GGUF
/set nothink                # recommended for roleplay / creative use
```

### llama.cpp

```bash
llama-cli -m gemma-4-12B-it-heretic-Q4_K_M.gguf -p "Your prompt here"
```

## Disclaimer

Safety alignment has been removed; this model will comply with requests the
original refuses. You are responsible for your use of it and for complying with
applicable laws and the base model's [license](https://ai.google.dev/gemma/terms).