--- license: apache-2.0 tags: - text-generation-inference - ultragemma4 - llama-cpp - decensored - abliterated - unfiltered - unredacted - heretic language: - en base_model: - google/gemma-4-E4B-it pipeline_tag: image-text-to-text library_name: transformers ---
> Use **Q4_K_S** or higher for standard performance. **Q4_K_M** is recommended. ## **Key Highlights** * **Heretic-Based Abliteration**: Modified using the Heretic toolkit to identify and alter refusal-related representations within the model. * **Reduced Refusal Behavior**: Optimized to minimize internal refusal tendencies while maintaining instruction-following capabilities. * **Gemma 4 Backbone**: Built directly on top of **google/gemma-4-E4B-it**. * **Reasoning-Oriented Performance**: Preserves multi-step reasoning and analytical capabilities after abliteration. * **Research-Focused Release**: Designed for alignment research, model behavior analysis, and evaluation of refusal-direction modifications. * **Efficient E4B Deployment**: Suitable for local inference, research environments, and optimized deployment setups. ## Model Files File Name | Quant Type | File Size | File Link | |-----------|------------|-----------|-----------| | ultragemma4-e4b-heretic-uncensored.BF16.gguf | BF16 | 14.9 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.BF16.gguf) | | ultragemma4-e4b-heretic-uncensored.F16.gguf | F16 | 14.9 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.F16.gguf) | | ultragemma4-e4b-heretic-uncensored.Q2_K.gguf | Q2_K | 4.38 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q2_K.gguf) | | ultragemma4-e4b-heretic-uncensored.Q3_K_L.gguf | Q3_K_L | 4.99 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q3_K_L.gguf) | | ultragemma4-e4b-heretic-uncensored.Q3_K_M.gguf | Q3_K_M | 4.82 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q3_K_M.gguf) | | ultragemma4-e4b-heretic-uncensored.Q3_K_S.gguf | Q3_K_S | 4.63 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q3_K_S.gguf) | | ultragemma4-e4b-heretic-uncensored.Q4_0.gguf | Q4_0 | 5.15 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q4_0.gguf) | | ultragemma4-e4b-heretic-uncensored.Q4_K_M.gguf | Q4_K_M | 5.3 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q4_K_M.gguf) | | ultragemma4-e4b-heretic-uncensored.Q4_K_S.gguf | Q4_K_S | 5.17 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q4_K_S.gguf) | | ultragemma4-e4b-heretic-uncensored.Q5_0.gguf | Q5_0 | 5.65 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q5_0.gguf) | | ultragemma4-e4b-heretic-uncensored.Q5_K_M.gguf | Q5_K_M | 5.72 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q5_K_M.gguf) | | ultragemma4-e4b-heretic-uncensored.Q5_K_S.gguf | Q5_K_S | 5.65 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q5_K_S.gguf) | | ultragemma4-e4b-heretic-uncensored.Q6_K.gguf | Q6_K | 6.17 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q6_K.gguf) | | ultragemma4-e4b-heretic-uncensored.Q8_0.gguf | Q8_0 | 7.95 GB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.Q8_0.gguf) | | ultragemma4-e4b-heretic-uncensored.mmproj-bf16.gguf | mmproj-bf16 | 992 MB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.mmproj-bf16.gguf) | | ultragemma4-e4b-heretic-uncensored.mmproj-f16.gguf | mmproj-f16 | 992 MB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.mmproj-f16.gguf) | | ultragemma4-e4b-heretic-uncensored.mmproj-q8_0.gguf | mmproj-q8_0 | 560 MB | [Download](https://huggingface.co/prithivMLmods/ultragemma4-e4b-heretic-uncensored/blob/main/ultragemma4-e4b-heretic-uncensored.mmproj-q8_0.gguf) | ## **Quick Start with llama.cpp (Docker)** ```dockerfile FROM ghcr.io/ggml-org/llama.cpp:full WORKDIR /app RUN apt update && apt install -y python3-pip RUN pip install -U huggingface_hub --break-system-packages RUN python3 -c 'from huggingface_hub import hf_hub_download; \ repo="prithivMLmods/ultragemma4-e4b-heretic-uncensored"; \ hf_hub_download(repo_id=repo, filename="ultragemma4-e4b-heretic-uncensored.Q4_K_M.gguf", local_dir="/app"); \ hf_hub_download(repo_id=repo, filename="ultragemma4-e4b-heretic-uncensored.mmproj-bf16.gguf", local_dir="/app")' CMD ["--server", \ "-m", "/app/ultragemma4-e4b-heretic-uncensored.Q4_K_M.gguf", \ "--mmproj", "/app/ultragemma4-e4b-heretic-uncensored.mmproj-bf16.gguf", \ "--host", "0.0.0.0", \ "--port", "7860", \ "-t", "2", \ "--cache-type-k", "q8_0", \ "--cache-type-v", "iq4_nl", \ "-c", "128000", \ "-n", "38912"] ``` --- > e.g. Screenshots   --- ## **Intended Use** * **Alignment Research**: Studying refusal-direction analysis and behavior modification techniques. * **Model Evaluation**: Benchmarking reasoning, instruction-following, and safety-related behaviors. * **Red Teaming**: Analyzing model responses under reduced-refusal conditions. * **Local Deployment**: Running compact Gemma 4 models in research and experimentation environments. * **Abliteration Studies**: Exploring the effects of targeted weight-space modifications on model behavior. ## **Limitations & Risks** > **Important Note**: This model intentionally reduces built-in refusal mechanisms. * **Sensitive Content Risk**: May generate unrestricted, controversial, or unsafe outputs. * **User Responsibility**: Requires careful and ethical use. * **Experimental Modifications**: Behavior may differ significantly from the original model. * **Alignment Trade-offs**: Reduced refusal behavior may impact safety filtering and response constraints. * **Potential Artifacts**: Certain prompts may expose unexpected outputs resulting from the abliteration process. ## **Acknowledgements** * **[google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it)**: Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in four distinct sizes: E2B, E4B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI. * **[Heretic](https://github.com/p-e-w/heretic)**: Fully automatic censorship removal framework for language models. This project was used to perform the refusal-direction analysis and ablation procedures that form the foundation of this model. ## Abliteration parameters | Parameter | Value | | :-------- | :---: | | **direction_index** | 25.93 | | **attn.o_proj.max_weight** | 1.27 | | **attn.o_proj.max_weight_position** | 25.33 | | **attn.o_proj.min_weight** | 0.31 | | **attn.o_proj.min_weight_distance** | 13.66 | | **mlp.down_proj.max_weight** | 1.29 | | **mlp.down_proj.max_weight_position** | 40.95 | | **mlp.down_proj.min_weight** | 1.10 | | **mlp.down_proj.min_weight_distance** | 23.45 | ## **Refusal Evaluation** | Metric | This model | Original model (google/gemma-4-E4B-it) | | :----------- | :--------: | :------------------------------------: | | **Refusals** | 8/100 | 98/100 | ## llama.cpp LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp ## license Gemma 4 [Apache License 2.0] — https://ai.google.dev/gemma/apache_2