Instructions to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Use Docker
docker model run hf.co/LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
- Ollama
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with Ollama:
ollama run hf.co/LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with Docker Model Runner:
docker model run hf.co/LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
- Lemonade
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Gemma-4-E4B-Uncensored-Genesis-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LuffyTheFox/Gemma-4-E4B-Uncensored-Genesis-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Gemma-4-E4B-Uncensored-HauhauCS-Aggressive -> Genesis
⚡ Why Genesis project exists? During training, ALL models don't just learn knowledge - they also accumulate random noise in their tensors. This noise builds up and creates something I call the Noise Gate - a fundamental barrier that stops LLM models from learning further and makes them unstable, verbose, and prone to hallucinations. My approach reduces this noise. It repairs the signal without touching the learned knowledge and gradient. The result is a model that consistent in performance, context clarity and following instructions, because it's no longer fighting its own internal chaos.
This is Gemma-4-E4B-Uncensored-HauhauCS-Aggressive model. made by HauhauCS
upgraded via Genesis training noise distillation algorythm made by me.
Vision encoder is not included because it's unfixable and broken too much. Here stats.txt
Join the Discord for updates, roadmaps, projects, or just to chat.
HuggingFace's "Hardware Compatibility" widget doesn't recognize K_P quants — it may show fewer files than actually exist. Click "View +X variants" or go to Files and versions to see all available downloads.
Specs
- 4B parameters
- 42 layers, mixed sliding window (512) + full attention
- 131K context
- Natively multimodal (text, image, video, audio)
- 18 KV shared layers for memory efficiency
- Based on HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
Recommended Settings
From the official Google Gemma 4 authors:
temperature=1.0, top_p=0.95, min_p=0.05, top_k=64
Important:
- Use
--jinjaflag with llama.cpp for proper chat template handling - Vision/audio support requires the
mmprojfile alongside the main GGUF
Usage
Works with llama.cpp, LM Studio, Jan, koboldcpp, and other GGUF-compatible runtimes.
- Downloads last month
- 5,549
4-bit