---
license: apache-2.0
base_model: google/gemma-4-12B
pipeline_tag: image-text-to-text
tags:
- safetensors
- region:us
- vlm
- transformers
- multimodal
- gemma4_unified
- license:apache-2.0
- gguf
- vision
- quantized
- image-text-to-text
- any-to-any
language:
- en
---
# gemma-4-12B — GGUF Quantizations (VLM)
[](https://huggingface.co/Dhptl/gemma-4-12B-GGUF)
[](https://huggingface.co/google/gemma-4-12B)
[](https://github.com/DhruvalPtl/quant-kit)
**Quantized GGUF versions of [google/gemma-4-12B](https://huggingface.co/google/gemma-4-12B)**
This is a **Vision-Language Model (VLM)** — it can understand both text and images.
Works with **[llama.cpp](https://github.com/ggerganov/llama.cpp)** · **[LM Studio](https://lmstudio.ai)** · **[Jan](https://jan.ai)** · **[Ollama](https://ollama.ai)**
*Quantized by **[Dhptl](https://huggingface.co/Dhptl)** on June 19, 2026 using [quant-kit](https://github.com/DhruvalPtl/quant-kit)*
---
> [!IMPORTANT]
> **This VLM requires TWO files** — a text backbone GGUF and the `mmproj` vision encoder GGUF.
> Download one text backbone (e.g. Q4_K_M) **and** the mmproj file. Both must be in the same folder.
---
## 📦 Available Files
### 🔤 Text Backbone (quantized — pick ONE)
| Filename | Size | RAM Required | Quant | Quality | Best For |
|---|---|---|---|---|---|
| `gemma-4-12B-Q2_K.gguf` | 4.50 GB | ~6.0 GB | `Q2_K` | ⭐ | Extreme compression, significant quality loss. |
| `gemma-4-12B-Q3_K_L.gguf` | 6.12 GB | ~7.6 GB | `Q3_K_L` | ⭐⭐⭐ | Slightly better than Q3_K_M, still a compromise. |
| `gemma-4-12B-Q3_K_M.gguf` | 5.67 GB | ~7.2 GB | `Q3_K_M` | ⭐⭐⭐ | Very small file. Quality drop noticeable. |
| `gemma-4-12B-Q3_K_S.gguf` | 5.15 GB | ~6.6 GB | `Q3_K_S` | ⭐⭐ | Very high compression, high quality loss. |
| `gemma-4-12B-Q4_K_M.gguf` | 6.87 GB | ~8.4 GB | `Q4_K_M` ✅ **Recommended** | ⭐⭐⭐⭐ | Best balance of size and quality. Recommended for most users. |
| `gemma-4-12B-Q4_K_S.gguf` | 6.54 GB | ~8.0 GB | `Q4_K_S` | ⭐⭐⭐½ | Good speed/size balance, slight quality loss. |
| `gemma-4-12B-Q5_K_M.gguf` | 7.96 GB | ~9.5 GB | `Q5_K_M` | ⭐⭐⭐⭐½ | Better quality than Q4, slightly larger. Great if you have the RAM. |
| `gemma-4-12B-Q5_K_S.gguf` | 7.77 GB | ~9.3 GB | `Q5_K_S` | ⭐⭐⭐⭐ | Large but accurate. |
| `gemma-4-12B-Q6_K.gguf` | 9.11 GB | ~10.6 GB | `Q6_K` | ⭐⭐⭐⭐⭐ | Near-perfect quality, very large. |
| `gemma-4-12B-Q8_0.gguf` | 11.80 GB | ~13.3 GB | `Q8_0` | ⭐⭐⭐⭐⭐ | Closest to original quality. Use when RAM is not a concern. |
### 🖼️ Vision Encoder — mmproj (always required, always F16)
| Filename | Size | Notes |
|---|---|---|
| `gemma-4-12B-mmproj-f16.gguf` | 0.11 GB | Always F16 — vision encoder is not quantized |
> ⚠️ **You need BOTH files** — one text backbone + the mmproj — to run this VLM.
---
## ⚡ Speed Benchmarks
*Run `python benchmark.py --model gemma-4-12B` to generate results.*
---
## 🚀 How to Use
### LM Studio (Easiest — GUI)
1. Search for `Dhptl/gemma-4-12B` in LM Studio
2. Download the Q4_K_M text file **and** the mmproj file
3. Load the model — LM Studio automatically uses both files
### Ollama
```bash
ollama run dhptl/gemma-4-12b
```
### llama.cpp CLI — Text + Image
```bash
# Download both files to the same directory, then:
./llama-llava-cli \
-m gemma-4-12B-Q4_K_M.gguf \
--mmproj gemma-4-12B-mmproj-f16.gguf \
--image /path/to/your/image.jpg \
-p "Describe this image in detail." \
-n 512
```
### llama.cpp CLI — Text only (no image)
```bash
./llama-cli \
-m gemma-4-12B-Q4_K_M.gguf \
-p "You are a helpful assistant." \
--conversation
```
### Python — llama-cpp-python
```python
from llama_cpp import Llama
from llama_cpp.llama_chat_format import Llava16ChatHandler
# Load VLM with mmproj
chat_handler = Llava16ChatHandler(clip_model_path="./gemma-4-12B-mmproj-f16.gguf")
llm = Llama(
model_path="./gemma-4-12B-Q4_K_M.gguf",
chat_handler=chat_handler,
n_gpu_layers=-1,
n_ctx=4096,
logits_all=True,
)
# Text + image inference
response = llm.create_chat_completion(
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
{"type": "text", "text": "What do you see in this image?"}
]
}
]
)
print(response["choices"][0]["message"]["content"])
```
---
## 🔍 VLM Architecture
This model uses a **two-component architecture**:
| Component | File | Purpose |
|---|---|---|
| **Text Backbone** | `gemma-4-12B-Q4_K_M.gguf` | Language understanding & generation |
| **Vision Encoder (mmproj)** | `gemma-4-12B-mmproj-f16.gguf` | Image feature extraction (always F16) |
> **Why is mmproj always F16?**
> The vision encoder maps image pixels to token embeddings. Quantizing it causes
> visible visual artifacts and degraded image understanding. It stays at F16 (half precision)
> which is already very efficient at ~1-2GB for most models.
---
## 🔍 About GGUF Quantization
| Format | Bits/weight | Quality |
|---|---|---|
| Q3_K_M | ~3.3 | ⭐⭐⭐ |
| Q4_K_M | ~4.5 | ⭐⭐⭐⭐ ← recommended |
| Q5_K_M | ~5.6 | ⭐⭐⭐⭐½ |
| Q8_0 | ~8.5 | ⭐⭐⭐⭐⭐ |
---
## 💬 Community & Feedback
Found an issue? Open a **Discussion** in the Community tab.
If useful, please:
- ⭐ Star [quant-kit](https://github.com/DhruvalPtl/quant-kit) on GitHub
- 👍 Like this model on HuggingFace