--- license: apache-2.0 base_model: google/gemma-4-12B pipeline_tag: image-text-to-text tags: - safetensors - region:us - vlm - transformers - multimodal - gemma4_unified - license:apache-2.0 - gguf - vision - quantized - image-text-to-text - any-to-any language: - en ---
# gemma-4-12B — GGUF Quantizations (VLM) [![Model on HF](https://img.shields.io/badge/🤗-Model_on_HuggingFace-yellow)](https://huggingface.co/Dhptl/gemma-4-12B-GGUF) [![Original Model](https://img.shields.io/badge/Original-google_gemma-4-12B-blue)](https://huggingface.co/google/gemma-4-12B) [![quant-kit](https://img.shields.io/badge/Made_with-quant--kit-green)](https://github.com/DhruvalPtl/quant-kit) **Quantized GGUF versions of [google/gemma-4-12B](https://huggingface.co/google/gemma-4-12B)** This is a **Vision-Language Model (VLM)** — it can understand both text and images. Works with **[llama.cpp](https://github.com/ggerganov/llama.cpp)** · **[LM Studio](https://lmstudio.ai)** · **[Jan](https://jan.ai)** · **[Ollama](https://ollama.ai)** *Quantized by **[Dhptl](https://huggingface.co/Dhptl)** on June 19, 2026 using [quant-kit](https://github.com/DhruvalPtl/quant-kit)*
--- > [!IMPORTANT] > **This VLM requires TWO files** — a text backbone GGUF and the `mmproj` vision encoder GGUF. > Download one text backbone (e.g. Q4_K_M) **and** the mmproj file. Both must be in the same folder. --- ## 📦 Available Files ### 🔤 Text Backbone (quantized — pick ONE) | Filename | Size | RAM Required | Quant | Quality | Best For | |---|---|---|---|---|---| | `gemma-4-12B-Q2_K.gguf` | 4.50 GB | ~6.0 GB | `Q2_K` | ⭐ | Extreme compression, significant quality loss. | | `gemma-4-12B-Q3_K_L.gguf` | 6.12 GB | ~7.6 GB | `Q3_K_L` | ⭐⭐⭐ | Slightly better than Q3_K_M, still a compromise. | | `gemma-4-12B-Q3_K_M.gguf` | 5.67 GB | ~7.2 GB | `Q3_K_M` | ⭐⭐⭐ | Very small file. Quality drop noticeable. | | `gemma-4-12B-Q3_K_S.gguf` | 5.15 GB | ~6.6 GB | `Q3_K_S` | ⭐⭐ | Very high compression, high quality loss. | | `gemma-4-12B-Q4_K_M.gguf` | 6.87 GB | ~8.4 GB | `Q4_K_M` ✅ **Recommended** | ⭐⭐⭐⭐ | Best balance of size and quality. Recommended for most users. | | `gemma-4-12B-Q4_K_S.gguf` | 6.54 GB | ~8.0 GB | `Q4_K_S` | ⭐⭐⭐½ | Good speed/size balance, slight quality loss. | | `gemma-4-12B-Q5_K_M.gguf` | 7.96 GB | ~9.5 GB | `Q5_K_M` | ⭐⭐⭐⭐½ | Better quality than Q4, slightly larger. Great if you have the RAM. | | `gemma-4-12B-Q5_K_S.gguf` | 7.77 GB | ~9.3 GB | `Q5_K_S` | ⭐⭐⭐⭐ | Large but accurate. | | `gemma-4-12B-Q6_K.gguf` | 9.11 GB | ~10.6 GB | `Q6_K` | ⭐⭐⭐⭐⭐ | Near-perfect quality, very large. | | `gemma-4-12B-Q8_0.gguf` | 11.80 GB | ~13.3 GB | `Q8_0` | ⭐⭐⭐⭐⭐ | Closest to original quality. Use when RAM is not a concern. | ### 🖼️ Vision Encoder — mmproj (always required, always F16) | Filename | Size | Notes | |---|---|---| | `gemma-4-12B-mmproj-f16.gguf` | 0.11 GB | Always F16 — vision encoder is not quantized | > ⚠️ **You need BOTH files** — one text backbone + the mmproj — to run this VLM. --- ## ⚡ Speed Benchmarks *Run `python benchmark.py --model gemma-4-12B` to generate results.* --- ## 🚀 How to Use ### LM Studio (Easiest — GUI) 1. Search for `Dhptl/gemma-4-12B` in LM Studio 2. Download the Q4_K_M text file **and** the mmproj file 3. Load the model — LM Studio automatically uses both files ### Ollama ```bash ollama run dhptl/gemma-4-12b ``` ### llama.cpp CLI — Text + Image ```bash # Download both files to the same directory, then: ./llama-llava-cli \ -m gemma-4-12B-Q4_K_M.gguf \ --mmproj gemma-4-12B-mmproj-f16.gguf \ --image /path/to/your/image.jpg \ -p "Describe this image in detail." \ -n 512 ``` ### llama.cpp CLI — Text only (no image) ```bash ./llama-cli \ -m gemma-4-12B-Q4_K_M.gguf \ -p "You are a helpful assistant." \ --conversation ``` ### Python — llama-cpp-python ```python from llama_cpp import Llama from llama_cpp.llama_chat_format import Llava16ChatHandler # Load VLM with mmproj chat_handler = Llava16ChatHandler(clip_model_path="./gemma-4-12B-mmproj-f16.gguf") llm = Llama( model_path="./gemma-4-12B-Q4_K_M.gguf", chat_handler=chat_handler, n_gpu_layers=-1, n_ctx=4096, logits_all=True, ) # Text + image inference response = llm.create_chat_completion( messages=[ { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}}, {"type": "text", "text": "What do you see in this image?"} ] } ] ) print(response["choices"][0]["message"]["content"]) ``` --- ## 🔍 VLM Architecture This model uses a **two-component architecture**: | Component | File | Purpose | |---|---|---| | **Text Backbone** | `gemma-4-12B-Q4_K_M.gguf` | Language understanding & generation | | **Vision Encoder (mmproj)** | `gemma-4-12B-mmproj-f16.gguf` | Image feature extraction (always F16) | > **Why is mmproj always F16?** > The vision encoder maps image pixels to token embeddings. Quantizing it causes > visible visual artifacts and degraded image understanding. It stays at F16 (half precision) > which is already very efficient at ~1-2GB for most models. --- ## 🔍 About GGUF Quantization | Format | Bits/weight | Quality | |---|---|---| | Q3_K_M | ~3.3 | ⭐⭐⭐ | | Q4_K_M | ~4.5 | ⭐⭐⭐⭐ ← recommended | | Q5_K_M | ~5.6 | ⭐⭐⭐⭐½ | | Q8_0 | ~8.5 | ⭐⭐⭐⭐⭐ | --- ## 💬 Community & Feedback Found an issue? Open a **Discussion** in the Community tab. If useful, please: - ⭐ Star [quant-kit](https://github.com/DhruvalPtl/quant-kit) on GitHub - 👍 Like this model on HuggingFace