File size: 2,917 Bytes
e25e051
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5fbdb5
e25e051
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
license: apache-2.0
base_model: google/gemma-4-12B
tags:
  - gguf
  - nobodywho
  - tool-calling
  - vision
  - gemma
pipeline_tag: image-text-to-text
library_name: gguf
---

# NobodyWho/Google_Gemma4-12B-GGUF

## Overview

GGUF quantization of Google's **Gemma 4 12B (Unified)** model, re-hosted for
[NobodyWho](https://github.com/nobodywho-ooo/nobodywho). The unsloth build already ships a
tool-calling setup and recommended sampling metadata (`general.sampling`: temp 1.0,
top_k 64, top_p 0.95), so nothing needs patching — the model is verified with NobodyWho's test
suite. The 12B Unified variant is the laptop-class Gemma 4 — stronger reasoning and multimodal
capability than the edge (E2B/E4B) models while staying well below the larger MoE/dense variants
in memory. Multimodal (text + image), multilingual, Apache 2.0.

## Model Capabilities

- **Text generation** — instruction-following chat, stronger reasoning
- **Tool calling** — native function calling with grammar-constrained output
- **Vision** — ⚠️ the 12B `mmproj` (vision **+ audio** encoder) currently **fails to load in
  NobodyWho** (llama.cpp MTMD/CLIP init error); needs a newer llama.cpp. Text + tool calling are
  unaffected. For vision today, use Gemma 4 E2B/E4B (verified working)
- **Long context** — 256k tokens
- **Multilingual** — 140+ languages

## Available Quantizations

| File | Approach | Tool-calling tests |
|------|----------|--------------------|
| `gemma-4-12b-it-BF16.gguf` | Sampling embedded upstream | not separately run |
| `gemma-4-12b-it-Q8_0.gguf` | Sampling embedded upstream | 14/14 |
| `gemma-4-12b-it-Q4_K_M.gguf` | Sampling embedded upstream | **14/14** |
| `mmproj-BF16.gguf` | Vision projection | — |


## Quick Start

Using the [NobodyWho](https://github.com/nobodywho-ooo/nobodywho) library:

```python
from nobodywho import Chat

chat = Chat("huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf")
response = chat.ask("What is the capital of Denmark?").completed()
print(response)  # The capital of Denmark is Copenhagen.
```

### Vision

```python
from nobodywho import Model, Chat, Prompt, Image, Text

model = Model(
    "huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf",
    projection_model_path="huggingface:NobodyWho/Google_Gemma4-12B-GGUF/mmproj-BF16.gguf",
)
chat = Chat(model=model, system_prompt="You are a helpful assistant.")
response = chat.ask(Prompt([
    Text("What is in this image?"),
    Image("./photo.png"),
])).completed()
print(response)
```

## Model Specifications

- **Parameters:** 12B (Unified)
- **Context length:** 262,144 tokens (256K)
- **License:** Apache 2.0
- **Base model:** google/gemma-4-12B
- **Architecture:** gemma4 (vision-capable)

## Licensing / Credits

Licensed under Apache 2.0 (unchanged from upstream). All model credit belongs to Google
DeepMind. GGUF quantizations provided by unsloth.