gergelyvagujhelyi commited on
Commit
e25e051
·
verified ·
1 Parent(s): de09bff

Add model card

Browse files
Files changed (1) hide show
  1. README.md +103 -0
README.md ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google/gemma-4-12B
4
+ tags:
5
+ - gguf
6
+ - nobodywho
7
+ - tool-calling
8
+ - vision
9
+ - gemma
10
+ pipeline_tag: image-text-to-text
11
+ library_name: gguf
12
+ ---
13
+
14
+ # NobodyWho/Google_Gemma4-12B-GGUF
15
+
16
+ ## Overview
17
+
18
+ GGUF quantization of Google's **Gemma 4 12B (Unified)** model, re-hosted for
19
+ [NobodyWho](https://github.com/nobodywho-ooo/nobodywho). The unsloth build already ships a
20
+ tool-calling setup and recommended sampling metadata (`general.sampling`: temp 1.0,
21
+ top_k 64, top_p 0.95), so nothing needs patching — the model is verified with NobodyWho's test
22
+ suite. The 12B Unified variant is the laptop-class Gemma 4 — stronger reasoning and multimodal
23
+ capability than the edge (E2B/E4B) models while staying well below the larger MoE/dense variants
24
+ in memory. Multimodal (text + image), multilingual, Apache 2.0.
25
+
26
+ ## Model Capabilities
27
+
28
+ - **Text generation** — instruction-following chat, stronger reasoning
29
+ - **Tool calling** — native function calling with grammar-constrained output
30
+ - **Vision** — ⚠️ the 12B `mmproj` (vision **+ audio** encoder) currently **fails to load in
31
+ NobodyWho** (llama.cpp MTMD/CLIP init error); needs a newer llama.cpp. Text + tool calling are
32
+ unaffected. For vision today, use Gemma 4 E2B/E4B (verified working)
33
+ - **Long context** — 256k tokens
34
+ - **Multilingual** — 140+ languages
35
+
36
+ ## Available Quantizations
37
+
38
+ | File | Approach | Tool-calling tests |
39
+ |------|----------|--------------------|
40
+ | `gemma-4-12b-it-BF16.gguf` | Sampling embedded upstream | not separately run |
41
+ | `gemma-4-12b-it-Q8_0.gguf` | Sampling embedded upstream | 14/14 |
42
+ | `gemma-4-12b-it-Q4_K_M.gguf` | Sampling embedded upstream | **14/14** |
43
+ | `mmproj-BF16.gguf` | Vision projection — ⚠️ does not load in NobodyWho yet | — |
44
+
45
+ > Tool calling verified on Q8_0 and Q4_K_M (14/14 each, June 2026; BF16 hosted but not separately tested — 24 GB).
46
+ > **Vision: the 12B `mmproj` fails to load** in the current NobodyWho build (llama.cpp MTMD/CLIP
47
+ > init error) — Gemma 4 E2B/E4B vision is verified working. Quant names follow the unsloth `gemma-4-12b-it-GGUF` repo.
48
+
49
+ ## Quick Start
50
+
51
+ Using the [NobodyWho](https://github.com/nobodywho-ooo/nobodywho) library:
52
+
53
+ ```python
54
+ from nobodywho import Chat
55
+
56
+ chat = Chat("huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf")
57
+ response = chat.ask("What is the capital of Denmark?").completed()
58
+ print(response) # The capital of Denmark is Copenhagen.
59
+ ```
60
+
61
+ ### Vision
62
+
63
+ > ⚠️ **Not working yet on 12B:** the `mmproj` fails to load in the current NobodyWho build. The
64
+ > snippet below is the intended API (it works for Gemma 4 E2B/E4B today).
65
+
66
+ ```python
67
+ from nobodywho import Model, Chat, Prompt, Image, Text
68
+
69
+ model = Model(
70
+ "huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf",
71
+ projection_model_path="huggingface:NobodyWho/Google_Gemma4-12B-GGUF/mmproj-BF16.gguf",
72
+ )
73
+ chat = Chat(model=model, system_prompt="You are a helpful assistant.")
74
+ response = chat.ask(Prompt([
75
+ Text("What is in this image?"),
76
+ Image("./photo.png"),
77
+ ])).completed()
78
+ print(response)
79
+ ```
80
+
81
+ ### llama-cpp-python
82
+
83
+ ```python
84
+ from llama_cpp import Llama
85
+
86
+ llm = Llama.from_pretrained(
87
+ repo_id="NobodyWho/Google_Gemma4-12B-GGUF",
88
+ filename="gemma-4-12b-it-Q4_K_M.gguf",
89
+ )
90
+ ```
91
+
92
+ ## Model Specifications
93
+
94
+ - **Parameters:** 12B (Unified)
95
+ - **Context length:** 262,144 tokens (256K)
96
+ - **License:** Apache 2.0
97
+ - **Base model:** google/gemma-4-12B
98
+ - **Architecture:** gemma4 (vision-capable)
99
+
100
+ ## Licensing / Credits
101
+
102
+ Licensed under Apache 2.0 (unchanged from upstream). All model credit belongs to Google
103
+ DeepMind. GGUF quantizations provided by unsloth.