AsbjornOlling's picture
Update README.md
b5fbdb5 verified
|
Raw
History Blame Contribute Delete
2.92 kB
metadata
license: apache-2.0
base_model: google/gemma-4-12B
tags:
  - gguf
  - nobodywho
  - tool-calling
  - vision
  - gemma
pipeline_tag: image-text-to-text
library_name: gguf

NobodyWho/Google_Gemma4-12B-GGUF

Overview

GGUF quantization of Google's Gemma 4 12B (Unified) model, re-hosted for NobodyWho. The unsloth build already ships a tool-calling setup and recommended sampling metadata (general.sampling: temp 1.0, top_k 64, top_p 0.95), so nothing needs patching — the model is verified with NobodyWho's test suite. The 12B Unified variant is the laptop-class Gemma 4 — stronger reasoning and multimodal capability than the edge (E2B/E4B) models while staying well below the larger MoE/dense variants in memory. Multimodal (text + image), multilingual, Apache 2.0.

Model Capabilities

  • Text generation — instruction-following chat, stronger reasoning
  • Tool calling — native function calling with grammar-constrained output
  • Vision — ⚠️ the 12B mmproj (vision + audio encoder) currently fails to load in NobodyWho (llama.cpp MTMD/CLIP init error); needs a newer llama.cpp. Text + tool calling are unaffected. For vision today, use Gemma 4 E2B/E4B (verified working)
  • Long context — 256k tokens
  • Multilingual — 140+ languages

Available Quantizations

File Approach Tool-calling tests
gemma-4-12b-it-BF16.gguf Sampling embedded upstream not separately run
gemma-4-12b-it-Q8_0.gguf Sampling embedded upstream 14/14
gemma-4-12b-it-Q4_K_M.gguf Sampling embedded upstream 14/14
mmproj-BF16.gguf Vision projection

Quick Start

Using the NobodyWho library:

from nobodywho import Chat

chat = Chat("huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf")
response = chat.ask("What is the capital of Denmark?").completed()
print(response)  # The capital of Denmark is Copenhagen.

Vision

from nobodywho import Model, Chat, Prompt, Image, Text

model = Model(
    "huggingface:NobodyWho/Google_Gemma4-12B-GGUF/gemma-4-12b-it-Q4_K_M.gguf",
    projection_model_path="huggingface:NobodyWho/Google_Gemma4-12B-GGUF/mmproj-BF16.gguf",
)
chat = Chat(model=model, system_prompt="You are a helpful assistant.")
response = chat.ask(Prompt([
    Text("What is in this image?"),
    Image("./photo.png"),
])).completed()
print(response)

Model Specifications

  • Parameters: 12B (Unified)
  • Context length: 262,144 tokens (256K)
  • License: Apache 2.0
  • Base model: google/gemma-4-12B
  • Architecture: gemma4 (vision-capable)

Licensing / Credits

Licensed under Apache 2.0 (unchanged from upstream). All model credit belongs to Google DeepMind. GGUF quantizations provided by unsloth.