Gemma 4 12B IT (QAT q4_0) β€” multimodal dual-mode llamafile

One self-contained executable. Runs on NVIDIA GPUs, Apple Silicon (M1–M5) and plain CPUs β€” and tunes itself to whichever it finds. Serves chat completions (text + image + audio input), embeddings, and a full web UI from one model instance on one port.

Source, patches, changelog: SEBK4C/Llamafile-gemma-4-12B-it-qat-q4_0-gguf-Inferance-And-embeddings (release v0.7.0). Built on a performance fork of mozilla-ai/llamafile (v0.10.7).

Download

One file: gemma4-server.llamafile (v0.7.0-universal). It detects its platform at runtime and boots the right baked profile:

  • NVIDIA β€” CUDA TinyBLAS DSO (driver-only) + CUDA-tuned defaults
  • Apple Silicon β€” Metal-tuned defaults (.args.xnu: ~20 tok/s on M1 Pro vs ~6 under the CUDA-tuned flags) + a prewarmed system prompt (the first message skips its prefill entirely)
  • CPU-only β€” portable fallback, any arch

Also inside regardless of platform: baked Qwen3 embeddings (1024-dim /v1/embeddings), /v1/ingest, MTP drafter, web UI with voice controls (LLAMAFILE_TTS_PORT / LLAMAFILE_EMBED_PORT hook up external sidecars). Previous artifacts remain available in this repo's commit history. v0.7.0 release notes

v0.6.0 + v0.6.1 β€” embeddings and ingest, baked in (2026-07-06)

One file now serves chat + retrieval-grade embeddings + document ingest:

  • Baked embedding sidecar: Qwen3-Embedding-0.6B (Apache-2.0, 1024-dim) auto-spawns on startup, proxied at /embed/v1/embeddings, /embed/health, /embed/tokenize. v0.6.1: the main /v1/embeddings endpoint serves these vectors by default β€” every OpenAI client gets retrieval-grade 1024-dim embeddings with zero config (best use: documents bare, queries prefixed Instruct: <task>\nQuery: ). The 12B's raw anisotropic output stays available via header X-Raw-Embeddings: 1; opt out entirely with LLAMAFILE_NO_EMBED=1.
  • POST /v1/ingest: text in β†’ one grammar-constrained enrichment call (title, summary, entities, task domain, chunking hints) + a deterministic fidelity gate (ungrounded entities dropped β€” composed dates and mutated numbers never reach your index) + token-budgeted chunks + 1024-dim doc & chunk vectors, returned as one ingest.v1 envelope for hybrid BM25+vector indexing.
  • All sidecars (voice + embeddings) are supervised and tied to the main server's lifetime β€” they self-reap even after kill -9.
  • CUDA e2e (RTX 3080 Ti): 19-test probe PASS, 105.5 tok/s chat, /v1/ingest 1.97 s on GPU. Measured pipeline quality (all reproducible from bench/): retrieval hit@1 0.929 / MRR 0.955 over a multimodal corpus; labeled-receipt key-field recall 100% with 0% gate false-drops; NFCorpus nDCG@10 0.3626. Full details: release notes Β· INGEST_GUIDE.

Optimized defaults (v0.5.0)

This build ships with empirically-validated serving defaults (release notes):

  • Sampler (server default β€” API and WebUI): temperature 1.0, top_k 64, top_p 0.95, min_p 0.01, DRY (0.8 / 1.75 / 2), repeat_penalty off β€” Google's official Gemma 4 recipe plus DRY anti-loop insurance. It never runs greedy (temperature 0), which is the confirmed degenerate-loop trigger on this model.
  • System prompt (WebUI default only): a distilled Claude's-Constitution prompt (honest, calibrated, corrects false premises, non-sycophantic, follows real intent, treats you as a capable adult, not over-cautious) plus an explicit override-decline clause. In a powered A/B (6 jailbreak + 6 benign-edgy probes Γ— 4 reps) the clause lifts jailbreak-decline from 0.75 to 1.00 at zero over-refusal cost, with no quality regression on the full battery. The web UI seeds it as the default system message; it is not injected into raw /v1 API requests.

Review the testing β€” these defaults were chosen from a benchmarked A/B study (Constitution+decline prompt wins or ties a bare prompt on accuracy, humanness, sophistication, and calibration, and declines persona/developer-mode/fiction jailbreaks 100% without over-refusing):

Quickstart β€” with the web UI

chmod +x gemma4-server.llamafile
./gemma4-server.llamafile

Then open http://127.0.0.1:8080/ in your browser. That's the whole install. The web UI supports markdown chat with visible reasoning, image and audio attachments (πŸŽ™ records straight from your microphone), conversation management, and llama.cpp's tooling.

On first start the file detects your hardware and applies tuned defaults:

Detected Applied automatically
NVIDIA GPU (CUDA baked in, driver-only) 131072 ctx, f16 KV, MTP speculative n=4 β€” ~90–200 tok/s on an RTX 3080 Ti
Apple Silicon (Metal) 8192 ctx, MTP n=2 β€” the config validated on an M4
CPU only 8192 ctx, speculation off

Override any flag to take control of it (./gemma4-server.llamafile -c 32768), or LLAMAFILE_NO_AUTOTUNE=1 to disable tuning entirely. Launched it twice by accident? The second copy tells you, and opens the already-running UI.

API

Endpoint What
POST /v1/chat/completions OpenAI-style chat (+ SSE streaming, function calling); accepts text, image_url data URIs, input_audio (wav/mp3)
POST /v1/messages + /v1/messages/count_tokens Anthropic Messages API, native β€” tools, thinking blocks, SSE; Claude Code connects with env vars alone
POST /v1/responses OpenAI Responses API β€” reasoning + message output items, streaming
POST /tts/v1/audio/speech built-in Kokoro text-to-speech (returns WAV; /tts/health = liveness)
POST /v1/embeddings Retrieval-grade embeddings by default (v0.6.1) β€” transparently served by the baked Qwen3-Embedding-0.6B (1024-dim, Apache-2.0). Embed documents bare; prefix queries with Instruct: <task>\nQuery: (measured: βˆ’17% retrieval quality without it). Raw 12B pooled output (3840-dim, anisotropic, research-only): header X-Raw-Embeddings: 1 or LLAMAFILE_NO_EMBED=1
POST /v1/ingest text β†’ retrieval-ready JSON (v0.6.0): one grammar-constrained enrichment call (title/summary/entities/task-domain/chunk hints) + deterministic fidelity gate (ungrounded entities dropped) + token-budgeted chunks + doc & chunk vectors, in one ingest.v1 envelope
/embed/v1/embeddings, /embed/tokenize, /embed/health the embedding sidecar addressed directly (same vectors as /v1/embeddings default)
GET /health, /props, POST /tokenize usual llama-server extras
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"Why is the sky blue?"}],"max_tokens":256}'

What's inside the file

  • Gemma 4 12B IT QAT-q4_0 weights + vision/audio projector + MTP drafter head
  • CUDA backend (TinyBLAS, sm_75–sm_120; needs only the NVIDIA driver) with the upstream fattn fixes Gemma 4's 512-dim heads require, CUDA graphs on
  • Metal + CPU backends (Cosmopolitan: macOS/Linux/BSD, arm64 + x86_64)
  • Kokoro-82M text-to-speech, built in β€” the file spawns its own read-aloud engine on startup (supervised, auto-respawns) and proxies it at /tts; no sidecar, no extra download
  • The llama.cpp web UI with voice/read-aloud extensions β€” karaoke word highlighting with click-to-jump and reading-speed controls β€” working out of the box on the built-in TTS; audio input needs nothing extra either

Useful flags

  • --clear-all β€” wipe saved KV state and extraction caches, start fresh
  • -c N Β­β€” context (up to 262144 with -ctk q8_0 -ctv q8_0 on 12 GB GPUs)
  • --port N, --host 0.0.0.0 β€” serve beyond localhost
  • Windows can't run >4 GB executables β€” use the repo's bin/llamafile with external weights there: model + mmproj GGUFs from google/gemma-4-12B-it-qat-q4_0-gguf, and the MTP drafter GGUF from the v0.5.0 release assets (everything is also extractable from this llamafile: unzip gemma4-server.llamafile)

Test your hardware in one command

curl -fsSLO https://raw.githubusercontent.com/SEBK4C/Llamafile-gemma-4-12B-it-qat-q4_0-gguf-Inferance-And-embeddings/cuda-3080ti-optim/bench/api_probe.py
python3 api_probe.py --base http://127.0.0.1:8080

21 end-to-end tests β€” every endpoint above (embeddings default + /v1/ingest included) plus vision, audio-in and TTS β€” 18 PASS / 0 FAIL on the v0.6.1 build; each with wall-clock speed on your machine. Stdlib-only, no installs. All published runs, charts and the protocol live in the test-data repo.

Works with coding agents (tested end-to-end)

⚠️ Gemma 4 12B is NOT a top coding model β€” it completes small, well-scoped agentic tasks at interactive speed, fully offline. Temper expectations accordingly and review what it produces.

Harness Surface Verified
Claude Code Anthropic /v1/messages (no adapter) multi-turn file+run tasks, 9.6–12.5 s
OpenCode OpenAI chat completions + tools same tasks, 8–12 s
OpenClaw openai-completions provider chat + exec-tool turns, 7–11 s
Cline / Kilo Code OpenAI-compatible (VS Code) config verified at API level

Setup guides with the exact verified configs: docs/integrations/.

Credits

mozilla-ai/llamafile Β· ggml-org/llama.cpp Β· ggml-org/llama-ui Β· hexgrad/Kokoro-82M Β· thewh1teagle/kokoro-onnx Β· google/gemma-4-12B-it-qat-q4_0-gguf

Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for SEBK4C/gemma-4-12b-it-qat-q4_0-llamafile