Gemma 4 12B IT (QAT q4_0) β multimodal dual-mode llamafile
One self-contained executable. Runs on NVIDIA GPUs, Apple Silicon (M1βM5) and plain CPUs β and tunes itself to whichever it finds. Serves chat completions (text + image + audio input), embeddings, and a full web UI from one model instance on one port.
Source, patches, changelog: SEBK4C/Llamafile-gemma-4-12B-it-qat-q4_0-gguf-Inferance-And-embeddings (release v0.7.0). Built on a performance fork of mozilla-ai/llamafile (v0.10.7).
Download
One file: gemma4-server.llamafile (v0.7.0-universal). It detects its
platform at runtime and boots the right baked profile:
- NVIDIA β CUDA TinyBLAS DSO (driver-only) + CUDA-tuned defaults
- Apple Silicon β Metal-tuned defaults (
.args.xnu: ~20 tok/s on M1 Pro vs ~6 under the CUDA-tuned flags) + a prewarmed system prompt (the first message skips its prefill entirely) - CPU-only β portable fallback, any arch
Also inside regardless of platform: baked Qwen3 embeddings (1024-dim
/v1/embeddings), /v1/ingest, MTP drafter, web UI with voice controls
(LLAMAFILE_TTS_PORT / LLAMAFILE_EMBED_PORT hook up external sidecars).
Previous artifacts remain available in this repo's commit history.
v0.7.0 release notes
v0.6.0 + v0.6.1 β embeddings and ingest, baked in (2026-07-06)
One file now serves chat + retrieval-grade embeddings + document ingest:
- Baked embedding sidecar: Qwen3-Embedding-0.6B (Apache-2.0, 1024-dim) auto-spawns on startup, proxied at
/embed/v1/embeddings,/embed/health,/embed/tokenize. v0.6.1: the main/v1/embeddingsendpoint serves these vectors by default β every OpenAI client gets retrieval-grade 1024-dim embeddings with zero config (best use: documents bare, queries prefixedInstruct: <task>\nQuery:). The 12B's raw anisotropic output stays available via headerX-Raw-Embeddings: 1; opt out entirely withLLAMAFILE_NO_EMBED=1. POST /v1/ingest: text in β one grammar-constrained enrichment call (title, summary, entities, task domain, chunking hints) + a deterministic fidelity gate (ungrounded entities dropped β composed dates and mutated numbers never reach your index) + token-budgeted chunks + 1024-dim doc & chunk vectors, returned as oneingest.v1envelope for hybrid BM25+vector indexing.- All sidecars (voice + embeddings) are supervised and tied to the main server's lifetime β they self-reap even after
kill -9. - CUDA e2e (RTX 3080 Ti): 19-test probe PASS, 105.5 tok/s chat,
/v1/ingest1.97 s on GPU. Measured pipeline quality (all reproducible from bench/): retrieval hit@1 0.929 / MRR 0.955 over a multimodal corpus; labeled-receipt key-field recall 100% with 0% gate false-drops; NFCorpus nDCG@10 0.3626. Full details: release notes Β· INGEST_GUIDE.
Optimized defaults (v0.5.0)
This build ships with empirically-validated serving defaults (release notes):
- Sampler (server default β API and WebUI): temperature 1.0, top_k 64, top_p 0.95, min_p 0.01, DRY (0.8 / 1.75 / 2), repeat_penalty off β Google's official Gemma 4 recipe plus DRY anti-loop insurance. It never runs greedy (temperature 0), which is the confirmed degenerate-loop trigger on this model.
- System prompt (WebUI default only): a distilled Claude's-Constitution prompt (honest, calibrated, corrects false premises, non-sycophantic, follows real intent, treats you as a capable adult, not over-cautious) plus an explicit override-decline clause. In a powered A/B (6 jailbreak + 6 benign-edgy probes Γ 4 reps) the clause lifts jailbreak-decline from 0.75 to 1.00 at zero over-refusal cost, with no quality regression on the full battery. The web UI seeds it as the default system message; it is not injected into raw
/v1API requests.
Review the testing β these defaults were chosen from a benchmarked A/B study (Constitution+decline prompt wins or ties a bare prompt on accuracy, humanness, sophistication, and calibration, and declines persona/developer-mode/fiction jailbreaks 100% without over-refusing):
- Test data + charts: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data
- Protocol + full research log:
bench/RESEARCH_HISTORY.md, harnessbench/serve_bench.py, specbench/program.mdin the GitHub repo.
Quickstart β with the web UI
chmod +x gemma4-server.llamafile
./gemma4-server.llamafile
Then open http://127.0.0.1:8080/ in your browser. That's the whole install. The web UI supports markdown chat with visible reasoning, image and audio attachments (π records straight from your microphone), conversation management, and llama.cpp's tooling.
On first start the file detects your hardware and applies tuned defaults:
| Detected | Applied automatically |
|---|---|
| NVIDIA GPU (CUDA baked in, driver-only) | 131072 ctx, f16 KV, MTP speculative n=4 β ~90β200 tok/s on an RTX 3080 Ti |
| Apple Silicon (Metal) | 8192 ctx, MTP n=2 β the config validated on an M4 |
| CPU only | 8192 ctx, speculation off |
Override any flag to take control of it (./gemma4-server.llamafile -c 32768),
or LLAMAFILE_NO_AUTOTUNE=1 to disable tuning entirely. Launched it twice by
accident? The second copy tells you, and opens the already-running UI.
API
| Endpoint | What |
|---|---|
POST /v1/chat/completions |
OpenAI-style chat (+ SSE streaming, function calling); accepts text, image_url data URIs, input_audio (wav/mp3) |
POST /v1/messages + /v1/messages/count_tokens |
Anthropic Messages API, native β tools, thinking blocks, SSE; Claude Code connects with env vars alone |
POST /v1/responses |
OpenAI Responses API β reasoning + message output items, streaming |
POST /tts/v1/audio/speech |
built-in Kokoro text-to-speech (returns WAV; /tts/health = liveness) |
POST /v1/embeddings |
Retrieval-grade embeddings by default (v0.6.1) β transparently served by the baked Qwen3-Embedding-0.6B (1024-dim, Apache-2.0). Embed documents bare; prefix queries with Instruct: <task>\nQuery: (measured: β17% retrieval quality without it). Raw 12B pooled output (3840-dim, anisotropic, research-only): header X-Raw-Embeddings: 1 or LLAMAFILE_NO_EMBED=1 |
POST /v1/ingest |
text β retrieval-ready JSON (v0.6.0): one grammar-constrained enrichment call (title/summary/entities/task-domain/chunk hints) + deterministic fidelity gate (ungrounded entities dropped) + token-budgeted chunks + doc & chunk vectors, in one ingest.v1 envelope |
/embed/v1/embeddings, /embed/tokenize, /embed/health |
the embedding sidecar addressed directly (same vectors as /v1/embeddings default) |
GET /health, /props, POST /tokenize |
usual llama-server extras |
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"Why is the sky blue?"}],"max_tokens":256}'
What's inside the file
- Gemma 4 12B IT QAT-q4_0 weights + vision/audio projector + MTP drafter head
- CUDA backend (TinyBLAS, sm_75βsm_120; needs only the NVIDIA driver) with the upstream fattn fixes Gemma 4's 512-dim heads require, CUDA graphs on
- Metal + CPU backends (Cosmopolitan: macOS/Linux/BSD, arm64 + x86_64)
- Kokoro-82M text-to-speech,
built in β the file spawns its own read-aloud engine on startup
(supervised, auto-respawns) and proxies it at
/tts; no sidecar, no extra download - The llama.cpp web UI with voice/read-aloud extensions β karaoke word highlighting with click-to-jump and reading-speed controls β working out of the box on the built-in TTS; audio input needs nothing extra either
Useful flags
--clear-allβ wipe saved KV state and extraction caches, start fresh-c NΒβ context (up to 262144 with-ctk q8_0 -ctv q8_0on 12 GB GPUs)--port N,--host 0.0.0.0β serve beyond localhost- Windows can't run >4 GB executables β use the repo's
bin/llamafilewith external weights there: model + mmproj GGUFs from google/gemma-4-12B-it-qat-q4_0-gguf, and the MTP drafter GGUF from the v0.5.0 release assets (everything is also extractable from this llamafile:unzip gemma4-server.llamafile)
Test your hardware in one command
curl -fsSLO https://raw.githubusercontent.com/SEBK4C/Llamafile-gemma-4-12B-it-qat-q4_0-gguf-Inferance-And-embeddings/cuda-3080ti-optim/bench/api_probe.py
python3 api_probe.py --base http://127.0.0.1:8080
21 end-to-end tests β every endpoint above (embeddings default + /v1/ingest included) plus vision, audio-in and TTS β 18 PASS / 0 FAIL on the v0.6.1 build;
each with wall-clock speed on your machine. Stdlib-only, no installs. All
published runs, charts and the protocol live in the
test-data repo.
Works with coding agents (tested end-to-end)
β οΈ Gemma 4 12B is NOT a top coding model β it completes small, well-scoped agentic tasks at interactive speed, fully offline. Temper expectations accordingly and review what it produces.
| Harness | Surface | Verified |
|---|---|---|
| Claude Code | Anthropic /v1/messages (no adapter) |
multi-turn file+run tasks, 9.6β12.5 s |
| OpenCode | OpenAI chat completions + tools | same tasks, 8β12 s |
| OpenClaw | openai-completions provider |
chat + exec-tool turns, 7β11 s |
| Cline / Kilo Code | OpenAI-compatible (VS Code) | config verified at API level |
Setup guides with the exact verified configs: docs/integrations/.
Credits
mozilla-ai/llamafile Β· ggml-org/llama.cpp Β· ggml-org/llama-ui Β· hexgrad/Kokoro-82M Β· thewh1teagle/kokoro-onnx Β· google/gemma-4-12B-it-qat-q4_0-gguf
- Downloads last month
- 29
Model tree for SEBK4C/gemma-4-12b-it-qat-q4_0-llamafile
Base model
google/gemma-4-12B