Gemma-4 12B Coder — SFT v5 (weights)
gemma-4 12B coder weights (safetensors) — for fine-tuning, merging, or quantizing.
Ready-to-serve GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF.
💡 Pick this for the best tool-calling (our gate winner). For an uncensored model, use SFT v5 + abliterated.
At a glance
| Type | Model weights (safetensors) |
| Techniques | sft-qlora |
| Tool-calling | ✅ 100% gate pass (recovery-shim path) |
| Status | ✅ Active / supported |
| Use | GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF |
Use it — GGUF quantizations
Ready-to-serve GGUF quants live at tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF
(llama.cpp / Ollama one-liners on that card). These are the safetensors weights, for
fine-tuning / merging / quantizing.
Tool-calling gate
Served the GGUF on llama.cpp (llama-server --jinja), prompted 7 tool-use cases + 1 no-tool abstain, scored whether a structured tool call was emitted. raw = llama.cpp native parse; shim = same outputs re-parsed for gemma-4 native markup. Tools folded into the prompt at eval time, matching training.
The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.
| Measured on | Pass rate |
|---|---|
this model — raw (--jinja) |
0.125 |
| this model — shim (prod path) | 1.000 |
Intended use & limitations
Built for code generation and agentic tool use; serve locally via llama.cpp / Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or fabricated — validate tool arguments before executing, and keep a human in the loop for anything consequential.
Where this sits in the family
- base (upstream) —
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
Provenance & reproduction
How this model was built — technique chain, training mix, and the exact knobs/pins, so the result is reproducible without any of our tooling.
Mechanics applied
| Step | Technique | What it does | Provenance |
|---|---|---|---|
| 1 | sft-qlora |
QLoRA supervised fine-tune to keep + improve native tool-calling | — |
1. sft-qlora
- tools_mode: mixed (xLAM schemas folded; conditional taught)
Training data & mix
Public sources; weights/row-caps are the exact balance.
| Dataset | Subset | Role | Weight | Max rows |
|---|---|---|---|---|
Agent-Ark/Toucan-1.5M |
Kimi-K2 | tool-dense multiturn | 2.0 | 1000 |
Nanbeige/ToolMind |
graph_syn_datasets/graphsyn.jsonl | reasoning-heavy (down-weighted vs v4 — the v4 culprit) | 1.0 | 500 |
NousResearch/hermes-function-calling-v1 |
func-calling.json | multiturn function calling | 1.5 | 600 |
NousResearch/hermes-function-calling-v1 |
func-calling-singleturn.json | terse single-call (up vs v4) | 2.0 | 600 |
Salesforce/xlam-function-calling-60k |
tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts) | terse, verifiable single-call (up vs v4) | 2.0 | 600 |
Pinned revisions (byte-exact reproduction):
Agent-Ark/Toucan-1.5M(Kimi-K2) @0df3cf37f2abefb380370cfb02eabea2a35ae782Nanbeige/ToolMind(graph_syn_datasets/graphsyn.jsonl) @8020ed1c03c367e4eb720ac3828ab4b0b95d8bafNousResearch/hermes-function-calling-v1(func-calling.json) @dae3e1d28cfbcf4b915c04ea1e072030529b4bdaNousResearch/hermes-function-calling-v1(func-calling-singleturn.json) @dae3e1d28cfbcf4b915c04ea1e072030529b4bdaSalesforce/xlam-function-calling-60k(tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts)) @26d14ebfe18b1f7b524bd39b404b50af5dc97866
Training hyperparameters
| Knob | Value |
|---|---|
| method | QLoRA (4-bit NF4 base, bf16 compute) |
| lora_r / lora_alpha / dropout | 32 / 32 / 0.05 |
| target_modules | all-linear |
| objective | train on assistant turns only (responses-only masking) |
| optimizer | adamw_8bit |
| lr / scheduler / warmup | 2e-4 / cosine / 0.03 |
| epochs | 1 |
| seq_len | 4096 |
| effective_batch | 16 |
| chat_template | gemma (native turn boundaries 105/106) |
Training environment
Exact pins the run trained against (the base arch needs a recent transformers).
| Package | Version |
|---|---|
| torch | 2.11.0 |
| transformers | 5.13.0.dev0 @ c21da1b (git pin) |
| peft | 0.19.1 |
| trl | 1.6.0 |
| datasets | 5.0.0 |
| bitsandbytes | 0.49.2 |
| accelerate | 1.14.0 |
| liger-kernel | 0.8.0 |
| attention | eager (no flash-attn) |
Part of the Gemma-4 12B Coder — active collection.
Something not right, or a request? Open a discussion — happy to help.
- Downloads last month
- 31
Model tree for tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5
Base model
google/gemma-4-12BDatasets used to train tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5
Salesforce/xlam-function-calling-60k
Agent-Ark/Toucan-1.5M
Collection including tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5
Evaluation results
- Tool-call pass rate (shim, prod path) on gemma4-coder-tool-evalself-reported1.000
- Tool-call pass rate (raw llama.cpp --jinja) on gemma4-coder-tool-evalself-reported0.125