Gemma-4 12B Coder — SFT v5 (weights)

gemma-4 12B coder weights (safetensors) — for fine-tuning, merging, or quantizing.

Ready-to-serve GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF.

💡 Pick this for the best tool-calling (our gate winner). For an uncensored model, use SFT v5 + abliterated.

At a glance

Type Model weights (safetensors)
Techniques sft-qlora
Tool-calling ✅ 100% gate pass (recovery-shim path)
Status ✅ Active / supported
Use GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF

Use it — GGUF quantizations

Ready-to-serve GGUF quants live at tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF (llama.cpp / Ollama one-liners on that card). These are the safetensors weights, for fine-tuning / merging / quantizing.

Tool-calling gate

Served the GGUF on llama.cpp (llama-server --jinja), prompted 7 tool-use cases + 1 no-tool abstain, scored whether a structured tool call was emitted. raw = llama.cpp native parse; shim = same outputs re-parsed for gemma-4 native markup. Tools folded into the prompt at eval time, matching training.

The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.

Measured on Pass rate
this model — raw (--jinja) 0.125
this model — shim (prod path) 1.000

Intended use & limitations

Built for code generation and agentic tool use; serve locally via llama.cpp / Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or fabricated — validate tool arguments before executing, and keep a human in the loop for anything consequential.

Where this sits in the family


Provenance & reproduction

How this model was built — technique chain, training mix, and the exact knobs/pins, so the result is reproducible without any of our tooling.

Mechanics applied

Step Technique What it does Provenance
1 sft-qlora QLoRA supervised fine-tune to keep + improve native tool-calling

1. sft-qlora

  • tools_mode: mixed (xLAM schemas folded; conditional taught)

Training data & mix

Public sources; weights/row-caps are the exact balance.

Dataset Subset Role Weight Max rows
Agent-Ark/Toucan-1.5M Kimi-K2 tool-dense multiturn 2.0 1000
Nanbeige/ToolMind graph_syn_datasets/graphsyn.jsonl reasoning-heavy (down-weighted vs v4 — the v4 culprit) 1.0 500
NousResearch/hermes-function-calling-v1 func-calling.json multiturn function calling 1.5 600
NousResearch/hermes-function-calling-v1 func-calling-singleturn.json terse single-call (up vs v4) 2.0 600
Salesforce/xlam-function-calling-60k tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts) terse, verifiable single-call (up vs v4) 2.0 600

Pinned revisions (byte-exact reproduction):

  • Agent-Ark/Toucan-1.5M (Kimi-K2) @ 0df3cf37f2abefb380370cfb02eabea2a35ae782
  • Nanbeige/ToolMind (graph_syn_datasets/graphsyn.jsonl) @ 8020ed1c03c367e4eb720ac3828ab4b0b95d8baf
  • NousResearch/hermes-function-calling-v1 (func-calling.json) @ dae3e1d28cfbcf4b915c04ea1e072030529b4bda
  • NousResearch/hermes-function-calling-v1 (func-calling-singleturn.json) @ dae3e1d28cfbcf4b915c04ea1e072030529b4bda
  • Salesforce/xlam-function-calling-60k (tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts)) @ 26d14ebfe18b1f7b524bd39b404b50af5dc97866

Training hyperparameters

Knob Value
method QLoRA (4-bit NF4 base, bf16 compute)
lora_r / lora_alpha / dropout 32 / 32 / 0.05
target_modules all-linear
objective train on assistant turns only (responses-only masking)
optimizer adamw_8bit
lr / scheduler / warmup 2e-4 / cosine / 0.03
epochs 1
seq_len 4096
effective_batch 16
chat_template gemma (native turn boundaries 105/106)

Training environment

Exact pins the run trained against (the base arch needs a recent transformers).

Package Version
torch 2.11.0
transformers 5.13.0.dev0 @ c21da1b (git pin)
peft 0.19.1
trl 1.6.0
datasets 5.0.0
bitsandbytes 0.49.2
accelerate 1.14.0
liger-kernel 0.8.0
attention eager (no flash-attn)

Part of the Gemma-4 12B Coder — active collection.

Something not right, or a request? Open a discussion — happy to help.

Downloads last month
31
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5

Finetuned
(13)
this model
Finetunes
1 model
Quantizations
1 model

Datasets used to train tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5

Collection including tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5

Evaluation results

  • Tool-call pass rate (shim, prod path) on gemma4-coder-tool-eval
    self-reported
    1.000
  • Tool-call pass rate (raw llama.cpp --jinja) on gemma4-coder-tool-eval
    self-reported
    0.125