Instructions to use SEBK4C/Winnow-12B-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use SEBK4C/Winnow-12B-LoRA with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it") model = PeftModel.from_pretrained(base_model, "SEBK4C/Winnow-12B-LoRA") - Notebooks
- Google Colab
- Kaggle
Winnow-12B LoRA (extracted)
The EldanRing/Winnow-12B typed-decision fine-tune, recovered as a rank-32 LoRA adapter for stock google/gemma-4-12B-it from the merged release. Serve one Gemma 4 12B and apply Winnow per request: plain Gemma for chat (text, images, audio), Winnow for Jev-style decisions, from the same loaded weights.
Winnow was published only as merged weights. This adapter was extracted from those released weights; it is not EldanRing's original training artifact. All credit for the fine-tune goes to EldanRing; to Google DeepMind for Gemma 4. Apache-2.0, like both.
Files
| file | what |
|---|---|
adapter_model.safetensors, adapter_config.json |
PEFT adapter, plain-SVD extraction (recommended), r=32, Ξ±=32 (scale 1), 328 modules |
rounding-aware/ |
the same, rounding-aware extraction (99.1 % of weights re-merge bit-exactly; see below) |
gguf/winnow-12b-lora-r32-svd-F16.gguf |
llama.cpp LoRA (262 MB), recommended |
gguf/winnow-12b-lora-r32-interval-F16.gguf |
llama.cpp LoRA, rounding-aware |
gguf/mmproj-gemma-4-12b-it-F32.gguf |
vision + audio projector for the base, F32 (an F16 v.patch_embd breaks vision) |
extraction/ |
extract_interval.py (both methods), per-matrix statistics and spectra (extraction-r32.json), figures |
eval/ |
benchmark summaries and per-decision agreement |
serving/ |
the server patch (apply_gjh.py for winnow-inference), build script, systemd unit, proxy and gateway snippets |
REPORT.md |
the full write-up |
Use
PEFT / transformers
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel
base = AutoModelForImageTextToText.from_pretrained("google/gemma-4-12B-it", revision="707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "SEBK4C/Winnow-12B-LoRA")
llama.cpp, adapter per request
llama-server -m gemma-4-12b-it-Q5_K_M.gguf --mmproj mmproj-gemma-4-12b-it-F32.gguf \
--lora winnow-12b-lora-r32-svd-F16.gguf --lora-init-without-apply
# per request: "lora": [{"id": 0, "scale": 1}] -> Winnow; omitted -> plain Gemma 4
Typed decisions (/v1/systemone) with the adapter: Winnow's own server
(winnow-inference) disables its decision endpoint when a LoRA is
loaded. serving/apply_gjh.py keeps it on, applies the adapter to the decision context, picks the adapter by model
name (winnow-12b β on, gemma-4-12b β off, on chat and decisions alike) and adds winnow.audio next to
winnow.images.
How it was extracted
349 of Gemma 4's 677 tensors are bit-identical in Winnow; everything that changed is in the 328 LoRA targets
(q/k/v/o/gate/up/down Γ 48 layers, no v_proj on the 8 global layers). But only 27 % of those weights changed at
all: most single-weight LoRA updates are smaller than a BF16 step and were rounded away when the adapter was
merged. Winnow β base is therefore the LoRA plus rounding noise of the same size.
- Plain SVD of the difference, rank 32: 82.5 % of weights re-merge bit-exactly.
- Rounding-aware: every released weight pins the true update to the interval of values that round to it; alternating projections between that box and the rank-32 set (clip, then truncated SVD) reach 99.1 % (rank 16 stalls at 82.6 %, which confirms rank 32).
Results
Near-lossless (Q8_0, one RTX 4090), Winnow's pinned suites, per-decision agreement with the merged release:
| Q8_0 | JevBench public | Kev-v9 clean | typed | Kev-v9 ECE | same decision as merged Winnow (JevBench / Kev / typed) | mean TV |
|---|---|---|---|---|---|---|
| stock Gemma 4 12B | 83.55 % | 78.20 % | 71.70 % | 0.170 | 91.3 / 88.8 / 85.6 % | 0.10β0.19 |
| merged Winnow (reference) | 86.15 % | 81.36 % | 70.25 % | 0.098 | β | β |
| this adapter (SVD) | 85.71 % | 81.26 % | 69.95 % | 0.096 | 98.7 / 99.1 / 97.9 % | 0.009β0.017 |
| rounding-aware variant | 86.15 % | 81.45 % | 69.10 % | 0.084 | 98.3 / 97.3 / 96.4 % | 0.023β0.045 |
| Winnow Q8, published | 85.71 % | 81.55 % | 70.00 % |
The SVD adapter reproduces Winnow's published Q8 scores and 98β99 % of its individual decisions. The rounding-aware variant matches 99 % of the weights bit-exactly yet behaves 2β3Γ further from Winnow: bit-exactness is the wrong objective (rounding noise averages out in the activations; a box-consistent rank-32 solution's error does not).
Deployed on a 12 GB RTX 3080 Ti (base Q5_K_M): SVD adapter 83.98 / 80.78 / 69.75 %, 96.1β96.5 % decision agreement with merged Winnow at the same precision; ~30 ms added per decision request for applying the LoRA at runtime.
What Winnow adds over stock Gemma is mostly calibration (Kev-v9 ECE 0.170 β 0.098, JevBench Brier 0.275 β 0.205)
and 3 points on Kev-v9. On 2,000 real Jev requests from our coding-agent harness, the adapter leaves the share of
decisions that match hosted Jev unchanged (90 % yes/no) but brings its probabilities much closer to Jev's
(mean TV 0.20 β 0.15).
In a coding-agent harness (bonsai-harness, Bonsai 2 27B as the coder)
Same seven text tasks, one run per cell; decisions via hosted Jev or this adapter on a 3080 Ti (the other model shadowing every decision on the identical state):
| Jev, standard | adapter, standard | Jev, menu mode | adapter, menu mode | |
|---|---|---|---|---|
| checks passed (of 115) | 59 | 59 | 73 | 85 |
| wall time | 96 min | 157 min | 179 min | 150 min |
Same score in the standard harness at 1.6Γ the time (one consumer GPU vs a hosted API: p50 0.9β2.8 s vs 0.25 s per
decision under four parallel agents); in menu mode, where the decider picks every step, the local adapter scored
higher. 93 % of yes/no gate decisions and ~70 % of menu picks agree with Jev. n = 1 per cell; details in REPORT.md.
Limits
- Extracted, not original: ~1β2 % of decisions differ from the merged model at Q8, ~4 % at Q5 (where Q5 noise dominates).
- Runtime LoRA costs ~30 ms per decision request on a 3080 Ti.
- Audio decisions work on real speech (LibriSpeech fixture: "Mister Quilter" p = 0.9999); synthetic TTS voices are the weak spot. Winnow's report does not evaluate audio, and neither did its training.
- Downloads last month
- 63
16-bit
