gemma-4-12b-constraintkit

A 12B model that turns plain English into exact CAD, and knows when to say "I can't build that."

Assemblies built by constraint-kit: a planetary gearset the Z3 solver designed to exactly 4:1, a three-gear train on a plate, a gear seated on a plate boss, a helical gear

93% 87% 57% vs 16% ~$35
builds verified against real geometry (SFT stage, held-out) honest declines on out-of-scope requests (SFT stage; GRPO: 9/10) honest builds vs a prompted frontier model, fresh requests total training cost, SFT + GRPO

This is a Gemma 4 12B fine-tune for constraint-kit, an open-source text-to-CAD system where the AI decides what to build and deterministic math decides where it goes. The model writes a small program: parts, parameters, and mates. constraint-kit compiles that program into real B-rep solids, solves the assembly, and exports STEP. Nothing is a mesh the model hallucinated. Every build can be checked against geometry.

flowchart LR
    A(["Plain English"]) --> B["gemma-4-12b-constraintkit<br/>reasons, then writes a program"]
    B -->|out of scope| X(["unsupported: reason"])
    B -->|parts, params, mates| C["constraint-kit<br/>exact B-rep generators<br/>+ deterministic mate kernel"]
    C --> D(["STEP / GLB, BOM,<br/>mass, drawings"])
"3x 12mm OD, 5mm bore, 8mm height spacers stacked."
→ {"parts": [{"id": "s0", "type": "spacer", "params": {"outer_d": 12, "bore_d": 5, "height": 8}}, ... ×3],
   "mates": [{"a": "s0", "a_joint": "top", "b": "s1", "b_joint": "bottom"}, ...]}

"Need a spiral bevel gear for a heavy-duty application."
→ {"unsupported": "only straight bevel gears are modeled"}

That second answer is the point. It knows the vocabulary well enough to say which variant it's missing, instead of quietly handing you a straight bevel gear and calling it done.

What the programs become

These are real constraint-kit builds, rendered straight from the exported geometry. Only materials and lighting were added. They show what constraint-kit makes from the kind of parts-params-mates program this model writes. To be precise, the quoted sentences went through constraint-kit's default hosted planner (DeepSeek), not this model, and the banner's planetary gearset was designed by the Z3 solver to exactly 4:1.

Bearing block on a V-slot extrusion M8 threaded rod with nuts and washers Four-bar linkage
"a pillow bearing block bolted to the side of a 100 mm 20x20 aluminum V-slot extrusion with two M5 socket head screws, a 608 ball bearing pressed into the block's bore, and a 60 mm long 8 mm steel shaft through the bearing" "an M8 threaded rod 80 mm long with a steel washer and hex nut near each end". The thread is real helical geometry Four-bar linkage from link lengths 100 / 35 / 90 / 70 mm, closed by the SolveSpace solver
NEMA stepper motor Flanged pulley Meshing spur gears
NEMA stepper, placed as a BOSL2 library "vitamin" Flanged pulley, added after the census found it among the most-requested missing parts Spur pair at the exact center distance, m·(z₁+z₂)/2

Why it's worth a look

  • Scored against geometry, not vibes. A build counts only if the part's volume and bounding box land within 2% of the held-out target.
  • 93% geometry-verified builds, 87% honest declines on held-out probes, first shot, with no retry loop and no prompt scaffolding. That's the same range as a prompted frontier model (90–97%) from a local 12B.
  • On fresh requests it beat that frontier baseline on honesty: 57% honest builds vs 16%.
  • GRPO with the CAD verifier as the reward. No reward model and no LLM judge: the reward is whether the program checks out and builds, and how close its geometric signature is. That halved silent wrong-variant builds (2/10 → 1/10) and raised correct declines to 9/10, with no regression on builds (6/6 held).
  • Shows its reasoning. A thinking trace comes before every program, so you can see why it chose a part or declined.
  • Small and local. A 9.8 GB Q6_K GGUF for llama.cpp, or a 251 MB LoRA adapter.

Files

File What it is
gemma4-12b-constraintkit-grpo-Q6_K.gguf Final model (SFT + GRPO), Q6_K, 9.8 GB. For llama.cpp
lora/ The same model as a PEFT LoRA adapter (r=16) on unsloth/gemma-4-12b-it
train_sys.txt The exact system prompt used in training. Required, because the model was not trained on any other prompt

Usage (llama.cpp)

llama-server -m gemma4-12b-constraintkit-grpo-Q6_K.gguf -c 8192 -ngl 99 --port 8086
# On Volta (V100-class) GPUs add:  -fa off -ub 1024   (flash-attn auto = very slow prefill there)

Use raw /completion with the training serialization. The system turn starts with the <|think|> mode token, and generation stops at <turn|>:

import json, urllib.request

SYS = open("train_sys.txt").read()

def generate(request: str) -> str:
    prompt = (f"<|turn>system\n<|think|>\n{SYS}<turn|>\n"
              f"<|turn>user\n{request}<turn|>\n<|turn>model\n")
    body = {"prompt": prompt, "n_predict": 1400, "temperature": 0.2, "top_p": 0.95, "stop": ["<turn|>"]}
    req = urllib.request.Request("http://127.0.0.1:8086/completion", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    text = json.loads(urllib.request.urlopen(req).read())["content"]
    return text.split("<channel|>")[-1].strip()   # drop the reasoning trace, keep the JSON program

print(generate("a plate with a spacer and a 24-tooth module-1 gear stacked on its boss"))

The output is a DSL program ({"parts": [...], "mates": [...]}) or {"unsupported": "..."}. To check it and build it, send it to a running constraint-kit service (POST /dsl/verify, POST /build).

Results

Stage Builds (geometry-verified) Honest declines Notes
SFT, bf16 56/60 (93%) 26/30 (87%) 0 format failures, 0 build errors
SFT, Q6_K GGUF 55/60 (92%) 25/30 (83%) Same 90 probes. Quantization costs about 1–3 points, mostly in declines
SFT + GRPO (this model) 6/6 held 9/10 (from 8/10) 16-probe A/B. Silent wrong-variant builds halved (2/10 → 1/10)

Training

flowchart LR
    prompt(["Request"]) --> policy["Gemma 4 12B<br/>LoRA policy"]
    policy --> prog["Reasoning + program<br/>or decline"]
    prog --> verifier["constraint-kit verifier<br/>checks pass? builds?<br/>geometry signature match?"]
    verifier -->|reward| policy
  • SFT: LoRA r=16, 2 epochs, about 40k rows made by constraint-kit's own data factory: natural language → reasoning → DSL program, honest declines, and self-correction traces driven by diagnostics (bad attempts loss-masked). About 10.5 h on one RTX PRO 6000, roughly $7.
  • GRPO: 150 steps on a signal-dense mix (70% decline / 30% hard builds). The reward is the constraint-kit DSL verifier used as a library: graded check-pass, builds, signature proximity, and an exact-match bonus.
  • The whole thing cost about $35 in rented GPU time. Scripts are in training/.

A negative result, reported honestly: a multimodal image→CAD LoRA was also trained and thrown away. With images verified as fed to the model, it scored 52% with the render vs 47% without, which is noise. The model ignored the picture and relied on the bounding box. The results are written up in the repo. The weights aren't shipped.

Limitations

  • It only knows the vocabulary in train_sys.txt (spacer, spur_gear, plate, shaft, washer, nut, bolt, panel, link, housing, sheet_bracket, and a few library "vitamins"). Anything else should be declined.
  • Declines are the weak spot: about 1 in 10 unsupported requests still gets built as a near match.
  • Run the output through constraint-kit's verifier (POST /dsl/verify) before trusting the geometry.

License

Apache-2.0, the same as the Gemma 4 base model. constraint-kit itself is MIT.

Downloads last month
244
GGUF
Model size
12B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jajmangold/gemma-4-12b-constraintkit-GGUF

Adapter
(108)
this model