Instructions to use jajmangold/gemma-4-12b-constraintkit-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K # Run inference directly in the terminal: llama cli -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K # Run inference directly in the terminal: llama cli -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K # Run inference directly in the terminal: ./llama-cli -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Use Docker
docker model run hf.co/jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
- LM Studio
- Jan
- Ollama
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with Ollama:
ollama run hf.co/jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
- Unsloth Desktop
- Pi
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with Docker Model Runner:
docker model run hf.co/jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
- Lemonade
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Run and chat with the model
lemonade run user.gemma-4-12b-constraintkit-GGUF-Q6_K
List all available models
lemonade list
- Hermes Agent
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jajmangold/gemma-4-12b-constraintkit-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jajmangold/gemma-4-12b-constraintkit-GGUF:Q6_K" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
gemma-4-12b-constraintkit
A 12B model that turns plain English into exact CAD, and knows when to say "I can't build that."
| 93% | 87% | 57% vs 16% | ~$35 |
|---|---|---|---|
| builds verified against real geometry (SFT stage, held-out) | honest declines on out-of-scope requests (SFT stage; GRPO: 9/10) | honest builds vs a prompted frontier model, fresh requests | total training cost, SFT + GRPO |
This is a Gemma 4 12B fine-tune for constraint-kit, an open-source text-to-CAD system where the AI decides what to build and deterministic math decides where it goes. The model writes a small program: parts, parameters, and mates. constraint-kit compiles that program into real B-rep solids, solves the assembly, and exports STEP. Nothing is a mesh the model hallucinated. Every build can be checked against geometry.
flowchart LR
A(["Plain English"]) --> B["gemma-4-12b-constraintkit<br/>reasons, then writes a program"]
B -->|out of scope| X(["unsupported: reason"])
B -->|parts, params, mates| C["constraint-kit<br/>exact B-rep generators<br/>+ deterministic mate kernel"]
C --> D(["STEP / GLB, BOM,<br/>mass, drawings"])
"3x 12mm OD, 5mm bore, 8mm height spacers stacked."
→ {"parts": [{"id": "s0", "type": "spacer", "params": {"outer_d": 12, "bore_d": 5, "height": 8}}, ... ×3],
"mates": [{"a": "s0", "a_joint": "top", "b": "s1", "b_joint": "bottom"}, ...]}
"Need a spiral bevel gear for a heavy-duty application."
→ {"unsupported": "only straight bevel gears are modeled"}
That second answer is the point. It knows the vocabulary well enough to say which variant it's missing, instead of quietly handing you a straight bevel gear and calling it done.
What the programs become
These are real constraint-kit builds, rendered straight from the exported geometry. Only materials and lighting were added. They show what constraint-kit makes from the kind of parts-params-mates program this model writes. To be precise, the quoted sentences went through constraint-kit's default hosted planner (DeepSeek), not this model, and the banner's planetary gearset was designed by the Z3 solver to exactly 4:1.
Why it's worth a look
- Scored against geometry, not vibes. A build counts only if the part's volume and bounding box land within 2% of the held-out target.
- 93% geometry-verified builds, 87% honest declines on held-out probes, first shot, with no retry loop and no prompt scaffolding. That's the same range as a prompted frontier model (90–97%) from a local 12B.
- On fresh requests it beat that frontier baseline on honesty: 57% honest builds vs 16%.
- GRPO with the CAD verifier as the reward. No reward model and no LLM judge: the reward is whether the program checks out and builds, and how close its geometric signature is. That halved silent wrong-variant builds (2/10 → 1/10) and raised correct declines to 9/10, with no regression on builds (6/6 held).
- Shows its reasoning. A thinking trace comes before every program, so you can see why it chose a part or declined.
- Small and local. A 9.8 GB Q6_K GGUF for llama.cpp, or a 251 MB LoRA adapter.
Files
| File | What it is |
|---|---|
gemma4-12b-constraintkit-grpo-Q6_K.gguf |
Final model (SFT + GRPO), Q6_K, 9.8 GB. For llama.cpp |
lora/ |
The same model as a PEFT LoRA adapter (r=16) on unsloth/gemma-4-12b-it |
train_sys.txt |
The exact system prompt used in training. Required, because the model was not trained on any other prompt |
Usage (llama.cpp)
llama-server -m gemma4-12b-constraintkit-grpo-Q6_K.gguf -c 8192 -ngl 99 --port 8086
# On Volta (V100-class) GPUs add: -fa off -ub 1024 (flash-attn auto = very slow prefill there)
Use raw /completion with the training serialization. The system turn starts with the <|think|> mode token,
and generation stops at <turn|>:
import json, urllib.request
SYS = open("train_sys.txt").read()
def generate(request: str) -> str:
prompt = (f"<|turn>system\n<|think|>\n{SYS}<turn|>\n"
f"<|turn>user\n{request}<turn|>\n<|turn>model\n")
body = {"prompt": prompt, "n_predict": 1400, "temperature": 0.2, "top_p": 0.95, "stop": ["<turn|>"]}
req = urllib.request.Request("http://127.0.0.1:8086/completion", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
text = json.loads(urllib.request.urlopen(req).read())["content"]
return text.split("<channel|>")[-1].strip() # drop the reasoning trace, keep the JSON program
print(generate("a plate with a spacer and a 24-tooth module-1 gear stacked on its boss"))
The output is a DSL program ({"parts": [...], "mates": [...]}) or {"unsupported": "..."}. To check it and
build it, send it to a running constraint-kit service (POST /dsl/verify, POST /build).
Results
| Stage | Builds (geometry-verified) | Honest declines | Notes |
|---|---|---|---|
| SFT, bf16 | 56/60 (93%) | 26/30 (87%) | 0 format failures, 0 build errors |
| SFT, Q6_K GGUF | 55/60 (92%) | 25/30 (83%) | Same 90 probes. Quantization costs about 1–3 points, mostly in declines |
| SFT + GRPO (this model) | 6/6 held | 9/10 (from 8/10) | 16-probe A/B. Silent wrong-variant builds halved (2/10 → 1/10) |
Training
flowchart LR
prompt(["Request"]) --> policy["Gemma 4 12B<br/>LoRA policy"]
policy --> prog["Reasoning + program<br/>or decline"]
prog --> verifier["constraint-kit verifier<br/>checks pass? builds?<br/>geometry signature match?"]
verifier -->|reward| policy
- SFT: LoRA r=16, 2 epochs, about 40k rows made by constraint-kit's own data factory: natural language → reasoning → DSL program, honest declines, and self-correction traces driven by diagnostics (bad attempts loss-masked). About 10.5 h on one RTX PRO 6000, roughly $7.
- GRPO: 150 steps on a signal-dense mix (70% decline / 30% hard builds). The reward is the constraint-kit DSL verifier used as a library: graded check-pass, builds, signature proximity, and an exact-match bonus.
- The whole thing cost about $35 in rented GPU time. Scripts are in
training/.
A negative result, reported honestly: a multimodal image→CAD LoRA was also trained and thrown away. With images verified as fed to the model, it scored 52% with the render vs 47% without, which is noise. The model ignored the picture and relied on the bounding box. The results are written up in the repo. The weights aren't shipped.
Limitations
- It only knows the vocabulary in
train_sys.txt(spacer, spur_gear, plate, shaft, washer, nut, bolt, panel, link, housing, sheet_bracket, and a few library "vitamins"). Anything else should be declined. - Declines are the weak spot: about 1 in 10 unsupported requests still gets built as a near match.
- Run the output through constraint-kit's verifier (
POST /dsl/verify) before trusting the geometry.
License
Apache-2.0, the same as the Gemma 4 base model. constraint-kit itself is MIT.
- Downloads last month
- 244
6-bit






