Instructions to use bssgoat/local-typed-decisions with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bssgoat/local-typed-decisions with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bssgoat/local-typed-decisions:Q4_K_M # Run inference directly in the terminal: llama cli -hf bssgoat/local-typed-decisions:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bssgoat/local-typed-decisions:Q4_K_M # Run inference directly in the terminal: llama cli -hf bssgoat/local-typed-decisions:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bssgoat/local-typed-decisions:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf bssgoat/local-typed-decisions:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bssgoat/local-typed-decisions:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf bssgoat/local-typed-decisions:Q4_K_M
Use Docker
docker model run hf.co/bssgoat/local-typed-decisions:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use bssgoat/local-typed-decisions with Ollama:
ollama run hf.co/bssgoat/local-typed-decisions:Q4_K_M
- Unsloth Desktop
- Pi
How to use bssgoat/local-typed-decisions with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bssgoat/local-typed-decisions:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bssgoat/local-typed-decisions:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bssgoat/local-typed-decisions with Docker Model Runner:
docker model run hf.co/bssgoat/local-typed-decisions:Q4_K_M
- Lemonade
How to use bssgoat/local-typed-decisions with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bssgoat/local-typed-decisions:Q4_K_M
Run and chat with the model
lemonade run user.local-typed-decisions-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use bssgoat/local-typed-decisions with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bssgoat/local-typed-decisions:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bssgoat/local-typed-decisions:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bssgoat/local-typed-decisions with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bssgoat/local-typed-decisions:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bssgoat/local-typed-decisions:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
local-typed-decisions โ GGUF models
local-typed-decisions makes typed decisions (Choice, Score, Noul = yes/no) with a probability for every
allowed answer, locally on a consumer GPU (tested on an AMD Radeon RX 7600 8 GB with llama.cpp Vulkan) or on CPU.
The model never writes free text for the caller: the answer is always one of the requested options.
This repository holds the GGUF weights used by the runtime. Code, tests, methodology and all reports: https://github.com/GitLuci/local-typed-decisions.
The comparison baseline throughout is Jev 1.13 by TypeSafe, used as reference via its API. This project is not affiliated with TypeSafe.
Files
| file | size | used by modes |
|---|---|---|
qwen3-8b-gguf/qwen3-8b-Q8_0.gguf |
8.7 GB | fast, slow |
qwen3-4b-gguf/qwen3-4b-Q8_0.gguf |
4.3 GB | ultra-fast, medium |
qwen3-8b-gguf/qwen3-8b-Q4_K_M.gguf |
5.0 GB | extra (lighter; not used by the measured modes) |
qwen3-4b-gguf/qwen3-4b-Q4_K_M.gguf |
2.5 GB | extra (lighter; not used by the measured modes) |
SHA-256 checksums are in SHA256SUMS. Always verify them: the same size does not mean the same file.
What was modified, and what was not
- The weights are Qwen3's, with no training or fine-tuning. The only change is quantization with llama.cpp
(commit
9588757) without an importance matrix: Q8_0 (the one used) and Q4_K_M. Q8_0 was validated against the original weights in FP32: same fast-path answer on 54/54 cases; as a thinker 41/45 like FP32, with 43/45 agreement. Q4_K_M missed the agreement criterion by one case and is kept only as an extra. - Everything else lives in the runtime code:
- a fixed system prompt with STATE / QUESTION / OPTIONS and "answer with the letter only";
- options are shown as letters and the answer is read from the probabilities of the letter tokens
(llama-server
n_probs), renormalised over the options; - in thinking modes the model first writes a
<think>block with a bounded budget (1024 tokens for the 4B, 128 for the 8B), then the letter is read; - the
routedmode sends each question to the mode that did best on that kind of question.
Modes
Modes are named by speed; sorted by test-3 accuracy.
| mode | model | how it decides | test-3 /900 | time per question (RX 7600) |
|---|---|---|---|---|
| Jev 1.13 (API, baseline) | โ | โ | 841 | network |
routed |
4B and 8B | picks the mode by the question's domain (map below) | 829 (post hoc) | 4.5 s mean, 0.5 s median |
medium |
4B Q8 | thinks up to 1024 tokens, then reads the letter | 816 | 7.3 s |
slow |
8B Q8 | thinks up to 128 tokens, then reads the letter | 804 | 11.6 s |
fast |
8B Q8 | reads the letter directly, one forward pass | 782 | 0.48 s |
ultra-fast |
4B Q8 | reads the letter directly, one forward pass | 726 | 0.29 s |
Times are per question, each on its own, with no context reuse between questions on the same text (medians over the
test-3 run; routed measured on test-2 including model switches).
Routing map (chosen on test-3, confirmed on test-2): numeric, sentence โ medium; factual,
deterministic, sentiment โ slow; subjective_tone, robotic_style, noul_refund, score_urgency, random โ
fast.
Results
All tables are sorted from best to worst, with the Jev baseline placed by its own score.
Stress test (test-3): 900 sealed questions, gold independent of Jev
| correct | ฮ vs Jev (95 % CI) | per question | |
|---|---|---|---|
| Jev 1.13 (API, reference) | 841 (93.4 %) | โ | โ |
routed (combined: each domain in its best mode)ยน |
829 (92.1 %) | โ1.3 | ~5.3 s (mean) |
medium (4B thinking) |
816 (90.7 %) | โ2.8 [โ5.0; โ0.6] | 7.3 s |
slow (8B thinking โค 128) |
804 (89.3 %) | โ4.1 [โ6.2; โ2.1] | 11.6 s |
fast (8B) |
782 (86.9 %) | โ6.6 [โ8.9; โ4.3] | 0.48 s |
ultra-fast (4B) |
726 (80.7 %) | โ12.8 [โ15.6; โ10.0] | 0.29 s |
Per domain (correct out of 100):
| domain | Jev | routed (mode used) | medium | slow | fast | ultra-fast |
|---|---|---|---|---|---|---|
| numeric | 80 | 98 (medium) | 98 | 77 | 77 | 72 |
| factual | 100 | 96 (slow) | 93 | 96 | 91 | 91 |
| deterministic | 95 | 93 (slow) | 92 | 93 | 82 | 88 |
| sentence | 98 | 90 (medium) | 90 | 85 | 81 | 82 |
| sentiment | 90 | 88 (slow) | 85 | 88 | 87 | 86 |
| subjective_tone | 97 | 92 (fast) | 87 | 92 | 92 | 52 |
| robotic_style | 88 | 84 (fast) | 84 | 84 | 84 | 76 |
| noul_refund | 93 | 92 (fast) | 91 | 93 | 92 | 84 |
| score_urgency | 100 | 96 (fast) | 96 | 96 | 96 | 95 |
| total | 841 | 829 | 816 | 804 | 782 | 726 |
ยน routed on test-3 is computed from the answers of the four modes, with a map chosen by looking at these same
results, so it is optimistic. The real end-to-end measurement of routed was done on test-2 (below), where it ties
Jev.
- No mode reaches the pre-registered "Jev level" criterion (lower CI bound above โ3 points);
mediumis the closest. mediumbeats Jev on arithmetic: 98 vs 80, +18 [9.7; 27.0], Holm-corrected โ the only "beats" in the test.ultra-fastfails on tone (52/100); do not use it for tone.
test-2: 156 labeled questions
| correct | ฮ vs Jev (points) | per question | |
|---|---|---|---|
routed (measured end to end with the released runtime, GPU) |
145 | +0.6 | 4.5 s (mean) |
| Jev | 144 | โ | โ |
slow |
144 | 0.0 | 11.4 s |
medium |
141 | โ1.9 | 7.4 s |
fast |
132 | โ7.7 | 0.45 s |
ultra-fast |
120 | โ15.4 | 0.28 s |
Comparison with similar systems
Same sealed sets, same gold and scoring; sorted by test-3 accuracy. Other systems were run by the maintainers on CPU
(AMD Ryzen 5 5600X), pre-registered, with the Jev, medium and fast reference values reproduced in the same run.
| system | test-3 /900 | ฮ vs Jev [95 % CI] | test-2 /156 | median s/question | hardware |
|---|---|---|---|---|---|
| Jev 1.13 (API, baseline) | 841 | โ | 144 | network | remote |
routed |
829 (post hoc) | โ1.3 (no CI) | 145 (measured) | 4.5 mean | RX 7600 |
medium |
816 | โ2.8 [โ5.0; โ0.6] | 141 | 7.3 | RX 7600 |
slow |
804 | โ4.1 [โ6.2; โ2.1] | 144 | 11.6 | RX 7600 |
fast |
782 | โ6.6 [โ8.9; โ4.3] | 132 | 0.48 | RX 7600 |
ultra-fast |
726 | โ12.8 [โ15.6; โ10.0] | 123 | 0.29 | RX 7600 |
| Laya multilingual (third-party) | 462 | โ42.1 [โ45.9; โ38.1] | 86 | 0.13 | CPU |
| DeBERTa-v3-large zero-shot | 374 | โ51.9 [โ55.3; โ48.2] | 59 | 12.8 | CPU |
| BART-large-MNLI zero-shot | 358 | โ53.7 [โ57.1; โ50.1] | 57 | 0.95 | CPU |
| mDeBERTa-v3 zero-shot (multilingual) | 353 | โ54.2 [โ57.7; โ50.8] | 55 | 4.2 | CPU |
| Julia-1 (third-party) | 322 | โ57.7 [โ61.3; โ54.0] | 62 | 0.06 | CPU |
Every other system differs from Jev, fast and medium with McNemar p < 1e-60 on test-3, with no errors or
truncations. CPU and GPU latencies are not directly comparable. GLiClass was not included: same zero-shot
label-matching paradigm, no Score or Noul equivalent, and it needs a third-party library. Full aggregate results:
reports/comparison/results.md in the code repository.
How it was tested
- Pre-registration: every test had its plan and decision rule committed before running.
- Sealed cases: test-3 was sealed by hash before any model ran; Jev ran once on the sealed cases.
- Gold independent of Jev, from three sources: (a) code-generated items with computed answers, confirmed by a blind verifier (380/380); (b) texts from public datasets (Banking77, GoEmotions); (c) texts written by two authors, labeled blind by two independent labelers (97.4 % agreement, ฮบ 0.973), with an adjudicator for disagreements.
- Three difficulty levels (N1โN3), fixed before the models ran, with per-domain quotas.
- Statistics: paired 95 % CI (the more conservative of Newcombe and bootstrap), McNemar, Holm.
- Reproducibility: a clean-machine run from a fresh clone rebuilt the same GGUF byte for byte and reproduced the 4B fast-path test-2 result exactly.
Run
- Get llama.cpp b11205 (Vulkan build, or CPU).
- In the code repository:
python scripts/fetch_models.pydownloads the two Q8_0 files from this repository and checks their SHA-256 (while this repository is private, log in first withhf auth login). - Start the server:
python -m typed_decisions serve --config examples/llama-routed.json --port 8000(--mode ultra-fast|fast|medium|slowfixes one mode). Each model's server starts on its first request (11โ22 s); one model is resident at a time. - In
routedmode each question carries its domain in thedomainsfield. Full guide:docs/RUNTIME.md.
Real smoke on the RX 7600: 18/18 requests answered, no orphan processes after a normal or an abrupt exit. After start-up each question takes: ultra-fast 0.5 s, fast 0.6 s, medium 6โ9 s, slow 12 s.
Memory:
- 8B with 30 layers on the GPU (
-ngl 30, context 2048): ~7.2 GB VRAM. Use--load-mode none(b11205 has no--no-mmap): RAM stays at ~2.6 GB instead of ~10.4 GB. - 4B fully on the GPU (
-ngl 99): ~5.2 GB VRAM.
Caveats
- The gold for the authored texts (c) was validated by LLM labelers, without a human audit.
- Numbers come from llama.cpp b11205 on the GPU. With another engine (e.g. llama-cpp-python on CPU) ~3 % of the answers change.
- The routing map was chosen by looking at test-3 and confirmed on test-2; claiming Jev level needs a new sealed set.
- Probabilities are not calibrated; read confidence as a ranking, not a frequency.
License
Apache-2.0, as the original Qwen3 weights (Qwen/Qwen3-4B, Qwen/Qwen3-8B).
- Downloads last month
- 68
4-bit
8-bit



