local-typed-decisions โ€” GGUF models

local-typed-decisions makes typed decisions (Choice, Score, Noul = yes/no) with a probability for every allowed answer, locally on a consumer GPU (tested on an AMD Radeon RX 7600 8 GB with llama.cpp Vulkan) or on CPU. The model never writes free text for the caller: the answer is always one of the requested options.

This repository holds the GGUF weights used by the runtime. Code, tests, methodology and all reports: https://github.com/GitLuci/local-typed-decisions.

The comparison baseline throughout is Jev 1.13 by TypeSafe, used as reference via its API. This project is not affiliated with TypeSafe.

Files

file size used by modes
qwen3-8b-gguf/qwen3-8b-Q8_0.gguf 8.7 GB fast, slow
qwen3-4b-gguf/qwen3-4b-Q8_0.gguf 4.3 GB ultra-fast, medium
qwen3-8b-gguf/qwen3-8b-Q4_K_M.gguf 5.0 GB extra (lighter; not used by the measured modes)
qwen3-4b-gguf/qwen3-4b-Q4_K_M.gguf 2.5 GB extra (lighter; not used by the measured modes)

SHA-256 checksums are in SHA256SUMS. Always verify them: the same size does not mean the same file.

What was modified, and what was not

  • The weights are Qwen3's, with no training or fine-tuning. The only change is quantization with llama.cpp (commit 9588757) without an importance matrix: Q8_0 (the one used) and Q4_K_M. Q8_0 was validated against the original weights in FP32: same fast-path answer on 54/54 cases; as a thinker 41/45 like FP32, with 43/45 agreement. Q4_K_M missed the agreement criterion by one case and is kept only as an extra.
  • Everything else lives in the runtime code:
    • a fixed system prompt with STATE / QUESTION / OPTIONS and "answer with the letter only";
    • options are shown as letters and the answer is read from the probabilities of the letter tokens (llama-server n_probs), renormalised over the options;
    • in thinking modes the model first writes a <think> block with a bounded budget (1024 tokens for the 4B, 128 for the 8B), then the letter is read;
    • the routed mode sends each question to the mode that did best on that kind of question.

Modes

Modes are named by speed; sorted by test-3 accuracy.

mode model how it decides test-3 /900 time per question (RX 7600)
Jev 1.13 (API, baseline) โ€” โ€” 841 network
routed 4B and 8B picks the mode by the question's domain (map below) 829 (post hoc) 4.5 s mean, 0.5 s median
medium 4B Q8 thinks up to 1024 tokens, then reads the letter 816 7.3 s
slow 8B Q8 thinks up to 128 tokens, then reads the letter 804 11.6 s
fast 8B Q8 reads the letter directly, one forward pass 782 0.48 s
ultra-fast 4B Q8 reads the letter directly, one forward pass 726 0.29 s

Times are per question, each on its own, with no context reuse between questions on the same text (medians over the test-3 run; routed measured on test-2 including model switches).

Routing map (chosen on test-3, confirmed on test-2): numeric, sentence โ†’ medium; factual, deterministic, sentiment โ†’ slow; subjective_tone, robotic_style, noul_refund, score_urgency, random โ†’ fast.

Results

All tables are sorted from best to worst, with the Jev baseline placed by its own score.

Stress test (test-3): 900 sealed questions, gold independent of Jev

test-3 accuracy

accuracy vs latency

correct ฮ” vs Jev (95 % CI) per question
Jev 1.13 (API, reference) 841 (93.4 %) โ€” โ€”
routed (combined: each domain in its best mode)ยน 829 (92.1 %) โˆ’1.3 ~5.3 s (mean)
medium (4B thinking) 816 (90.7 %) โˆ’2.8 [โˆ’5.0; โˆ’0.6] 7.3 s
slow (8B thinking โ‰ค 128) 804 (89.3 %) โˆ’4.1 [โˆ’6.2; โˆ’2.1] 11.6 s
fast (8B) 782 (86.9 %) โˆ’6.6 [โˆ’8.9; โˆ’4.3] 0.48 s
ultra-fast (4B) 726 (80.7 %) โˆ’12.8 [โˆ’15.6; โˆ’10.0] 0.29 s

Per domain (correct out of 100):

test-3 per domain

domain Jev routed (mode used) medium slow fast ultra-fast
numeric 80 98 (medium) 98 77 77 72
factual 100 96 (slow) 93 96 91 91
deterministic 95 93 (slow) 92 93 82 88
sentence 98 90 (medium) 90 85 81 82
sentiment 90 88 (slow) 85 88 87 86
subjective_tone 97 92 (fast) 87 92 92 52
robotic_style 88 84 (fast) 84 84 84 76
noul_refund 93 92 (fast) 91 93 92 84
score_urgency 100 96 (fast) 96 96 96 95
total 841 829 816 804 782 726

ยน routed on test-3 is computed from the answers of the four modes, with a map chosen by looking at these same results, so it is optimistic. The real end-to-end measurement of routed was done on test-2 (below), where it ties Jev.

  • No mode reaches the pre-registered "Jev level" criterion (lower CI bound above โˆ’3 points); medium is the closest.
  • medium beats Jev on arithmetic: 98 vs 80, +18 [9.7; 27.0], Holm-corrected โ€” the only "beats" in the test.
  • ultra-fast fails on tone (52/100); do not use it for tone.

test-2: 156 labeled questions

test-2 accuracy

correct ฮ” vs Jev (points) per question
routed (measured end to end with the released runtime, GPU) 145 +0.6 4.5 s (mean)
Jev 144 โ€” โ€”
slow 144 0.0 11.4 s
medium 141 โˆ’1.9 7.4 s
fast 132 โˆ’7.7 0.45 s
ultra-fast 120 โˆ’15.4 0.28 s

Comparison with similar systems

Same sealed sets, same gold and scoring; sorted by test-3 accuracy. Other systems were run by the maintainers on CPU (AMD Ryzen 5 5600X), pre-registered, with the Jev, medium and fast reference values reproduced in the same run.

system test-3 /900 ฮ” vs Jev [95 % CI] test-2 /156 median s/question hardware
Jev 1.13 (API, baseline) 841 โ€” 144 network remote
routed 829 (post hoc) โˆ’1.3 (no CI) 145 (measured) 4.5 mean RX 7600
medium 816 โˆ’2.8 [โˆ’5.0; โˆ’0.6] 141 7.3 RX 7600
slow 804 โˆ’4.1 [โˆ’6.2; โˆ’2.1] 144 11.6 RX 7600
fast 782 โˆ’6.6 [โˆ’8.9; โˆ’4.3] 132 0.48 RX 7600
ultra-fast 726 โˆ’12.8 [โˆ’15.6; โˆ’10.0] 123 0.29 RX 7600
Laya multilingual (third-party) 462 โˆ’42.1 [โˆ’45.9; โˆ’38.1] 86 0.13 CPU
DeBERTa-v3-large zero-shot 374 โˆ’51.9 [โˆ’55.3; โˆ’48.2] 59 12.8 CPU
BART-large-MNLI zero-shot 358 โˆ’53.7 [โˆ’57.1; โˆ’50.1] 57 0.95 CPU
mDeBERTa-v3 zero-shot (multilingual) 353 โˆ’54.2 [โˆ’57.7; โˆ’50.8] 55 4.2 CPU
Julia-1 (third-party) 322 โˆ’57.7 [โˆ’61.3; โˆ’54.0] 62 0.06 CPU

Every other system differs from Jev, fast and medium with McNemar p < 1e-60 on test-3, with no errors or truncations. CPU and GPU latencies are not directly comparable. GLiClass was not included: same zero-shot label-matching paradigm, no Score or Noul equivalent, and it needs a third-party library. Full aggregate results: reports/comparison/results.md in the code repository.

How it was tested

  • Pre-registration: every test had its plan and decision rule committed before running.
  • Sealed cases: test-3 was sealed by hash before any model ran; Jev ran once on the sealed cases.
  • Gold independent of Jev, from three sources: (a) code-generated items with computed answers, confirmed by a blind verifier (380/380); (b) texts from public datasets (Banking77, GoEmotions); (c) texts written by two authors, labeled blind by two independent labelers (97.4 % agreement, ฮบ 0.973), with an adjudicator for disagreements.
  • Three difficulty levels (N1โ€“N3), fixed before the models ran, with per-domain quotas.
  • Statistics: paired 95 % CI (the more conservative of Newcombe and bootstrap), McNemar, Holm.
  • Reproducibility: a clean-machine run from a fresh clone rebuilt the same GGUF byte for byte and reproduced the 4B fast-path test-2 result exactly.

Run

  1. Get llama.cpp b11205 (Vulkan build, or CPU).
  2. In the code repository: python scripts/fetch_models.py downloads the two Q8_0 files from this repository and checks their SHA-256 (while this repository is private, log in first with hf auth login).
  3. Start the server: python -m typed_decisions serve --config examples/llama-routed.json --port 8000 (--mode ultra-fast|fast|medium|slow fixes one mode). Each model's server starts on its first request (11โ€“22 s); one model is resident at a time.
  4. In routed mode each question carries its domain in the domains field. Full guide: docs/RUNTIME.md.

Real smoke on the RX 7600: 18/18 requests answered, no orphan processes after a normal or an abrupt exit. After start-up each question takes: ultra-fast 0.5 s, fast 0.6 s, medium 6โ€“9 s, slow 12 s.

Memory:

  • 8B with 30 layers on the GPU (-ngl 30, context 2048): ~7.2 GB VRAM. Use --load-mode none (b11205 has no --no-mmap): RAM stays at ~2.6 GB instead of ~10.4 GB.
  • 4B fully on the GPU (-ngl 99): ~5.2 GB VRAM.

Caveats

  • The gold for the authored texts (c) was validated by LLM labelers, without a human audit.
  • Numbers come from llama.cpp b11205 on the GPU. With another engine (e.g. llama-cpp-python on CPU) ~3 % of the answers change.
  • The routing map was chosen by looking at test-3 and confirmed on test-2; claiming Jev level needs a new sealed set.
  • Probabilities are not calibrated; read confidence as a ranking, not a frequency.

License

Apache-2.0, as the original Qwen3 weights (Qwen/Qwen3-4B, Qwen/Qwen3-8B).

Downloads last month
68
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bssgoat/local-typed-decisions

Finetuned
Qwen/Qwen3-4B
Quantized
(331)
this model