mys commited on
Commit
1e9e8ba
·
verified ·
1 Parent(s): f530b72

Add ggmlc-compiled Laya GGUFs (F16, Q8_0, UD_Q4_K_M)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,6 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ laya_typed_decisions_f16.gguf filter=lfs diff=lfs merge=lfs -text
37
+ laya_typed_decisions_q8_0.gguf filter=lfs diff=lfs merge=lfs -text
38
+ laya_typed_decisions_ud_q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,75 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: convaiinnovations/laya-typed-decisions
4
+ tags:
5
+ - ggmlc
6
+ - gguf
7
+ - laya
8
+ - jev
9
+ - modernbert
10
+ - decision
11
+ - system-1
12
+ language:
13
+ - en
14
+ ---
15
+
16
+ # Laya Typed-Decisions GGUF (ggmlc)
17
+
18
+ **Specialist System 1 decision model** compiled from [convaiinnovations/laya-typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions) (ModernBERT-large, 421M, context 1024). Fine-tuned for invoice match, SOC alerts, customer-service next action, and agent-trace harnesses.
19
+
20
+ These files are **not** llama.cpp / `llama-cli` GGUFs. They are produced by **[ggmlc](https://github.com/monatis/ggmlc)**, a neural network compiler that lowers PyTorch, JAX, Flax, and Keras models to high-performance GGML execution. Loading them in llama.cpp will fail.
21
+
22
+ Typed questions (`choice` / `score` / `noul`) are scored in **one encoder pass**. There is no autoregressive token generation.
23
+
24
+ Source, CLI, and binaries: **[examples/laya](https://github.com/monatis/ggmlc/tree/main/examples/laya)**
25
+
26
+ Other families: [laya-GGUF](https://huggingface.co/mys/laya-GGUF) (English) · [laya-multilingual-GGUF](https://huggingface.co/mys/laya-multilingual-GGUF)
27
+
28
+ ## Files
29
+
30
+ | File | Quant | Size | Notes |
31
+ | :--- | :--- | ---: | :--- |
32
+ | `laya_typed_decisions_f16.gguf` | F16 | ~811 MB | Full precision. |
33
+ | `laya_typed_decisions_q8_0.gguf` | Q8_0 | ~434 MB | Smaller, still accurate. |
34
+ | `laya_typed_decisions_ud_q4_k_m.gguf` | UD_Q4_K_M | ~404 MB | Smallest. 1D norms/biases stay F32. |
35
+
36
+ English ModernBERT tokenizer (`[CLS]/[SEP]/[PAD]/[MASK]`). `max_len=1024`.
37
+
38
+ ```bash
39
+ huggingface-cli download mys/laya-typed-decisions-GGUF laya_typed_decisions_f16.gguf --local-dir .
40
+ ```
41
+
42
+ ## Run with `laya`
43
+
44
+ Download a binary from [ggmlc releases](https://github.com/monatis/ggmlc/releases/latest) (`laya.exe` / `laya`). `--device` defaults to `auto` (CUDA or Metal if present, else CPU).
45
+
46
+ `--models-dir` routing does **not** pick this family automatically unless `--family typed-decisions` or the question ids match a specialist workflow (`invoice`, `security`, `customer_service`, `harness`).
47
+
48
+ ```bash
49
+ laya help
50
+ laya list-presets
51
+ laya info laya_typed_decisions_f16.gguf
52
+
53
+ laya decide laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph
54
+ laya decide laya_typed_decisions_f16.gguf --preset security --json
55
+ laya decide laya_typed_decisions_f16.gguf --preset customer_service
56
+ laya decide laya_typed_decisions_f16.gguf --preset harness
57
+ laya serve laya_typed_decisions_f16.gguf --port 8080 --device auto --cuda-graph
58
+ laya bench laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph
59
+ ```
60
+
61
+ Force it from a mixed `--models-dir`:
62
+
63
+ ```bash
64
+ laya decide --models-dir . --family typed-decisions --preset invoice
65
+ ```
66
+
67
+ `serve` starts Decision Studio (`GET /`) and `POST /api/decide`. `daemon` is newline JSON-RPC on stdin/stdout.
68
+
69
+ ## What this is
70
+
71
+ [Laya](https://github.com/NandhaKishorM/laya) is the open reproduction of TypeSafe **Jev**: given a *state* and typed questions, it returns calibrated probabilities instead of generating tokens. This checkpoint is the specialist sibling of the general English model — same tokenizer family, longer context, trained on typed-decision workflows.
72
+
73
+ ## License
74
+
75
+ Apache 2.0, same as the upstream Laya weights. Compiler: [ggmlc](https://github.com/monatis/ggmlc) (MIT).
laya_typed_decisions_f16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:71bf44211496be01cb8057b5ce906aaba62cd0192115c3abb5b42b2f0147ee2c
3
+ size 849812128
laya_typed_decisions_q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6eb6ef58bc4f99bd9db40381604108af782ffc314aa00da6641f429991ca1f18
3
+ size 455179648
laya_typed_decisions_ud_q4_k_m.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:42919cae9c11f744a16b8f3136a3356c586a558108080050e832ba14bbb0010c
3
+ size 423581888