File size: 3,415 Bytes
1e9e8ba
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
---

license: apache-2.0
base_model: convaiinnovations/laya-typed-decisions
tags:
  - ggmlc
  - gguf
  - laya
  - jev
  - modernbert
  - decision
  - system-1
language:
  - en
---


# Laya Typed-Decisions GGUF (ggmlc)

**Specialist System 1 decision model** compiled from [convaiinnovations/laya-typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions) (ModernBERT-large, 421M, context 1024). Fine-tuned for invoice match, SOC alerts, customer-service next action, and agent-trace harnesses.

These files are **not** llama.cpp / `llama-cli` GGUFs. They are produced by **[ggmlc](https://github.com/monatis/ggmlc)**, a neural network compiler that lowers PyTorch, JAX, Flax, and Keras models to high-performance GGML execution. Loading them in llama.cpp will fail.

Typed questions (`choice` / `score` / `noul`) are scored in **one encoder pass**. There is no autoregressive token generation.

Source, CLI, and binaries: **[examples/laya](https://github.com/monatis/ggmlc/tree/main/examples/laya)**

Other families: [laya-GGUF](https://huggingface.co/mys/laya-GGUF) (English) · [laya-multilingual-GGUF](https://huggingface.co/mys/laya-multilingual-GGUF)

## Files

| File | Quant | Size | Notes |
| :--- | :--- | ---: | :--- |
| `laya_typed_decisions_f16.gguf` | F16 | ~811 MB | Full precision. |
| `laya_typed_decisions_q8_0.gguf` | Q8_0 | ~434 MB | Smaller, still accurate. |

| `laya_typed_decisions_ud_q4_k_m.gguf` | UD_Q4_K_M | ~404 MB | Smallest. 1D norms/biases stay F32. |

English ModernBERT tokenizer (`[CLS]/[SEP]/[PAD]/[MASK]`). `max_len=1024`.

```bash

huggingface-cli download mys/laya-typed-decisions-GGUF laya_typed_decisions_f16.gguf --local-dir .

```

## Run with `laya`

Download a binary from [ggmlc releases](https://github.com/monatis/ggmlc/releases/latest) (`laya.exe` / `laya`). `--device` defaults to `auto` (CUDA or Metal if present, else CPU).

`--models-dir` routing does **not** pick this family automatically unless `--family typed-decisions` or the question ids match a specialist workflow (`invoice`, `security`, `customer_service`, `harness`).

```bash

laya help

laya list-presets

laya info laya_typed_decisions_f16.gguf



laya decide laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph

laya decide laya_typed_decisions_f16.gguf --preset security --json

laya decide laya_typed_decisions_f16.gguf --preset customer_service

laya decide laya_typed_decisions_f16.gguf --preset harness

laya serve laya_typed_decisions_f16.gguf --port 8080 --device auto --cuda-graph

laya bench laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph

```

Force it from a mixed `--models-dir`:

```bash

laya decide --models-dir . --family typed-decisions --preset invoice

```

`serve` starts Decision Studio (`GET /`) and `POST /api/decide`. `daemon` is newline JSON-RPC on stdin/stdout.

## What this is

[Laya](https://github.com/NandhaKishorM/laya) is the open reproduction of TypeSafe **Jev**: given a *state* and typed questions, it returns calibrated probabilities instead of generating tokens. This checkpoint is the specialist sibling of the general English model — same tokenizer family, longer context, trained on typed-decision workflows.

## License

Apache 2.0, same as the upstream Laya weights. Compiler: [ggmlc](https://github.com/monatis/ggmlc) (MIT).