File size: 3,415 Bytes
1e9e8ba | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 | ---
license: apache-2.0
base_model: convaiinnovations/laya-typed-decisions
tags:
- ggmlc
- gguf
- laya
- jev
- modernbert
- decision
- system-1
language:
- en
---
# Laya Typed-Decisions GGUF (ggmlc)
**Specialist System 1 decision model** compiled from [convaiinnovations/laya-typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions) (ModernBERT-large, 421M, context 1024). Fine-tuned for invoice match, SOC alerts, customer-service next action, and agent-trace harnesses.
These files are **not** llama.cpp / `llama-cli` GGUFs. They are produced by **[ggmlc](https://github.com/monatis/ggmlc)**, a neural network compiler that lowers PyTorch, JAX, Flax, and Keras models to high-performance GGML execution. Loading them in llama.cpp will fail.
Typed questions (`choice` / `score` / `noul`) are scored in **one encoder pass**. There is no autoregressive token generation.
Source, CLI, and binaries: **[examples/laya](https://github.com/monatis/ggmlc/tree/main/examples/laya)**
Other families: [laya-GGUF](https://huggingface.co/mys/laya-GGUF) (English) · [laya-multilingual-GGUF](https://huggingface.co/mys/laya-multilingual-GGUF)
## Files
| File | Quant | Size | Notes |
| :--- | :--- | ---: | :--- |
| `laya_typed_decisions_f16.gguf` | F16 | ~811 MB | Full precision. |
| `laya_typed_decisions_q8_0.gguf` | Q8_0 | ~434 MB | Smaller, still accurate. |
| `laya_typed_decisions_ud_q4_k_m.gguf` | UD_Q4_K_M | ~404 MB | Smallest. 1D norms/biases stay F32. |
English ModernBERT tokenizer (`[CLS]/[SEP]/[PAD]/[MASK]`). `max_len=1024`.
```bash
huggingface-cli download mys/laya-typed-decisions-GGUF laya_typed_decisions_f16.gguf --local-dir .
```
## Run with `laya`
Download a binary from [ggmlc releases](https://github.com/monatis/ggmlc/releases/latest) (`laya.exe` / `laya`). `--device` defaults to `auto` (CUDA or Metal if present, else CPU).
`--models-dir` routing does **not** pick this family automatically unless `--family typed-decisions` or the question ids match a specialist workflow (`invoice`, `security`, `customer_service`, `harness`).
```bash
laya help
laya list-presets
laya info laya_typed_decisions_f16.gguf
laya decide laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph
laya decide laya_typed_decisions_f16.gguf --preset security --json
laya decide laya_typed_decisions_f16.gguf --preset customer_service
laya decide laya_typed_decisions_f16.gguf --preset harness
laya serve laya_typed_decisions_f16.gguf --port 8080 --device auto --cuda-graph
laya bench laya_typed_decisions_f16.gguf --preset invoice --device auto --cuda-graph
```
Force it from a mixed `--models-dir`:
```bash
laya decide --models-dir . --family typed-decisions --preset invoice
```
`serve` starts Decision Studio (`GET /`) and `POST /api/decide`. `daemon` is newline JSON-RPC on stdin/stdout.
## What this is
[Laya](https://github.com/NandhaKishorM/laya) is the open reproduction of TypeSafe **Jev**: given a *state* and typed questions, it returns calibrated probabilities instead of generating tokens. This checkpoint is the specialist sibling of the general English model — same tokenizer family, longer context, trained on typed-decision workflows.
## License
Apache 2.0, same as the upstream Laya weights. Compiler: [ggmlc](https://github.com/monatis/ggmlc) (MIT).
|