Sandhi-1.0-8M-int8-executorch
The int8 ExecuTorch build of Sandhi-1.0-8M, a small character-level model that analyses a Malayalam word offline on a phone: sandhi split, morphemes, root lemma, grammatical features, IPA and syllables, and a confidence level.
Files
| File | What |
|---|---|
sandhi_int8.pte |
int8 ExecuTorch program (XNNPACK), 9,594,204 bytes. Three graphs: encoder, cross, step |
sandhi_meta.json |
The runtime contract: vocabulary, label sets, tensor shapes and dtypes of each graph, thresholds, limits, checksums |
code/normalize.py |
The one normalize() function. Input must be normalized with this before the model sees it. |
golden/ |
Test vectors: 1,017 full analyses (golden_analyse.jsonl) and normalization cases (golden_normalize.jsonl), for checking a port against the reference |
reports/ |
The quantization and export reports, as measured |
LICENSE, NOTICE |
Apache-2.0, and attribution for the training sources |
How to run it
This is a .pte program, not a transformers model. Load it with the ExecuTorch runtime (built with the XNNPACK backend) and follow sandhi_meta.json:
encoder: takes[<FULL>] + one id per character(int64) and returns boundary logits, syllable logits, the multi-labelotherhead, the classifier-head logits, and amemorytensor.cross: takesmemoryand returns the cross-attention keys and values.step: one decoder step with a fixed-size KV cache. The caller drives greedy decoding for two rows,root.lemmaandipa, up to 40 steps.
Words are limited to 39 characters. Thresholds and the confidence wording are in sandhi_meta.json. The first thing to do with any port is to run golden/: a correct port reproduces golden_analyse.jsonl. A reference Kotlin implementation lives in the app source at https://github.com/nithinmanoj10/sandhi (android/).
Example
The code that drives this .pte through the ExecuTorch runtime is the package ml-Sandhi. It takes its vocabulary and shapes from sandhi_meta.json, so it needs only the files in this repository.
pip install "ml-sandhi[int8] @ git+https://github.com/nithinmanoj10/ml-Sandhi"
import json
from ml_sandhi import Sandhi
sandhi = Sandhi.from_pretrained_int8() # downloads sandhi_int8.pte and sandhi_meta.json on first use
result = sandhi.analyse("കമ്പ്യൂട്ടറുകളിലൂടെ")
print(json.dumps(result, ensure_ascii=False, indent=2))
Output (fields shown; the rest is omitted). On this word it is identical to the fp32 model's:
{
"input": "കമ്പ്യൂട്ടറുകളിലൂടെ",
"source": "model",
"split": "കമ്പ്യൂട്ടറ + ുകള + ിലൂടെ",
"root": {"lemma": "കമ്പ്യൂട്ടർ", "pos": "n", "type": "known"},
"features": {"case": "perlative"},
"ipa": "kampjuːʈʈarukaɭiluːʈe",
"syllables": ["ക", "മ്പ്യൂ", "ട്ട", "റു", "ക", "ളി", "ലൂ", "ടെ"],
"confidence_text": "Fairly sure"
}
കമ്പ്യൂട്ടറുകളിലൂടെ ("through the computers") is a loanword plus two endings, and is not in the dictionary, so source is model. The lemma കമ്പ്യൂട്ടർ is not a prefix of the surface: sandhi rewrote the stem. The full result also carries morphemes, display and alternatives.
What quantization cost
Measured against the fp32 model (reports/phase3_quantization_v3.md, reports/phase3_export.md):
- On 3,000 dev words the int8 program agrees with fp32 on 99.93% of splits, 99.87% of lemmas, 99.90% of IPA, 99.93% of feature sets and 99.87% of confidence labels.
- The
torch.aoestimate is more pessimistic: split −0.24 points overall and −1.97 points on words of four or more morphemes. Treat that as the conservative bound. - Size: 33.1 MiB fp32 → 9.1 MiB int8.
License
Apache-2.0. Training sources are attributed in NOTICE. mlmorph and Mlphon were dev-time label generators and are not distributed here.
- Downloads last month
- 9
Model tree for nithinmanoj10/Sandhi-1.0-8M-int8-executorch
Base model
nithinmanoj10/Sandhi-1.0-8M
