--- license: mit base_model: segment-any-text/sat-3l-sm library_name: openvino tags: - openvino - int8 - nncf - quantization - sentence-segmentation - sat - wtpsplit language: - multilingual --- # sat-3l-sm - native OpenVINO INT8 Native OpenVINO INT8 build of the wtpsplit **SaT** (`sat-3l-sm`) multilingual sentence segmenter. A drop-in replacement for the onnxruntime SaT that runs on OpenVINO with no `onnxruntime` dependency, quantized for CPU latency. - **Base** - `segment-any-text/sat-3l-sm` (wtpsplit SaT, 3-layer subword XLM-R), converted ONNX → OpenVINO IR, then INT8 post-training quantization with NNCF - **Runtime** - OpenVINO CPU, compiled with the LATENCY performance hint (single-sentence-at-a-time use); ~half the FP32 size (205 MB vs 408 MB) - **Calibration** - the public **VitaminC** dev set (claims + evidence); calibration inputs captured from a reference run so NNCF sees the true activation distribution - **Inputs** - `input_ids` (int64), `attention_mask` (float32); output `logits` (per-token boundary scores) - **Tokenizer** - `facebookAI/xlm-roberta-base` ## Quantization fidelity (acceptance gate) Measured against the ONNX reference on a held-out, public VitaminC sample (300 texts, 31k characters) at the `sat-3l-sm` default threshold 0.25: - **Pearson r = 0.9986** on per-character newline-probabilities (newline-prob MAE 1.9e-04) - **97.4%** sentence-split agreement, **98.7%** of texts split identically - The FP16 IR is the lossless reference point (r = 1.0); INT8 is the size win ## Files - `openvino_model.xml` + `openvino_model.bin` - the INT8 OpenVINO IR - `config.json`, `openvino_config.json` - model + quantization config - `tokenizer.json`, `tokenizer_config.json`, `special_tokens_map.json` - the xlm-roberta-base tokenizer ## Usage ```python import openvino as ov from huggingface_hub import hf_hub_download xml = hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.xml") hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.bin") core = ov.Core() model = core.compile_model(core.read_model(xml), "CPU", {"PERFORMANCE_HINT": "LATENCY"}) # input_ids: int64 [batch, seq], attention_mask: float32 [batch, seq] -> logits logits = list(model({"input_ids": ids, "attention_mask": mask}).values())[0] ``` The full tokenise → sliding-window → threshold → sentence-reconstruction pipeline follows wtpsplit's SaT; this repo only swaps the inference backend to OpenVINO INT8. ## License and attribution MIT. Derived from [wtpsplit-lite](https://github.com/superlinear-ai/wtpsplit-lite) (Superlinear) and the SaT model [`segment-any-text/sat-3l-sm`](https://huggingface.co/segment-any-text/sat-3l-sm). See the base model and wtpsplit for the segmentation method.