sat-3l-sm - native OpenVINO INT8
Native OpenVINO INT8 build of the wtpsplit SaT (sat-3l-sm) multilingual sentence segmenter. A drop-in replacement for the onnxruntime SaT that runs on OpenVINO with no onnxruntime dependency, quantized for CPU latency.
- Base -
segment-any-text/sat-3l-sm(wtpsplit SaT, 3-layer subword XLM-R), converted ONNX โ OpenVINO IR, then INT8 post-training quantization with NNCF - Runtime - OpenVINO CPU, compiled with the LATENCY performance hint (single-sentence-at-a-time use); ~half the FP32 size (205 MB vs 408 MB)
- Calibration - the public VitaminC dev set (claims + evidence); calibration inputs captured from a reference run so NNCF sees the true activation distribution
- Inputs -
input_ids(int64),attention_mask(float32); outputlogits(per-token boundary scores) - Tokenizer -
facebookAI/xlm-roberta-base
Quantization fidelity (acceptance gate)
Measured against the ONNX reference on a held-out, public VitaminC sample (300 texts, 31k characters) at the sat-3l-sm default threshold 0.25:
- Pearson r = 0.9986 on per-character newline-probabilities (newline-prob MAE 1.9e-04)
- 97.4% sentence-split agreement, 98.7% of texts split identically
- The FP16 IR is the lossless reference point (r = 1.0); INT8 is the size win
Files
openvino_model.xml+openvino_model.bin- the INT8 OpenVINO IRconfig.json,openvino_config.json- model + quantization configtokenizer.json,tokenizer_config.json,special_tokens_map.json- the xlm-roberta-base tokenizer
Usage
import openvino as ov
from huggingface_hub import hf_hub_download
xml = hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.xml")
hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.bin")
core = ov.Core()
model = core.compile_model(core.read_model(xml), "CPU", {"PERFORMANCE_HINT": "LATENCY"})
# input_ids: int64 [batch, seq], attention_mask: float32 [batch, seq] -> logits
logits = list(model({"input_ids": ids, "attention_mask": mask}).values())[0]
The full tokenise โ sliding-window โ threshold โ sentence-reconstruction pipeline follows wtpsplit's SaT; this repo only swaps the inference backend to OpenVINO INT8.
License and attribution
MIT. Derived from wtpsplit-lite (Superlinear) and the SaT model segment-any-text/sat-3l-sm. See the base model and wtpsplit for the segmentation method.
- Downloads last month
- 4
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for stellars/sat-3l-sm-openvino-int8
Base model
segment-any-text/sat-3l-sm