sat-3l-sm - native OpenVINO INT8

Native OpenVINO INT8 build of the wtpsplit SaT (sat-3l-sm) multilingual sentence segmenter. A drop-in replacement for the onnxruntime SaT that runs on OpenVINO with no onnxruntime dependency, quantized for CPU latency.

  • Base - segment-any-text/sat-3l-sm (wtpsplit SaT, 3-layer subword XLM-R), converted ONNX โ†’ OpenVINO IR, then INT8 post-training quantization with NNCF
  • Runtime - OpenVINO CPU, compiled with the LATENCY performance hint (single-sentence-at-a-time use); ~half the FP32 size (205 MB vs 408 MB)
  • Calibration - the public VitaminC dev set (claims + evidence); calibration inputs captured from a reference run so NNCF sees the true activation distribution
  • Inputs - input_ids (int64), attention_mask (float32); output logits (per-token boundary scores)
  • Tokenizer - facebookAI/xlm-roberta-base

Quantization fidelity (acceptance gate)

Measured against the ONNX reference on a held-out, public VitaminC sample (300 texts, 31k characters) at the sat-3l-sm default threshold 0.25:

  • Pearson r = 0.9986 on per-character newline-probabilities (newline-prob MAE 1.9e-04)
  • 97.4% sentence-split agreement, 98.7% of texts split identically
  • The FP16 IR is the lossless reference point (r = 1.0); INT8 is the size win

Files

  • openvino_model.xml + openvino_model.bin - the INT8 OpenVINO IR
  • config.json, openvino_config.json - model + quantization config
  • tokenizer.json, tokenizer_config.json, special_tokens_map.json - the xlm-roberta-base tokenizer

Usage

import openvino as ov
from huggingface_hub import hf_hub_download

xml = hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.xml")
hf_hub_download("stellars/sat-3l-sm-openvino-int8", "openvino_model.bin")
core = ov.Core()
model = core.compile_model(core.read_model(xml), "CPU", {"PERFORMANCE_HINT": "LATENCY"})
# input_ids: int64 [batch, seq], attention_mask: float32 [batch, seq] -> logits
logits = list(model({"input_ids": ids, "attention_mask": mask}).values())[0]

The full tokenise โ†’ sliding-window โ†’ threshold โ†’ sentence-reconstruction pipeline follows wtpsplit's SaT; this repo only swaps the inference backend to OpenVINO INT8.

License and attribution

MIT. Derived from wtpsplit-lite (Superlinear) and the SaT model segment-any-text/sat-3l-sm. See the base model and wtpsplit for the segmentation method.

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for stellars/sat-3l-sm-openvino-int8

Finetuned
(1)
this model