DZAIR-ONNX-INT8
Dynamically quantized 8-bit ONNX build of algerian-nlp/DZAIR, the 105.3M-parameter encoder for Algerian Darija. A quarter of the disk of the fp32 graph, for deployments where binary size is the binding constraint.
Fidelity and latency
Measured on an Apple Silicon host
(macOS arm64) against the torch fp32 release. The full report ships as
onnx_report.json.
| ONNX fp32 | ONNX int8 | |
|---|---|---|
model.onnx |
424.0 MB | 108.4 MB (3.89× smaller) |
| Cosine against torch fp32 | 1.000000 | 0.999579 |
| Max absolute difference | 3.9e-06 | 0.207 |
| Latency at batch 8 × length 128 | 165.2 ms / 6,198 tok/s | 170.8 ms / 5,996 tok/s |
Release gate: cosine ≥ 0.98. Quantization is symmetric per-channel QInt8 on
weights with dynamic per-token activation scaling, applied to the MatMul and
Gemm operators by onnxruntime.quantization.quantize_dynamic.
Int8 is not faster on this host. It measured 3% slower than the fp32 graph (170.8 ms against 165.2 ms). Dynamic quantization pays a per-inference scaling cost that ARM CPUs do not always win back; on x86 with VNNI the trade usually goes the other way. Choose this build for size, and measure latency on your own target before assuming a speedup.
Downstream accuracy under int8 has not been measured. The fidelity number above is embedding-level cosine, not task accuracy; no fine-tuning run exists for this build, so no accuracy figure is quoted for it.
Usage
onnxruntime alone — no torch, no transformers at serving time. The graph
takes input_ids and attention_mask and returns last_hidden_state, with
both batch and sequence axes dynamic.
import numpy as np
import onnxruntime
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
REPO = "algerian-nlp/DZAIR-ONNX-INT8"
tokenizer = AutoTokenizer.from_pretrained("algerian-nlp/DZAIR", trust_remote_code=True)
session = onnxruntime.InferenceSession(
hf_hub_download(REPO, "model.onnx"), providers=["CPUExecutionProvider"]
)
texts = ["واش راك خويا لاباس عليك؟", "rani rayeh lel dar bech nchouf la famille ya kho"]
encoded = tokenizer([t.lower() for t in texts], padding=True, return_tensors="np")
hidden = session.run(
None,
{"input_ids": encoded["input_ids"], "attention_mask": encoded["attention_mask"]},
)[0]
mask = encoded["attention_mask"][..., None]
embeddings = (hidden * mask).sum(axis=1) / mask.sum(axis=1)
print(hidden.shape, embeddings.shape) # (2, N, 768) (2, 768)
The tokenizer comes from the base repo and is also shipped here, so the graph can be served without it once ids are produced upstream.
optimum does not work here. ORTModelForFeatureExtraction fails to
import against transformers 5.x — optimum 2.1.0 still reads
is_offline_mode from transformers.utils, which 5.x removed (verified
2026-09-19). Use the onnxruntime session above, or pin transformers<5
if you need the optimum wrapper.
Input rules. Lowercase Latin spans before encoding; the tokenizer wraps
each row as [CLS] ... [SEP] and pads on the right. Do not transliterate
Arabizi phoneme digits (3, 7, 9) — they are atomic pieces in this vocabulary.
Files
| file | size | contents |
|---|---|---|
model.onnx |
108.4 MB | dynamically quantized int8 graph, opset 18 |
config.json |
1 KB | architecture of the source model |
tokenizer.model, tokenizer_config.json |
about 1.0 MB | 48k SentencePiece Unigram via DebertaV2Tokenizer, specials at ids 0–4, right padding |
tokenizer_rules.yaml |
2 KB | versioned normalisation rules |
onnx_report.json |
1 KB | the fidelity and latency measurements above, with per-file SHA-256 in the fp32 export report |
Licence
Apache-2.0, inherited from the base model. Read the licence composition on the base card before redistributing derivatives.
- Downloads last month
- 107
Model tree for algerian-nlp/DZAIR-ONNX-INT8
Base model
algerian-nlp/DZAIR