DZAIR-ONNX-INT8

Dynamically quantized 8-bit ONNX build of algerian-nlp/DZAIR, the 105.3M-parameter encoder for Algerian Darija. A quarter of the disk of the fp32 graph, for deployments where binary size is the binding constraint.

Fidelity and latency

Measured on an Apple Silicon host (macOS arm64) against the torch fp32 release. The full report ships as onnx_report.json.

ONNX fp32 ONNX int8
model.onnx 424.0 MB 108.4 MB (3.89× smaller)
Cosine against torch fp32 1.000000 0.999579
Max absolute difference 3.9e-06 0.207
Latency at batch 8 × length 128 165.2 ms / 6,198 tok/s 170.8 ms / 5,996 tok/s

Release gate: cosine ≥ 0.98. Quantization is symmetric per-channel QInt8 on weights with dynamic per-token activation scaling, applied to the MatMul and Gemm operators by onnxruntime.quantization.quantize_dynamic.

Int8 is not faster on this host. It measured 3% slower than the fp32 graph (170.8 ms against 165.2 ms). Dynamic quantization pays a per-inference scaling cost that ARM CPUs do not always win back; on x86 with VNNI the trade usually goes the other way. Choose this build for size, and measure latency on your own target before assuming a speedup.

Downstream accuracy under int8 has not been measured. The fidelity number above is embedding-level cosine, not task accuracy; no fine-tuning run exists for this build, so no accuracy figure is quoted for it.

Usage

onnxruntime alone — no torch, no transformers at serving time. The graph takes input_ids and attention_mask and returns last_hidden_state, with both batch and sequence axes dynamic.

import numpy as np
import onnxruntime
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer

REPO = "algerian-nlp/DZAIR-ONNX-INT8"
tokenizer = AutoTokenizer.from_pretrained("algerian-nlp/DZAIR", trust_remote_code=True)
session = onnxruntime.InferenceSession(
    hf_hub_download(REPO, "model.onnx"), providers=["CPUExecutionProvider"]
)

texts = ["واش راك خويا لاباس عليك؟", "rani rayeh lel dar bech nchouf la famille ya kho"]
encoded = tokenizer([t.lower() for t in texts], padding=True, return_tensors="np")
hidden = session.run(
    None,
    {"input_ids": encoded["input_ids"], "attention_mask": encoded["attention_mask"]},
)[0]

mask = encoded["attention_mask"][..., None]
embeddings = (hidden * mask).sum(axis=1) / mask.sum(axis=1)
print(hidden.shape, embeddings.shape)  # (2, N, 768) (2, 768)

The tokenizer comes from the base repo and is also shipped here, so the graph can be served without it once ids are produced upstream.

optimum does not work here. ORTModelForFeatureExtraction fails to import against transformers 5.x — optimum 2.1.0 still reads is_offline_mode from transformers.utils, which 5.x removed (verified 2026-09-19). Use the onnxruntime session above, or pin transformers<5 if you need the optimum wrapper.

Input rules. Lowercase Latin spans before encoding; the tokenizer wraps each row as [CLS] ... [SEP] and pads on the right. Do not transliterate Arabizi phoneme digits (3, 7, 9) — they are atomic pieces in this vocabulary.

Files

file size contents
model.onnx 108.4 MB dynamically quantized int8 graph, opset 18
config.json 1 KB architecture of the source model
tokenizer.model, tokenizer_config.json about 1.0 MB 48k SentencePiece Unigram via DebertaV2Tokenizer, specials at ids 0–4, right padding
tokenizer_rules.yaml 2 KB versioned normalisation rules
onnx_report.json 1 KB the fidelity and latency measurements above, with per-file SHA-256 in the fp32 export report

Licence

Apache-2.0, inherited from the base model. Read the licence composition on the base card before redistributing derivatives.

Downloads last month
107
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for algerian-nlp/DZAIR-ONNX-INT8

Quantized
(2)
this model