bert-tiny β Native C++ Port (Sajal Labs)
This is not a new model. The weights, architecture, and pretraining are
entirely prajjwal1/bert-tiny by Prajjwal
Bhargava (MIT license) β a compact pretrained BERT encoder introduced in
Turc et al. 2019 ("Well-Read Students Learn Better") and ported to
Hugging Face for Bhargava et al. 2021 ("Generalization in NLI"). Please
cite both papers if you use this model (citations below).
What Sajal Labs added: a from-scratch native C++ port of the encoder
and the WordPiece tokenizer β no PyTorch, no transformers, no Python at
inference time β with rigorous equivalence and benchmark validation against
the original. See Sajal Labs,
experiments
exp11
through
exp14,
for full methodology.
Model details (unchanged from the original)
- Architecture: BERT encoder, 2 layers, hidden=128, heads=2, intermediate=512
- Vocabulary: 30522 WordPiece tokens (bert-base-uncased vocab)
- Parameters: 4,385,920
- Precision: fp32
- Base model license: MIT (prajjwal1/bert-tiny)
- Port license: MIT (Sajal Labs' C++ code)
What was verified (Sajal Labs' contribution)
Tokenizer: 17/17 real test sentences β including contractions ("don't"),
hyphenation ("COVID-19"), an out-of-vocabulary word forcing an 11-piece
subword split, and an email address β produced byte-identical token IDs
to the original BertTokenizerFast. This is an exact-match bar, not a
tolerance: tokenization is deterministic.
Encoder: max absolute error 9.54e-06 (hidden states),
2.19e-06 (pooled [CLS] output), cosine similarity
~1.0, across 10 real sentences of varying length (4-25 tokens) β fp32-scale
agreement, consistent with floating-point non-associativity between two
independent implementations (not a bug; see the repo's research/papers.md).
Benchmark (single request, Apple M4 CPU β full data in benchmark_results.json)
True end-to-end cold invocation (process spawn β raw text in β prediction out
β process exit, external wall-clock). Primary comparison: native vs. ONNX
Runtime paired with the lean, standalone tokenizers library β the
best-case Python deployment, not the easiest target to beat:
| Implementation | Cold invocation p50 |
|---|---|
| Native C++ (Sajal runtime) | 11.37ms |
| ONNX Runtime + lean tokenizer | 95.72ms (8.4x slower) |
| ONNX Runtime + π€ transformers tokenizer | 2495.61ms (219.6x slower) |
| PyTorch + π€ transformers | 4973.91ms (437.6x slower) |
The last two rows are real and worth knowing, but they're the easy targets
(heavier Python stacks) β the 8.4x number above is
the one that holds up against someone who already optimized their Python
deployment correctly. Worth knowing before you read too much into "ONNX
Runtime" as a single number: its own cold-start
number depends heavily on which tokenizer library it's paired with β using
transformers for convenience costs ~25x more than using the lean, standalone
tokenizers library for the exact same token IDs. Native sidesteps that
whole dependency-choice question by construction. Full discussion in
exp14.
Honest scope note on the warm-loop numbers: native's advantage is
not unconditional the way cold-invocation is β Sajal Labs found it
depends on model width (hidden_size), with a measured crossover around
hiddenβ250 on this hardware
(exp12/exp13).
This model's hidden=128
sits comfortably below that, so native keeps a real warm-loop edge too β
but that's a property of this model's size, not a general claim.
How to use
Native (Sajal runtime, zero Python)
sajal run <this_directory> "The quick brown fox jumps over the lazy dog."
PyTorch / transformers (the original)
from transformers import BertModel, BertTokenizerFast
model = BertModel.from_pretrained("prajjwal1/bert-tiny")
tokenizer = BertTokenizerFast.from_pretrained("prajjwal1/bert-tiny")
ONNX Runtime
import onnxruntime as ort
session = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
# feed input_ids/attention_mask/token_type_ids from any WordPiece tokenizer
Citations (required if you use this model)
@article{turc2019distillation,
title={Well-Read Students Learn Better: On the Importance of Pre-training Compact Models},
author={Turc, Iulia and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
journal={arXiv preprint arXiv:1908.08962v2},
year={2019}
}
@misc{bhargava2021generalization,
title={Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics},
author={Bhargava, Prajjwal and Drozd, Aleksandr and Rogers, Anna},
year={2021},
eprint={2110.01518},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Files
model.safetensorsβ the original weights, HF/PyTorch-ecosystem formatmodel.onnx+model.onnx.dataβ ONNX Runtime-compatible export (weights externalized to the.datafile; both are required together)word_embeddings.bin,position_embeddings.bin,token_type_embeddings.bin,emb_ln_*.bin,layer{i}_*.bin,pooler_*.binβ raw native Sajal runtime formatvocab.txt,tokenizer_config.txtβ WordPiece vocabulary + config for the native tokenizer portconfig.jsonβ architecture metadatabenchmark_results.jsonβ full machine-readable benchmark/equivalence data behind the numbers above
- Downloads last month
- 36
Model tree for sajalmadan09/bert-tiny-native-cpp
Base model
prajjwal1/bert-tiny