Token Classification
Transformers
Safetensors
Persian
bert
feature-extraction
persian
word-importance
salience
dhh
asr-evaluation
ace-metric
distillation
Instructions to use Reza2kn/ShenavaSanj-v1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Reza2kn/ShenavaSanj-v1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Reza2kn/ShenavaSanj-v1.0")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("Reza2kn/ShenavaSanj-v1.0") model = AutoModel.from_pretrained("Reza2kn/ShenavaSanj-v1.0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| language: fa | |
| library_name: transformers | |
| pipeline_tag: token-classification | |
| base_model: HooshvareLab/bert-base-parsbert-uncased | |
| tags: [persian, word-importance, salience, dhh, asr-evaluation, ace-metric, distillation] | |
| # 🎙️ ShenavaSanj-v1.0 | |
| > **English + فارسی** · Part of [Shenava 1.0](https://huggingface.co/collections/Reza2kn/shenava-10-open-streaming-persian-asr-and-captioning) · [Project hub](https://github.com/Reza2kn/shenava-1) · [SLT paper submission](https://openreview.net/forum?id=QTa6ax9PU3) | |
| ## 🌟 At a glance | معرفی سریع | |
| | | English | فارسی | | |
| |---|---|---| | |
| | 🎯 Role | Persian word-importance scorer behind S3 / Semantic WER. | امتیازدهندهٔ اهمیت معنایی شنواسنج؛ این مخزن یکی از مصنوعات رسمی خانوادهٔ Shenava-1 است. | | |
| | 🧠 Family | ShenavaSanj semantic importance scorer | امتیازدهندهٔ اهمیت معنایی شنواسنج | | |
| | 📦 Format | Transformers token-classification checkpoint | چکپوینت دستهبندی توکن در Transformers | | |
| | 📐 Scale | approximately 200M parameters | اندازه: approximately 200M parameters | | |
| | 📥 Input | Persian text tokenized with the repository tokenizer | متن فارسی پس از توکنسازی با توکنایزر مخزن | | |
| | 📤 Output | Per-token importance scores used by S³ / semantic WER | امتیاز اهمیت برای هر توکن جهت محاسبهٔ S³ و WER معنایی | | |
| | ⚖️ License | Apache License 2.0 | مجوز Apache 2.0 | | |
| ## 🇬🇧 English documentation | |
| ### 🧭 Overview | |
| Persian word-importance scorer behind S3 / Semantic WER. This repository is an official Shenava-1 release artifact, not an isolated checkpoint. It belongs to a Persian-first stack covering training data, streaming ASR, semantic evaluation, on-device exports, captioning applications, and reproducible benchmarks. Use the collection link above to locate sibling model sizes, deployment formats, datasets, and evaluation assets. | |
| The artifact is optimized for Persian speech and the conventions used by the Shenava/VisualEars pipeline. A model file alone is not the entire inference system: audio preparation, tokenizer assets, streaming state, decoding, Persian text normalization, and inverse text normalization can materially affect observed output. | |
| ### ✅ Intended uses | |
| - Persian ASR research, benchmarking, and reproducible comparison inside the Shenava-1 evaluation protocol. | |
| - Offline or streaming transcription when the selected runtime and graph support that mode. | |
| - On-device captioning, accessibility prototypes, and Persian speech interfaces. | |
| - Conversion or runtime integration work that preserves the source model’s tokenizer, decoding assumptions, and numerical checks. | |
| ### 🚫 Out-of-scope or unsafe uses | |
| - Do not treat transcripts as guaranteed verbatim records for legal, medical, emergency, or other high-stakes decisions. | |
| - Do not infer identity, health, ethnicity, intent, or other sensitive traits from speech or model errors. | |
| - Do not compare formats using different text normalization, test subsets, or decoding settings and present the result as model quality. | |
| - Do not assume robustness to every Persian accent, code-switching pattern, recording channel, or adversarial acoustic condition. | |
| ### 📁 Repository contents | |
| This snapshot contains **11 files** totaling approximately **655.43 MB**. Common file groups: `.json` × 5, `no extension` × 2, `.md` × 1, `.pt` × 1, `.safetensors` × 1, `.txt` × 1. | |
| Largest or representative artifacts: | |
| - `model.safetensors` | |
| - `tokenizer.json` | |
| - `vocab.txt` | |
| - `head.pt` | |
| - `tokenizer_config.json` | |
| The repository card and `LICENSE` are part of the release. Runtime-specific configuration, tokenizer, vocabulary, metadata, and state files should be kept beside the main weights when present. | |
| ### 🚀 Download and integration | |
| ```python | |
| from transformers import AutoModelForTokenClassification, AutoTokenizer | |
| repo_id = "Reza2kn/ShenavaSanj-v1.0" | |
| tokenizer = AutoTokenizer.from_pretrained(repo_id) | |
| model = AutoModelForTokenClassification.from_pretrained(repo_id) | |
| inputs = tokenizer("این یک جملهٔ فارسی برای سنجش اهمیت واژههاست.", return_tensors="pt") | |
| scores = model(**inputs).logits | |
| ``` | |
| For NeMo checkpoints, restore through `nemo.collections.asr.models.ASRModel.restore_from(...)` rather than assuming a CTC-only class. For converted artifacts, follow the graph metadata and the runtime-specific notes retained later in this card. Validate one known clip against the source checkpoint before shipping a conversion. | |
| ### 📏 Evaluation | |
| Report at least WER and CER using the same Persian normalization rules, plus S³ when semantic importance matters. Shenava’s public [Triple Threat leaderboard](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) combines the Golden6669 and FLEURS-fa splits. Record the exact repository revision, decoder settings, chunk/context configuration, precision, device, and normalization code. | |
| Deployment exports should be checked for numerical and transcription parity against their parent repository, [ `Reza2kn/visualears-fastconformer-fa-full-ab` ](https://huggingface.co/Reza2kn/visualears-fastconformer-fa-full-ab). Runtime speed is hardware-specific; publish latency, real-time factor, warm-up policy, thread count, and audio duration together. | |
| ### ⚠️ Limitations and responsible use | |
| ASR quality varies with accent, age, speaking style, background noise, distance, clipping, reverberation, telephony bandwidth, overlapping speech, and code-switching. Persian orthography also permits multiple acceptable written forms. WER or CER can therefore penalize a semantically correct alternative, while a low aggregate score can still hide loss of a critical word. Review meaning-critical outputs and expose uncertainty in accessibility-facing products. | |
| ### 🔁 Reproducibility checklist | |
| 1. Pin the Hub revision and runtime/library versions. | |
| 2. Resample audio deterministically and document channel mixing. | |
| 3. Keep tokenizer and decoding assets from this repository together. | |
| 4. Record streaming chunk, left/right context, cache reset, and endpointing behavior. | |
| 5. Apply one documented Persian normalization/ITN pipeline to references and hypotheses. | |
| 6. Publish failed cases and condition-level results, not only a single average. | |
| ## 🇮🇷 مستندات فارسی | |
| ### 🧭 معرفی | |
| امتیازدهندهٔ اهمیت معنایی شنواسنج است. این مخزن یک مصنوع رسمی از انتشار Shenava-1 است و باید همراه با دادههای آموزشی، توکنایزر، روش رمزگشایی، نرمالسازی فارسی و تنظیمات اجرای جریانی دیده شود. پیوند مجموعه در بالای صفحه، نسخههای همخانواده، قالبهای استقرار، دادهها و معیارهای ارزیابی را یکجا نشان میدهد. | |
| هدف پروژه ارائهٔ زیرساخت باز و قابل بازتولید برای بازشناسی گفتار و زیرنویس فارسی است. نتیجهٔ نهایی فقط به وزن مدل وابسته نیست؛ نرخ نمونهبرداری، کانال صوت، وضعیت کش، روش رمزگشایی، تبدیل اعداد گفتاری و یکسانسازی نیمفاصله نیز بر خروجی اثر دارند. | |
| ### ✅ کاربردهای پیشنهادی | |
| - پژوهش، بنچمارک و مقایسهٔ منصفانهٔ ASR فارسی با پروتکل یکسان. | |
| - رونویسی آفلاین یا جریانی، در صورتی که قالب و زماناجرای انتخابی از آن پشتیبانی کند. | |
| - زیرنویس روی دستگاه، ابزارهای دسترسپذیری و رابطهای گفتاری فارسی. | |
| - تبدیل مدل و یکپارچهسازی با زماناجراهای مختلف همراه با آزمون برابری خروجی. | |
| ### 🚫 کاربردهای نامناسب | |
| - خروجی را در تصمیمهای پزشکی، حقوقی، اضطراری یا پرخطر بهعنوان سند قطعی به کار نبرید. | |
| - از خطا یا صدای کاربر برای استنباط هویت، سلامت، قومیت، نیت یا ویژگی حساس استفاده نکنید. | |
| - نتایجی را که با زیرمجموعه، نرمالسازی یا رمزگشایی متفاوت ساخته شدهاند مقایسهٔ مستقیم ننامید. | |
| - پوشش کامل همهٔ لهجهها، گفتار آمیخته، کانالها و شرایط صوتی را فرض نکنید. | |
| ### 📁 محتوای مخزن | |
| این نسخه شامل **11 فایل** با حجم تقریبی **655.43 MB** است. گروههای رایج فایل: `.json` × 5, `no extension` × 2, `.md` × 1, `.pt` × 1, `.safetensors` × 1, `.txt` × 1. | |
| فایلهای شاخص: | |
| - `model.safetensors` | |
| - `tokenizer.json` | |
| - `vocab.txt` | |
| - `head.pt` | |
| - `tokenizer_config.json` | |
| فایلهای توکنایزر، واژگان، پیکربندی، وضعیت جریانی و فراداده را در صورت وجود کنار وزن اصلی نگه دارید. | |
| ### 🚀 دریافت و استفاده | |
| برای استفاده، توکنایزر و مدل را از همان شناسهٔ مخزن بارگذاری کنید و logits هر توکن را به امتیاز اهمیت تبدیل کنید. نسخهٔ کتابخانه را در محیط بازتولید ثبت کنید. | |
| برای چکپوینت NeMo از `ASRModel.restore_from(...)` استفاده کنید و مدل را صرفاً CTC فرض نکنید. برای خروجیهای تبدیلشده، یک کلیپ مرجع را با مدل مبدأ مقایسه کنید و سپس استقرار را انجام دهید. | |
| ### 📏 ارزیابی | |
| حداقل WER و CER را با نرمالسازی فارسی یکسان گزارش کنید و در سناریوهای حساس به معنا، S³ را نیز بیاورید. در [جدول Triple Threat](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) دو بخش Golden6669 و FLEURS-fa با هم سنجیده میشوند. شناسهٔ دقیق نسخه، تنظیمات دیکودر، کانتکست، دقت عددی، سختافزار و کد نرمالسازی را ثبت کنید. | |
| ### ⚠️ محدودیتها و استفادهٔ مسئولانه | |
| لهجه، سن، سبک گفتار، نویز، فاصله، کلیپشدن، پژواک، کانال تلفنی، همپوشانی گویندگان و کدسوئیچینگ میتوانند کیفیت را تغییر دهند. چند نگارش فارسی ممکن است از نظر معنایی درست باشند، اما WER/CER یکی را خطا حساب کند. در محصولات دسترسپذیری، واژههای کلیدی را جداگانه بازبینی و عدم قطعیت را به کاربر نشان دهید. | |
| ### 🔁 چکلیست بازتولید | |
| ۱. نسخهٔ دقیق مخزن و کتابخانهها را ثابت کنید. ۲. تبدیل نرخ نمونه و کانال را مستند کنید. ۳. توکنایزر و داراییهای رمزگشایی همین مخزن را نگه دارید. ۴. اندازهٔ قطعه، کانتکست، بازنشانی کش و endpointing را ثبت کنید. ۵. یک خط لولهٔ نرمالسازی/ITN مشترک به مرجع و خروجی اعمال کنید. ۶. خطاهای نمونهای و نتایج هر شرایط را در کنار میانگین منتشر کنید. | |
| ## 📚 Citation, links, and license | استناد، پیوندها و مجوز | |
| - 🤗 [Shenava-1 collection](https://huggingface.co/collections/Reza2kn/shenava-10-open-streaming-persian-asr-and-captioning) | |
| - 🧰 [Project repository](https://github.com/Reza2kn/shenava-1) | |
| - 📄 [SLT paper submission](https://openreview.net/forum?id=QTa6ax9PU3) | |
| - 📊 [Persian ASR Triple Threat](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) | |
| ```bibtex | |
| @misc{shenava1_shenavasanj_v1_0, | |
| title = {ShenavaSanj-v1.0: a Shenava-1 Persian speech artifact}, | |
| author = {Reza2kn}, | |
| year = {2026}, | |
| url = {https://huggingface.co/Reza2kn/ShenavaSanj-v1.0} | |
| } | |
| ``` | |
| Released under the **Apache License 2.0**. این مخزن با **مجوز Apache 2.0** منتشر شده است. | |
| --- | |
| ## 📎 Retained technical notes | یادداشتهای فنی پیشین | |
| The pre-existing technical card is retained below for revision-specific commands, measurements, and artifact details. The bilingual sections above define the common Shenava-1 documentation contract. | |
| یادداشت فنی قبلی برای فرمانها، اندازهگیریها و جزئیات همان نسخه در ادامه حفظ شده است. بخشهای دوزبانهٔ بالا قرارداد مستندسازی مشترک Shenava-1 را تعریف میکنند. | |
| # ShenavaSanj v1.0 (شنواسنج) | |
| Persian **word-importance** model — scores how *important* each word is for understanding an | |
| utterance (0 = filler/function word, 1 = essential content). Built to power a **DHH-oriented | |
| semantic error metric** (ACE-style, importance-weighted WER) for Persian ASR — so a missed *keyword* | |
| is penalized far more than a missed *filler*. | |
| - **Student:** ParsBERT (`HooshvareLab/bert-base-parsbert-uncased`, 110M) + token-regression head. | |
| - **Teacher:** `google/gemma-4-31b-it` (fp4) with DHH-framed prompt + 5 Persian few-shot anchors. | |
| - **Distillation:** soft-label regression (HuberLoss, δ=0.1) on the teacher's continuous [0,1] scores; | |
| first-subword alignment (NER-style). **No human annotation.** | |
| - **Data:** 26,490 unique conversational Persian utterances from `shekar-ai/neyshekar-v4-persian-asr-fa`. | |
| Teacher labels: [`Reza2kn/neyshekar-fa-wimp-teacher-labels`](https://huggingface.co/datasets/Reza2kn/neyshekar-fa-wimp-teacher-labels). | |
| ## Validation | |
| - **Student vs teacher** (held-out 1,324): token-ρ **0.934**, per-utterance-ρ **0.916**, MSE 0.0114. | |
| - **Provenance:** the teacher was validated on English DHH gold (Kafle & Huenerfauth LREC-2018): | |
| pooled token-ρ ≈ 0.80 vs ~0.84 human inter-annotator ceiling. Persian cross-model agreement | |
| (teacher vs gemini-3.5-flash) per-utt ρ ≈ 0.89. | |
| ## Usage | |
| ```python | |
| import torch, torch.nn as nn | |
| from transformers import AutoTokenizer, AutoModel | |
| from huggingface_hub import hf_hub_download | |
| REPO = "Reza2kn/ShenavaSanj-v1.0" | |
| tok = AutoTokenizer.from_pretrained(REPO) | |
| enc = AutoModel.from_pretrained(REPO).eval() | |
| head = nn.Linear(enc.config.hidden_size, 1) | |
| head.load_state_dict(torch.load(hf_hub_download(REPO, "head.pt"), map_location="cpu")); head.eval() | |
| @torch.no_grad() | |
| def importance(text): | |
| words = text.split() | |
| e = tok(words, is_split_into_words=True, return_tensors="pt", truncation=True, max_length=64) | |
| p = torch.sigmoid(head(enc(**e).last_hidden_state)).squeeze(-1)[0].tolist() | |
| out, seen = [], set() | |
| for ti, wid in enumerate(e.word_ids(0)): | |
| if wid is not None and wid not in seen: | |
| seen.add(wid); out.append(round(p[ti], 3)) | |
| return list(zip(words, out)) | |
| print(importance("خب یعنی چی الان؟")) | |
| # [('خب', 0.02), ('یعنی', 0.28), ('چی', 0.83), ('الان؟', 0.71)] | |
| ``` | |
| Scores are per **whitespace token**. For the ACE-style weighted-WER metric, weight each | |
| reference-word error by its ShenavaSanj importance and normalize by total reference importance. | |