--- license: cc-by-nc-4.0 language: - ar tags: - automatic-speech-recognition - wav2vec2 - ctc - quran - tajweed - phoneme-recognition - ipa base_model: quranic-orgz/omni-ctc-ipa-v7 pipeline_tag: automatic-speech-recognition --- # omni-ctc-ipa-v7-attraux-aux0p1-20260928 Wav2Vec2-CTC Quranic phoneme recognizer. Fine-tunes **`omni-ctc-ipa-v7`** with an **attribute auxiliary loss (AttrAux)** at weight **0.1**, part of the `attraux-v7` research roadmap (Stage 2 aux-weight dose sweep, seed 42, trained 2026-09-28). ## Model description - **Architecture**: `Wav2Vec2ForCTC`, 24 transformer layers, hidden size 1024, 71-token IPA-style phoneme vocabulary (Tajweed-aware: ghunna/tafkheem/gemination/qalqalah symbols included). - **Base checkpoint**: `omni-ctc-ipa-v7` (frozen feature encoder, transformer layers 0-11 frozen during this fine-tune, layers 12-23 trainable). - **Training data**: `hetchyy/everyayah-phonemes` (207,114 max train rows), 2,400 steps, batch size 1 x grad-accum 16, LR 1e-5, 240 warmup steps, fp16. - **AttrAux mechanism**: an auxiliary BCE loss over four pooled attribute channels (ghunna/tafkheem/gemination/qalqalah), derived by pooling the existing CTC logits (not a separate model head), weighted at **0.1** relative to the primary CTC loss, with attribute-group class rebalancing enabled (`--attr-aux-group-rebalance`). - **Best training-time eval PER**: 1.970% (`hetchyy/everyayah-phonemes` dev split). ## Insights This checkpoint is one of two survivors (out of 17 evaluated `attraux-v7` variants plus baselines) that beat the joint bar **pooled PER < 6.10% AND pooled tajweed macro-F1 > 97.70%** on a professional-reciter benchmark suite (quranmd + tadabur + mufti_malhan + quranlab, 187,417 pooled phoneme units): | Metric (pooled, 4 professional sets) | Value | |---|---| | PER | 6.09% | | Substitution / Deletion / Insertion | 2.45% / 1.05% / 2.59% | | Ghunna F1 | 97.16% | | Qalqalah F1 | 95.64% (best qalqalah score across all 17 evaluated models) | | Tafkheem F1 | 97.85% | | Gemination F1 | 98.01% | | Sifat-core F1 | 98.01% | | Macro F1 (all families) | 97.72% | **Caveat — read before use as evidence of AttrAux working:** the parent research roadmap (`docs/roadmaps/attraux-v7/ROADMAP.md`) evaluated this exact aux-weight sweep (w=0.05/0.1/0.4) against a matched no-aux control on the roadmap's primary metric — learner-correction precision (LCP) on a speaker-split QuranMB.v2 dev set with a paired bootstrap CI — and found **no arm, including this one, produced a statistically significant LCP gain over control** (Stage 2 verdict: REGRESSION/NO EFFECT for all three weights; overall roadmap KILL on 2026-09-29 after Stages 1-5' exhausted every tested AttrAux configuration). The pooled-PER/tajweed-F1 numbers above measure professional-recitation transcription quality, which this checkpoint is good at, but should **not** be read as evidence that its AttrAux training signal improved mispronunciation-detection precision — that hypothesis was tested and rejected on a different, dedicated benchmark. In short: a solid general-purpose Quranic phoneme transcriber, marginally ahead of its own base model (`omni-ctc-ipa-v7`: PER 6.37%, F1 97.50%) and of every other `attraux-v7` variant tried, but its AttrAux loss did not measurably improve mispronunciation-detection quality over a plain no-aux fine-tune at the same recipe. ## Intended use Quranic recitation phoneme transcription (near-offline/offline ASR) and Tajweed rule classification (ghunna/tafkheem/gemination/qalqalah) from audio. Not validated as a mispronunciation-detection/correction system — use `omni-ctc-ipa-v7` or a control checkpoint from the same sweep for that use case pending further evidence. ## Training hyperparameters ``` dataset: hetchyy/everyayah-phonemes base_model: omni-ctc-ipa-v7 max_steps: 2400 batch_size: 1, grad_accum: 16 learning_rate: 1e-5, warmup_steps: 240 freeze_transformer_layers: 12 (of 24) attr_aux_weight: 0.1 attr_aux_warmup_steps: 240 attr_aux_group_rebalance: true seed: 42 fp16: true ```