omni-ctc-ipa-v7-attraux-aux0p4-20260928

Wav2Vec2-CTC Quranic phoneme recognizer. Fine-tunes omni-ctc-ipa-v7 with an attribute auxiliary loss (AttrAux) at weight 0.4, part of the attraux-v7 research roadmap (Stage 2 aux-weight dose sweep, seed 42, trained 2026-09-28).

Model description

  • Architecture: Wav2Vec2ForCTC, 24 transformer layers, hidden size 1024, 71-token IPA-style phoneme vocabulary (Tajweed-aware: ghunna/tafkheem/gemination/qalqalah symbols included).
  • Base checkpoint: omni-ctc-ipa-v7 (frozen feature encoder, transformer layers 0-11 frozen during this fine-tune, layers 12-23 trainable).
  • Training data: hetchyy/everyayah-phonemes (207,114 max train rows), 2,400 steps, batch size 1 x grad-accum 16, LR 1e-5, 240 warmup steps, fp16.
  • AttrAux mechanism: an auxiliary BCE loss over four pooled attribute channels (ghunna/tafkheem/gemination/qalqalah), derived by pooling the existing CTC logits (not a separate model head), weighted at 0.4 relative to the primary CTC loss β€” the highest dose tested in this sweep β€” with attribute-group class rebalancing enabled (--attr-aux-group-rebalance).
  • Best training-time eval PER: 1.787% (hetchyy/everyayah-phonemes dev split; lowest of the three aux-weight doses tried).

Insights

This checkpoint is one of two survivors (out of 17 evaluated attraux-v7 variants plus baselines) that beat the joint bar pooled PER < 6.10% AND pooled tajweed macro-F1 > 97.70% on a professional-reciter benchmark suite (quranmd + tadabur + mufti_malhan + quranlab, 187,417 pooled phoneme units) β€” and is the best-scoring attraux-v7 checkpoint on both axes among all 17 models benchmarked:

Metric (pooled, 4 professional sets) Value
PER 6.07%
Substitution / Deletion / Insertion 2.38% / 1.09% / 2.60%
Ghunna F1 96.99%
Qalqalah F1 95.73% (best qalqalah score across all 17 evaluated models)
Tafkheem F1 97.93% (best tafkheem score across all 17 evaluated models)
Gemination F1 98.06%
Sifat-core F1 98.04% (best sifat-core score across all 17 evaluated models)
Macro F1 (all families) 97.74% (best of all 17 evaluated models)

Caveat β€” read before use as evidence of AttrAux working: the parent research roadmap (docs/roadmaps/attraux-v7/ROADMAP.md) evaluated this exact aux-weight sweep (w=0.05/0.1/0.4) against a matched no-aux control on the roadmap's primary metric β€” learner-correction precision (LCP) on a speaker-split QuranMB.v2 dev set with a paired bootstrap CI β€” and found no arm, including this one at the highest tested dose, produced a statistically significant LCP gain over control (Stage 2 verdict: REGRESSION/NO EFFECT for all three weights; overall roadmap KILL on 2026-09-29 after Stages 1-5' exhausted every tested AttrAux configuration, including hard-negative frame weighting, density oversampling, and FP-mined replay). The pooled-PER/tajweed-F1 numbers above measure professional-recitation transcription quality β€” where this is the strongest checkpoint tried β€” but should not be read as evidence that a higher AttrAux dose improved mispronunciation-detection precision; that hypothesis was tested directly and rejected.

In short: the best general-purpose Quranic phoneme transcriber found in this sweep, ahead of its own base model (omni-ctc-ipa-v7: PER 6.37%, F1 97.50%) and every other attraux-v7/legacy checkpoint benchmarked, but its higher AttrAux loss weight did not measurably improve mispronunciation-detection quality over a plain no-aux fine-tune at the same recipe.

Intended use

Quranic recitation phoneme transcription (near-offline/offline ASR) and Tajweed rule classification (ghunna/tafkheem/gemination/qalqalah) from audio. Not validated as a mispronunciation-detection/correction system β€” use omni-ctc-ipa-v7 or a control checkpoint from the same sweep for that use case pending further evidence.

Training hyperparameters

dataset: hetchyy/everyayah-phonemes
base_model: omni-ctc-ipa-v7
max_steps: 2400
batch_size: 1, grad_accum: 16
learning_rate: 1e-5, warmup_steps: 240
freeze_transformer_layers: 12 (of 24)
attr_aux_weight: 0.4
attr_aux_warmup_steps: 240
attr_aux_group_rebalance: true
seed: 42
fp16: true
Downloads last month
39
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support