omni-ctc-ipa-v7-attraux-aux0p1-20260928
Wav2Vec2-CTC Quranic phoneme recognizer. Fine-tunes omni-ctc-ipa-v7 with an
attribute auxiliary loss (AttrAux) at weight 0.1, part of the attraux-v7
research roadmap (Stage 2 aux-weight dose sweep, seed 42, trained 2026-09-28).
Model description
- Architecture:
Wav2Vec2ForCTC, 24 transformer layers, hidden size 1024, 71-token IPA-style phoneme vocabulary (Tajweed-aware: ghunna/tafkheem/gemination/qalqalah symbols included). - Base checkpoint:
omni-ctc-ipa-v7(frozen feature encoder, transformer layers 0-11 frozen during this fine-tune, layers 12-23 trainable). - Training data:
hetchyy/everyayah-phonemes(207,114 max train rows), 2,400 steps, batch size 1 x grad-accum 16, LR 1e-5, 240 warmup steps, fp16. - AttrAux mechanism: an auxiliary BCE loss over four pooled attribute channels
(ghunna/tafkheem/gemination/qalqalah), derived by pooling the existing CTC logits
(not a separate model head), weighted at 0.1 relative to the primary CTC loss,
with attribute-group class rebalancing enabled (
--attr-aux-group-rebalance). - Best training-time eval PER: 1.970% (
hetchyy/everyayah-phonemesdev split).
Insights
This checkpoint is one of two survivors (out of 17 evaluated attraux-v7 variants
plus baselines) that beat the joint bar pooled PER < 6.10% AND pooled tajweed
macro-F1 > 97.70% on a professional-reciter benchmark suite (quranmd + tadabur +
mufti_malhan + quranlab, 187,417 pooled phoneme units):
| Metric (pooled, 4 professional sets) | Value |
|---|---|
| PER | 6.09% |
| Substitution / Deletion / Insertion | 2.45% / 1.05% / 2.59% |
| Ghunna F1 | 97.16% |
| Qalqalah F1 | 95.64% (best qalqalah score across all 17 evaluated models) |
| Tafkheem F1 | 97.85% |
| Gemination F1 | 98.01% |
| Sifat-core F1 | 98.01% |
| Macro F1 (all families) | 97.72% |
Caveat โ read before use as evidence of AttrAux working: the parent research
roadmap (docs/roadmaps/attraux-v7/ROADMAP.md) evaluated this exact aux-weight sweep
(w=0.05/0.1/0.4) against a matched no-aux control on the roadmap's primary metric โ
learner-correction precision (LCP) on a speaker-split QuranMB.v2 dev set with a paired
bootstrap CI โ and found no arm, including this one, produced a statistically
significant LCP gain over control (Stage 2 verdict: REGRESSION/NO EFFECT for all
three weights; overall roadmap KILL on 2026-09-29 after Stages 1-5' exhausted every
tested AttrAux configuration). The pooled-PER/tajweed-F1 numbers above measure
professional-recitation transcription quality, which this checkpoint is good at, but
should not be read as evidence that its AttrAux training signal improved
mispronunciation-detection precision โ that hypothesis was tested and rejected on a
different, dedicated benchmark.
In short: a solid general-purpose Quranic phoneme transcriber, marginally ahead of its
own base model (omni-ctc-ipa-v7: PER 6.37%, F1 97.50%) and of every other
attraux-v7 variant tried, but its AttrAux loss did not measurably improve
mispronunciation-detection quality over a plain no-aux fine-tune at the same recipe.
Intended use
Quranic recitation phoneme transcription (near-offline/offline ASR) and Tajweed rule
classification (ghunna/tafkheem/gemination/qalqalah) from audio. Not validated as a
mispronunciation-detection/correction system โ use omni-ctc-ipa-v7 or a control
checkpoint from the same sweep for that use case pending further evidence.
Training hyperparameters
dataset: hetchyy/everyayah-phonemes
base_model: omni-ctc-ipa-v7
max_steps: 2400
batch_size: 1, grad_accum: 16
learning_rate: 1e-5, warmup_steps: 240
freeze_transformer_layers: 12 (of 24)
attr_aux_weight: 0.1
attr_aux_warmup_steps: 240
attr_aux_group_rebalance: true
seed: 42
fp16: true
- Downloads last month
- 14