--- language: en license: mit tags: - pronunciation-assessment - phoneme-scoring - wavlm - speech - children-speech base_model: microsoft/wavlm-large pipeline_tag: audio-classification --- # WavLM Phoneme Scorer Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English pronunciation assessment, trained on 11K children's speech samples. Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors. ## Quick Start ### Install ```bash pip install torch torchaudio transformers g2p-en huggingface_hub ``` ### Python API ```python from pipeline_v2 import PronunciationAssessorV2 # Auto-download model from HuggingFace assessor = PronunciationAssessorV2.from_pretrained() result = assessor.assess("audio.mp3", "Hello, Peter.") print(result["overall_score"]) # 85.2 print(result["n_errors"]) # 0 for word in result["words"]: print(f"{word['word']}: score={word['score']}") for ph in word["phonemes"]: err = " <- ERROR" if ph["error"] else "" print(f" /{ph['phone']}/ score={ph['score']} pherr={ph['pherr_prob']:.2f}{err}") ``` ### CLI ```bash # Model downloads automatically on first run python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter." ``` Output: ``` ============================================================ Text: "Hello, Peter." Overall Score: 85.2/100 (errors: 0/8) ============================================================ ✓ Hello score= 87.7 errors=0/4 /hh / score= 98.6 GOP= -0.97 pherr=0.05 /ah / score= 73.1 GOP= -7.40 pherr=0.43 /l / score= 88.6 GOP= +4.00 pherr=0.29 /ow / score= 90.7 GOP= -6.05 pherr=0.13 ✓ Peter score= 82.6 errors=0/4 /p / score= 95.7 GOP= +5.40 pherr=0.08 /iy / score= 90.2 GOP= +3.70 pherr=0.12 /t / score= 72.5 GOP= +0.50 pherr=0.55 /er / score= 71.8 GOP= -1.40 pherr=0.61 ``` ### Download Model Manually ```bash # Via huggingface-cli huggingface-cli download Jianshu001/wavlm-phoneme-scorer wavlm_finetuned.pt --local-dir . # Via Python from huggingface_hub import hf_hub_download hf_hub_download(repo_id="Jianshu001/wavlm-phoneme-scorer", filename="wavlm_finetuned.pt", local_dir=".") # Then use with local path assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt") ``` ## Architecture ``` Reference Text ──→ G2P ──→ Expected phoneme sequence │ Audio ──→ CTC model ──→ Viterbi Forced Alignment ──→ Frame segments │ │ └──→ WavLM-Large (fine-tuned) ──→ Hidden states ──→ Pool per segment │ + phone embedding (32d) + GOP score (1d) + n_frames (1d) │ MLP (1058 → 512 → 512 → 256) ├── score_head → phoneme score (0-100) └── pherr_head → error probability (0-1) ``` - **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total) - **Alignment**: `facebook/wav2vec2-xlsr-53-espeak-cv-ft` (frozen, CTC-based Viterbi forced alignment) - **Scoring head**: MLP with BatchNorm, GELU, Dropout, multi-task (score regression + error classification) ## Performance Evaluated on test set (8062 phonemes from 1727 audio files, children's speech): | Metric | GOP Baseline (v1.0) | This Model | |--------|-------------------|------------| | Phoneme Error AUC-ROC | 0.738 | **0.870** | | Phoneme Error F1 | 0.476 | **0.595** | | Phoneme Error Precision | 0.379 | **0.592** | | Phoneme Error Recall | 0.638 | **0.598** | | Phone Score Pearson | 0.372 | **0.645** | | Phone Score MAE | 27.44 | **16.47** | ## Training Data - 11,601 audio recordings of English learners (children) - 53,926 phonemes with professional human evaluation labels - Labels include per-phoneme scores (0-100) and error flags (pherr 0/1) ## Training Details - Fine-tuned top 6 of 24 WavLM transformer layers - Differential learning rate: backbone 1e-5, head 5e-4 - AdamW optimizer, cosine annealing, early stopping (patience=8) - Gradient accumulation (4 steps), gradient clipping (max_norm=1.0) - Batch size 64, trained for 24 epochs (early stopped) - Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error) - Train/Val/Test split by audio file: 40K/5K/8K phonemes ## Files | File | Description | Size | |------|-------------|------| | `wavlm_finetuned.pt` | Full checkpoint (backbone + head state dict) | 1.2GB | | `pipeline_v2.py` | Inference pipeline with `from_pretrained()` support | 18KB | | `finetune_wavlm.py` | Training script (reproducing the fine-tuning) | 25KB | ## Full Repository See the complete project (data, evaluation, all experiments) at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction) (branch: `feature/wavlm-pipeline`)