Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
license: mit
|
| 4 |
+
tags:
|
| 5 |
+
- pronunciation-assessment
|
| 6 |
+
- phoneme-scoring
|
| 7 |
+
- wavlm
|
| 8 |
+
- speech
|
| 9 |
+
- children-speech
|
| 10 |
+
base_model: microsoft/wavlm-large
|
| 11 |
+
pipeline_tag: audio-classification
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# WavLM Phoneme Scorer
|
| 15 |
+
|
| 16 |
+
Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English pronunciation assessment, trained on 11K children's speech samples.
|
| 17 |
+
|
| 18 |
+
Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
|
| 19 |
+
|
| 20 |
+
## Architecture
|
| 21 |
+
|
| 22 |
+
```
|
| 23 |
+
Audio βββ CTC model (alignment) βββ Frame-level phoneme segments
|
| 24 |
+
β
|
| 25 |
+
ββββ WavLM-Large (fine-tuned top 6 layers) βββ Hidden states per segment
|
| 26 |
+
β
|
| 27 |
+
+ phone embedding (32d)
|
| 28 |
+
+ GOP score (1d)
|
| 29 |
+
+ n_frames (1d)
|
| 30 |
+
β
|
| 31 |
+
MLP (1058 β 512 β 512 β 256)
|
| 32 |
+
β
|
| 33 |
+
βββ score_head β phoneme score (0-100)
|
| 34 |
+
βββ pherr_head β error probability (0-1)
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
- **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
|
| 38 |
+
- **Alignment**: `facebook/wav2vec2-xlsr-53-espeak-cv-ft` (frozen, CTC-based Viterbi forced alignment)
|
| 39 |
+
- **Scoring head**: MLP with BatchNorm, GELU, Dropout, multi-task (score regression + error classification)
|
| 40 |
+
|
| 41 |
+
## Performance
|
| 42 |
+
|
| 43 |
+
Evaluated on test set (8062 phonemes from 1727 audio files):
|
| 44 |
+
|
| 45 |
+
| Metric | GOP Baseline (v1.0) | This Model |
|
| 46 |
+
|--------|-------------------|------------|
|
| 47 |
+
| Phoneme Error AUC-ROC | 0.738 | **0.870** |
|
| 48 |
+
| Phoneme Error F1 | 0.476 | **0.595** |
|
| 49 |
+
| Phoneme Error Precision | 0.379 | **0.592** |
|
| 50 |
+
| Phone Score Pearson | 0.372 | **0.645** |
|
| 51 |
+
| Phone Score MAE | 27.44 | **16.47** |
|
| 52 |
+
|
| 53 |
+
## Usage
|
| 54 |
+
|
| 55 |
+
```python
|
| 56 |
+
from pipeline_v2 import PronunciationAssessorV2
|
| 57 |
+
|
| 58 |
+
assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
|
| 59 |
+
result = assessor.assess("audio.mp3", "Hello, Peter.")
|
| 60 |
+
|
| 61 |
+
# result["overall_score"] β 82.3
|
| 62 |
+
# result["words"][0]["phonemes"][0]["error"] β False
|
| 63 |
+
# result["words"][0]["phonemes"][0]["pherr_prob"] β 0.05
|
| 64 |
+
```
|
| 65 |
+
|
| 66 |
+
CLI:
|
| 67 |
+
```bash
|
| 68 |
+
python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
|
| 69 |
+
```
|
| 70 |
+
|
| 71 |
+
See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction)
|
| 72 |
+
|
| 73 |
+
## Training Data
|
| 74 |
+
|
| 75 |
+
- 11,601 audio recordings of English learners (children)
|
| 76 |
+
- 53,926 phonemes with professional human evaluation labels
|
| 77 |
+
- Labels include per-phoneme scores (0-100) and error flags
|
| 78 |
+
|
| 79 |
+
## Training Details
|
| 80 |
+
|
| 81 |
+
- Fine-tuned top 6 of 24 WavLM transformer layers
|
| 82 |
+
- Differential learning rate: backbone 1e-5, head 5e-4
|
| 83 |
+
- AdamW optimizer, cosine annealing, early stopping (patience=8)
|
| 84 |
+
- Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
|
| 85 |
+
- Batch size 64, trained for 24 epochs (early stopped)
|
| 86 |
+
- Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
|
| 87 |
+
|
| 88 |
+
## Files
|
| 89 |
+
|
| 90 |
+
- `wavlm_finetuned.pt` β Full checkpoint (backbone + head state dict, 1.2GB)
|
| 91 |
+
- `pipeline_v2.py` β Inference pipeline
|
| 92 |
+
- `finetune_wavlm.py` β Training script
|