Jianshu001 commited on
Commit
bdae6d5
Β·
verified Β·
1 Parent(s): 7b9b23e

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +92 -0
README.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ license: mit
4
+ tags:
5
+ - pronunciation-assessment
6
+ - phoneme-scoring
7
+ - wavlm
8
+ - speech
9
+ - children-speech
10
+ base_model: microsoft/wavlm-large
11
+ pipeline_tag: audio-classification
12
+ ---
13
+
14
+ # WavLM Phoneme Scorer
15
+
16
+ Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English pronunciation assessment, trained on 11K children's speech samples.
17
+
18
+ Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
19
+
20
+ ## Architecture
21
+
22
+ ```
23
+ Audio ──→ CTC model (alignment) ──→ Frame-level phoneme segments
24
+ β”‚
25
+ └──→ WavLM-Large (fine-tuned top 6 layers) ──→ Hidden states per segment
26
+ β”‚
27
+ + phone embedding (32d)
28
+ + GOP score (1d)
29
+ + n_frames (1d)
30
+ β”‚
31
+ MLP (1058 β†’ 512 β†’ 512 β†’ 256)
32
+ β”‚
33
+ β”œβ”€β”€ score_head β†’ phoneme score (0-100)
34
+ └── pherr_head β†’ error probability (0-1)
35
+ ```
36
+
37
+ - **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
38
+ - **Alignment**: `facebook/wav2vec2-xlsr-53-espeak-cv-ft` (frozen, CTC-based Viterbi forced alignment)
39
+ - **Scoring head**: MLP with BatchNorm, GELU, Dropout, multi-task (score regression + error classification)
40
+
41
+ ## Performance
42
+
43
+ Evaluated on test set (8062 phonemes from 1727 audio files):
44
+
45
+ | Metric | GOP Baseline (v1.0) | This Model |
46
+ |--------|-------------------|------------|
47
+ | Phoneme Error AUC-ROC | 0.738 | **0.870** |
48
+ | Phoneme Error F1 | 0.476 | **0.595** |
49
+ | Phoneme Error Precision | 0.379 | **0.592** |
50
+ | Phone Score Pearson | 0.372 | **0.645** |
51
+ | Phone Score MAE | 27.44 | **16.47** |
52
+
53
+ ## Usage
54
+
55
+ ```python
56
+ from pipeline_v2 import PronunciationAssessorV2
57
+
58
+ assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
59
+ result = assessor.assess("audio.mp3", "Hello, Peter.")
60
+
61
+ # result["overall_score"] β†’ 82.3
62
+ # result["words"][0]["phonemes"][0]["error"] β†’ False
63
+ # result["words"][0]["phonemes"][0]["pherr_prob"] β†’ 0.05
64
+ ```
65
+
66
+ CLI:
67
+ ```bash
68
+ python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
69
+ ```
70
+
71
+ See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction)
72
+
73
+ ## Training Data
74
+
75
+ - 11,601 audio recordings of English learners (children)
76
+ - 53,926 phonemes with professional human evaluation labels
77
+ - Labels include per-phoneme scores (0-100) and error flags
78
+
79
+ ## Training Details
80
+
81
+ - Fine-tuned top 6 of 24 WavLM transformer layers
82
+ - Differential learning rate: backbone 1e-5, head 5e-4
83
+ - AdamW optimizer, cosine annealing, early stopping (patience=8)
84
+ - Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
85
+ - Batch size 64, trained for 24 epochs (early stopped)
86
+ - Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
87
+
88
+ ## Files
89
+
90
+ - `wavlm_finetuned.pt` β€” Full checkpoint (backbone + head state dict, 1.2GB)
91
+ - `pipeline_v2.py` β€” Inference pipeline
92
+ - `finetune_wavlm.py` β€” Training script