Jianshu001's picture
Upload README.md with huggingface_hub
643bd25 verified
|
Raw History Blame Contribute Delete
5.36 kB
---
language: en
license: mit
tags:
- pronunciation-assessment
- phoneme-scoring
- wavlm
- speech
- children-speech
base_model: microsoft/wavlm-large
pipeline_tag: audio-classification
---
# WavLM Phoneme Scorer
Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English pronunciation assessment, trained on 11K children's speech samples.
Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
## Quick Start
### Install
```bash
pip install torch torchaudio transformers g2p-en huggingface_hub
```
### Python API
```python
from pipeline_v2 import PronunciationAssessorV2
# Auto-download model from HuggingFace
assessor = PronunciationAssessorV2.from_pretrained()
result = assessor.assess("audio.mp3", "Hello, Peter.")
print(result["overall_score"]) # 85.2
print(result["n_errors"]) # 0
for word in result["words"]:
print(f"{word['word']}: score={word['score']}")
for ph in word["phonemes"]:
err = " <- ERROR" if ph["error"] else ""
print(f" /{ph['phone']}/ score={ph['score']} pherr={ph['pherr_prob']:.2f}{err}")
```
### CLI
```bash
# Model downloads automatically on first run
python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
```
Output:
```
============================================================
Text: "Hello, Peter."
Overall Score: 85.2/100 (errors: 0/8)
============================================================
βœ“ Hello score= 87.7 errors=0/4
/hh / score= 98.6 GOP= -0.97 pherr=0.05
/ah / score= 73.1 GOP= -7.40 pherr=0.43
/l / score= 88.6 GOP= +4.00 pherr=0.29
/ow / score= 90.7 GOP= -6.05 pherr=0.13
βœ“ Peter score= 82.6 errors=0/4
/p / score= 95.7 GOP= +5.40 pherr=0.08
/iy / score= 90.2 GOP= +3.70 pherr=0.12
/t / score= 72.5 GOP= +0.50 pherr=0.55
/er / score= 71.8 GOP= -1.40 pherr=0.61
```
### Download Model Manually
```bash
# Via huggingface-cli
huggingface-cli download Jianshu001/wavlm-phoneme-scorer wavlm_finetuned.pt --local-dir .
# Via Python
from huggingface_hub import hf_hub_download
hf_hub_download(repo_id="Jianshu001/wavlm-phoneme-scorer", filename="wavlm_finetuned.pt", local_dir=".")
# Then use with local path
assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
```
## Architecture
```
Reference Text ──→ G2P ──→ Expected phoneme sequence
β”‚
Audio ──→ CTC model ──→ Viterbi Forced Alignment ──→ Frame segments
β”‚ β”‚
└──→ WavLM-Large (fine-tuned) ──→ Hidden states ──→ Pool per segment
β”‚
+ phone embedding (32d)
+ GOP score (1d)
+ n_frames (1d)
β”‚
MLP (1058 β†’ 512 β†’ 512 β†’ 256)
β”œβ”€β”€ score_head β†’ phoneme score (0-100)
└── pherr_head β†’ error probability (0-1)
```
- **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
- **Alignment**: `facebook/wav2vec2-xlsr-53-espeak-cv-ft` (frozen, CTC-based Viterbi forced alignment)
- **Scoring head**: MLP with BatchNorm, GELU, Dropout, multi-task (score regression + error classification)
## Performance
Evaluated on test set (8062 phonemes from 1727 audio files, children's speech):
| Metric | GOP Baseline (v1.0) | This Model |
|--------|-------------------|------------|
| Phoneme Error AUC-ROC | 0.738 | **0.870** |
| Phoneme Error F1 | 0.476 | **0.595** |
| Phoneme Error Precision | 0.379 | **0.592** |
| Phoneme Error Recall | 0.638 | **0.598** |
| Phone Score Pearson | 0.372 | **0.645** |
| Phone Score MAE | 27.44 | **16.47** |
## Training Data
- 11,601 audio recordings of English learners (children)
- 53,926 phonemes with professional human evaluation labels
- Labels include per-phoneme scores (0-100) and error flags (pherr 0/1)
## Training Details
- Fine-tuned top 6 of 24 WavLM transformer layers
- Differential learning rate: backbone 1e-5, head 5e-4
- AdamW optimizer, cosine annealing, early stopping (patience=8)
- Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
- Batch size 64, trained for 24 epochs (early stopped)
- Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
- Train/Val/Test split by audio file: 40K/5K/8K phonemes
## Files
| File | Description | Size |
|------|-------------|------|
| `wavlm_finetuned.pt` | Full checkpoint (backbone + head state dict) | 1.2GB |
| `pipeline_v2.py` | Inference pipeline with `from_pretrained()` support | 18KB |
| `finetune_wavlm.py` | Training script (reproducing the fine-tuning) | 25KB |
## Full Repository
See the complete project (data, evaluation, all experiments) at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction) (branch: `feature/wavlm-pipeline`)