File size: 1,966 Bytes
1352e38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
---
license: other
language:
  - en
tags:
  - text-to-speech
  - xtts
  - english
  - nigerian-pidgin
  - pidgin
  - coqui
library_name: coqui-tts
---

# XTTS-v2 — English + Nigerian Pidgin (lr=5e-5, 10 epochs, ~35 min data)

Fine-tuned [Coqui XTTS-v2](https://huggingface.co/coqui/XTTS-v2) for English and Nigerian Pidgin
(both synthesized with language id `en`), trained on `vaghawan/en-pidgin-high-similarity-11-35mins` (speaker `en_pidgin_voice`).

Pidgin has no native XTTS language code — train and infer Pidgin text as `en`.

## Files

| File | Purpose |
|------|---------|
| `best_model.pth` | Fine-tuned GPT checkpoint (epoch 10, lr=5e-5) |
| `config.json` | Model config |
| `vocab.json` | Stock XTTS English BPE vocabulary (no extend) |
| `references/en_pidgin_voice.wav` | Speaker reference |
| `infer_english_and_pidgin.py` | English + Pidgin sample inference |
| `infer.py` | Shared inference helpers |
| `xtts_hausa_patch.py` | Runtime language helpers (also used for `en`) |
| `env_config.py` | Config loader |
| `config.env.example` | Example settings (copy to `config.env`) |
| `requirements.txt` | Python dependencies |

## Quick start

```bash
pip install -r requirements.txt
cp config.env.example config.env

python infer_english_and_pidgin.py \
  --model-dir . \
  --speaker-wav references/en_pidgin_voice.wav \
  --out-dir outputs/samples_en_pidgin

# Single line (via infer.py):
python infer.py --model-dir . --language en \
  --text "How far? I dey fine." \
  --speaker-wav references/en_pidgin_voice.wav \
  --out outputs/out.wav
```

## Training details

- Base model: Coqui XTTS-v2 (stock English vocab — no `extend_vocab.py`)
- Epochs: 10
- Learning rate: 5e-5
- Language id: `en` (English and Pidgin)
- Speakers: `en_pidgin_voice`
- Data: ~35 minutes high-similarity English/Pidgin

## License

XTTS-v2 uses the [Coqui Public Model License](https://huggingface.co/coqui/XTTS-v2).
Check each dataset license before redistribution.