Update README.md
Browse files
README.md
CHANGED
|
@@ -67,19 +67,4 @@ Analysis runs at 100 frames/s (16 kHz, hop 160).
|
|
| 67 |
- **Pitch** — 128 mel-spaced bins; bin 0 is unvoiced.
|
| 68 |
- **Loudness** — 64 A-weighted dB bins, about 1.05 bins per dB.
|
| 69 |
- **Duration** — per-phoneme frame counts, at most 191 frames (1.91 s) per
|
| 70 |
-
phoneme; an edited timeline is capped at 2001 frames (20 s).
|
| 71 |
-
|
| 72 |
-
## Training data
|
| 73 |
-
|
| 74 |
-
English speech from Emilia, GigaSpeech, LibriTTS / LibriTTS-R, HiFi-TTS and
|
| 75 |
-
LibriLight, with forced-aligned ARPAbet phoneme annotations.
|
| 76 |
-
|
| 77 |
-
## Licence and intended use
|
| 78 |
-
|
| 79 |
-
**CC BY-NC 4.0 — non-commercial.** The training mix includes Emilia, which is
|
| 80 |
-
distributed under CC BY-NC 4.0, so these weights inherit that restriction. The
|
| 81 |
-
source code is MIT.
|
| 82 |
-
|
| 83 |
-
These models clone a speaker's voice from a few seconds of reference audio.
|
| 84 |
-
Only use recordings you have the right to use, disclose synthetic speech as
|
| 85 |
-
synthetic, and do not use this to impersonate anyone without their consent.
|
|
|
|
| 67 |
- **Pitch** — 128 mel-spaced bins; bin 0 is unvoiced.
|
| 68 |
- **Loudness** — 64 A-weighted dB bins, about 1.05 bins per dB.
|
| 69 |
- **Duration** — per-phoneme frame counts, at most 191 frames (1.91 s) per
|
| 70 |
+
phoneme; an edited timeline is capped at 2001 frames (20 s).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|