Genericize teacher reference in model card
Browse files
README.md
CHANGED
|
@@ -94,7 +94,7 @@ No real Hinglish speech was used for training. The corpus was manufactured, filt
|
|
| 94 |
in five steps.
|
| 95 |
|
| 96 |
### 1. Teacher
|
| 97 |
-
Speech was generated by **
|
| 98 |
code-switching. A human verified a sample of teacher clips as natural before distillation, so the
|
| 99 |
teacher defines the quality bar. Four Indian voices were used as the fixed voice set.
|
| 100 |
|
|
@@ -174,14 +174,14 @@ embedded English is recovered as English), UTMOS (naturalness), and resemblyzer
|
|
| 174 |
|
| 175 |
- **License: non-commercial.** This is a derivative of XTTS-v2 under the Coqui Public Model License
|
| 176 |
(CPML); it inherits CPML and is for research / non-commercial use. The training audio was
|
| 177 |
-
generated by
|
| 178 |
- **Four fixed voices** from one teacher engine; the voices are a fairly homogeneous family.
|
| 179 |
- **Accent is ~3% below the teacher** and sentence-initial English markers (e.g. "Wait") are
|
| 180 |
occasionally mispronounced.
|
| 181 |
- **Operational:** spell numbers as words (digits error out), chunk text over ~150 chars.
|
| 182 |
- **Absolute naturalness is not human-MOS-certified;** the evidence is relative parity to a
|
| 183 |
human-verified teacher, plus an English-biased UTMOS proxy.
|
| 184 |
-
- **Single teacher:** style and artifact diversity are bounded by
|
| 185 |
|
| 186 |
## Provenance and intended use
|
| 187 |
|
|
|
|
| 94 |
in five steps.
|
| 95 |
|
| 96 |
### 1. Teacher
|
| 97 |
+
Speech was generated by **the teacher TTS**, which does native intra-sentence Hindi-English
|
| 98 |
code-switching. A human verified a sample of teacher clips as natural before distillation, so the
|
| 99 |
teacher defines the quality bar. Four Indian voices were used as the fixed voice set.
|
| 100 |
|
|
|
|
| 174 |
|
| 175 |
- **License: non-commercial.** This is a derivative of XTTS-v2 under the Coqui Public Model License
|
| 176 |
(CPML); it inherits CPML and is for research / non-commercial use. The training audio was
|
| 177 |
+
generated by teacher TTS; review their terms before any redistribution or commercial use.
|
| 178 |
- **Four fixed voices** from one teacher engine; the voices are a fairly homogeneous family.
|
| 179 |
- **Accent is ~3% below the teacher** and sentence-initial English markers (e.g. "Wait") are
|
| 180 |
occasionally mispronounced.
|
| 181 |
- **Operational:** spell numbers as words (digits error out), chunk text over ~150 chars.
|
| 182 |
- **Absolute naturalness is not human-MOS-certified;** the evidence is relative parity to a
|
| 183 |
human-verified teacher, plus an English-biased UTMOS proxy.
|
| 184 |
+
- **Single teacher:** style and artifact diversity are bounded by teacher TTS.
|
| 185 |
|
| 186 |
## Provenance and intended use
|
| 187 |
|