poto-tts — Ghanaian English on Kokoro
English speech that pronounces Ghanaian words properly: Kwabena, Achimota, the Okuapenhene, Nyankpani, Yaw, cedi, banku.
pip install poto-tts
poto-tts "Kwabena went to Achimota" -o out.wav
from poto_tts import load
tts = load() # pulls this repo on first use
tts.save("The Okuapenhene met Nana Bawumia", "out.wav")
Hear it against standard Kokoro: the Space.
What it does, and what it does not
Not Ghanaian-accented voices. These are Kokoro's British speakers. What changes is what they say, not who they sound like.
standard Kokoro poto-tts
Kwabena kwˈAbnə kwabˈɪna
Achimota əʧɪmˈOTə ˌatʃimˈota
Okuapenhene ˈOkjuˌApənhˌin ˌokwapɛnhˈɛnɛ
Yaw jˈɔ jˈaw
Ewe jˈuː ˈɛvɛ
Ewe is the clearest case: standard Kokoro reads it as the English word you.
How
Pronunciation is data, not code. 44,321 Ghanaian words — names, places, titles, Twi and Ga loans, food, money, everyday coinage — are compiled into espeak's own dictionary, each carrying the lexicon's IPA, and read as British English. Your text reaches the model unchanged, and every other word is pronounced as British English would.
Two things contribute, and they compose. Standard Kokoro phonemises with
misaki, whose English lexicon holds no Ghanaian
words at all — it returns a placeholder for Kwabena, Achimota, Okuapenhene and
Akple. sherpa-onnx phonemises with espeak-ng, which at least reads Ghanaian spelling
as spelling. The dictionary is then where espeak's remaining guesses are replaced by the
recorded pronunciation, and where you add a name it has never seen.
All seven Akan vowels reach the model: ɛ and ɔ stay distinct from e and o.
On a 400-word sample checked against the lexicon, standard Kokoro says 2.8% of Ghanaian words correctly and poto-tts 98.5%.
Voices
Eight British speakers. sid is what sherpa-onnx wants.
| female | male | |
|---|---|---|
| name | Grace, Comfort, Mercy, Patience | Emmanuel, Isaac, Ebenezer, Bright |
| sid | 20, 21, 22, 23 | 26, 27, 24, 25 |
British only: the entries are read with espeak's British phoneme table, so the phonemes
are non-rhotic and use /a/ where American English has /æ/. Kokoro's twenty American
speakers were trained on American phonemes; they are still in the model and reachable
by their own names (load(voice="af_heart")) but not offered.
Output is 24 kHz mono.
Without Python — Android, iOS, WebAssembly, C++
Any sherpa-onnx runtime loads these files directly, offline. Send plain text with
lang=en.
| file | |
|---|---|
onnx/model.onnx |
the model |
voices.bin |
speaker embeddings |
tokens.txt |
phoneme → id |
espeak-ng-data/ |
the Ghanaian part. Ship this one. |
A stock espeak-ng-data gives you a working voice that mispronounces every Ghanaian
name, silently. That directory is the deliverable.
Integration guide, with Kotlin, Swift and C++ snippets, the speaker ids, and how to trim
espeak-ng-data from 28 MB to 2.3 MB:
docs/MOBILE.md.
Adding a pronunciation
The lexicon will miss somebody's name. Put it in a TSV and rebuild the dictionary:
pip install 'poto-tts[lexicon]'
printf 'Owusu\to w u s u\n' > my_words.tsv
poto-tts dict --out espeak-ng-data --ghanaian-stress --extra my_words.tsv
Licensing
- Kokoro weights: Apache-2.0 (hexgrad/Kokoro-82M, converted by k2-fsa). Commercial use is fine.
espeak-ng-data: GPL-3.0, as in any espeak-based deployment. Worth reading before an app-store submission — see the mobile guide.- Pronunciations derive from ghana-english-g2p (MIT).
Method and source: GhanaNLP/poto-tts.
- Downloads last month
- 47