Kokoro v1.0: synthesize.cpp GGUF

GGUF conversions of the official Kokoro-82M v1.0 checkpoint for synthesize.cpp.

Ported from hexgrad/kokoro revision dfb907a02bba8152ca444717ca5d78747ccb4bec and validated on 2026-07-26 against the pinned upstream PyTorch implementation.

A self-contained Kokoro inference package converted from the official Kokoro-82M v1.0 checkpoint. It is a StyleTTS 2 decoder with an iSTFTNet generator, runs on CPU and CUDA through the same public library interface, and carries all 54 preset Voices. The model is language-blind: it consumes IPA phoneme token IDs and a style vector, so the Text Frontend is the caller's concern rather than the package's.

Downloads

Profile Download Size Tensor storage SHA-256
F32 kokoro-v1-0-F32.gguf 352.8 MB (352,813,952 bytes) 511 F32 23cde0e3b2a3082fa97a84aed746d8c0cee79eca1ee9599c6e13b7e94ba0328b
F16 kokoro-v1-0-F16.gguf 246.5 MB (246,451,648 bytes) 384 F32 + 127 F16 951e4be979b8af5e72f2353de947c65372c84790b500dd0f53384c958ee2596e
Q8_MIXED kokoro-v1-0-Q8_MIXED.gguf 216.1 MB (216,109,696 bytes) 384 F32 + 14 F16 + 113 Q8_0 0d72f3778125a8f54c23468c6a2534f114d529121f9788e46ec45dc7d21b923c

All profiles use the same Kokoro architecture and public synthesize.cpp API. The profile name describes a versioned storage policy, not the language or Execution Backend.

Validation status

validation_level: port_validated

7 graph stages were replayed for 15 cases on DGX Spark CPU, NVIDIA GB10 CUDA 13.3. Duration structure was exact in every case. CUDA placement contained zero executable CPU fallback nodes.

Profile CPU waveform correlation DGX Spark CUDA waveform correlation
F32 0.987437 0.987973
F16 0.987445 0.988863
Q8_MIXED 0.98479 0.986341

Correlation is the honest measure for this family rather than a sample-wise drift. Kokoro's excitation is a sine whose phase accumulates across the whole utterance, so a difference in F0 of a few parts in ten thousand becomes radians of phase by the last syllable: the waveforms diverge by construction while the speech does not. The figure quoted is the worst case over the suite. What is bit-exact, on every profile and both backends, is the predicted durations.

Quality evaluation has not been run. These results establish that the port, Voice selection, deterministic request path, and CPU/CUDA execution work. They do not claim perceptual equivalence, naturalness, intelligibility, or speaker similarity.

A listening audit found no obvious regression. One maintainer compared a small set against the reference and reported nothing audible. That is release evidence, not a measurement: no rated comparison, no panel, no score, and it does not change the validation level. It says a defect large enough to hear was not found in what was heard.

Voices and input

This package exposes 54 preset speaker IDs, preset-catalog. It produces 24000 Hz mono F32 audio. No default speaker is invented; every request must select a Voice.

This package accepts UTF-8 phoneme strings through the built-in synthesize.symbol_map frontend, and also accepts exact token IDs. The frontend validates UTF-8, maps each Unicode scalar through the symbol table embedded in the GGUF, and applies the model's blank-insertion rule.

The built-in frontend does not perform grapheme-to-phoneme conversion or text normalization. Callers starting from raw text must currently run a compatible G2P frontend externally. The runtime does not silently invoke eSpeak or download a frontend.

Usage

Build synthesize.cpp and synthesize a deterministic request:

git clone https://github.com/handy-computer/synthesize.cpp.git
cd synthesize.cpp
cmake -S . -B build -DSYNTH_BUILD_CLI=ON
cmake --build build -j

hf download jiangzhuo9357/kokoro-v1-0-gguf kokoro-v1-0-F16.gguf \
  --local-dir models/kokoro-v1-0

build/bin/synthesize-cli \
  --model models/kokoro-v1-0/kokoro-v1-0-F16.gguf \
  --output output.wav \
  --phonemes "ðə skˈI əbˈʌv ðə pˈɔɹt wʌz ðə kˈʌləɹ ʌv tˈɛləvˌɪʒən, tˈund tə ɐ dˈɛd ʧˈænᵊl." \
  --language en \
  --voice af_heart \
  --seed 0

The same local GGUF can be loaded through the public C ABI and wrapped by C++, Rust, or Python. Model loading never contacts Hugging Face.

License and checkpoint provenance

Both the pinned source repository and the weights carry an explicit Apache-2.0 grant: the repository ships a LICENSE at the ported revision, and the model card declares license: apache-2.0. The Voice packs are distributed in the same weights repository at the same revision.

The upstream card lists CC BY training sources — Koniwa tnc and SIWIS — which are credited here in accordance with that licence.

Both the source and the weights carry an explicit Apache-2.0 grant, so no redistribution assumption is required for this variant. The upstream model card lists CC BY training sources — Koniwa tnc and SIWIS — whose attribution is carried forward here.

Only English is declared. The checkpoint also ships Voices for British English, Japanese, Mandarin Chinese, Spanish, French, Hindi, Italian and Brazilian Portuguese; those Voices load and run, but no language beyond en has its own validation cases, so none is advertised.


Original upstream project card

Reproduced from the pinned upstream repository card for offline provenance. The upstream repository remains authoritative.

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.

🐈 GitHub: https://github.com/hexgrad/kokoro

🚀 Demo: https://hf.co/spaces/hexgrad/Kokoro-TTS

As of April 2025, the market rate of Kokoro served over API is under $1 per million characters of text input, or under $0.06 per hour of audio output. (On average, 1000 characters of input is about 1 minute of output.) Sources: ArtificialAnalysis/Replicate at 65 cents per M chars and DeepInfra at 80 cents per M chars.

This is an Apache-licensed model, and Kokoro has been deployed in numerous projects and commercial APIs. We welcome the deployment of the model in real use cases.

Fake websites like kokorottsai_com (snapshot: https://archive.ph/nRRnk) and kokorotts_net (snapshot: https://archive.ph/60opa) are likely scams masquerading under the banner of a popular model.

Any website containing "kokoro" in its root domain (e.g. kokorottsai_com, kokorotts_net) is NOT owned by and NOT affiliated with this model page or its author, and attempts to imply otherwise are red flags.

Releases

Model Published Training Data Langs & Voices SHA256
v1.0 2025 Jan 27 Few hundred hrs 8 & 54 496dba11
v0.19 2024 Dec 25 <100 hrs 1 & 10 3b0c392f
Training Costs v0.19 v1.0 Total
in A100 80GB GPU hours 500 500 1000
average hourly rate $0.80/h $1.20/h $1/h
in USD $400 $600 $1000

Usage

You can run this basic cell on Google Colab. Listen to samples. For more languages and details, see Advanced Usage.

!pip install -q kokoro>=0.9.2 soundfile
!apt-get -qq -y install espeak-ng > /dev/null 2>&1
from kokoro import KPipeline
from IPython.display import display, Audio
import soundfile as sf
import torch
pipeline = KPipeline(lang_code='a')
text = '''
[Kokoro](/kˈOkəɹO/) is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, [Kokoro](/kˈOkəɹO/) can be deployed anywhere from production environments to personal projects.
'''
generator = pipeline(text, voice='af_heart')
for i, (gs, ps, audio) in enumerate(generator):
    print(i, gs, ps)
    display(Audio(data=audio, rate=24000, autoplay=i==0))
    sf.write(f'{i}.wav', audio, 24000)

Under the hood, kokoro uses misaki, a G2P library at https://github.com/hexgrad/misaki

Model Facts

Architecture:

Architected by: Li et al @ https://github.com/yl4579/StyleTTS2

Trained by: @rzvzn on Discord

Languages: Multiple

Model SHA256 Hash: 496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4

Training Details

Data: Kokoro was trained exclusively on permissive/non-copyrighted audio data and IPA phoneme labels. Examples of permissive/non-copyrighted audio include:

  • Public domain audio
  • Audio licensed under Apache, MIT, etc
  • Synthetic audio[1] generated by closed[2] TTS models from large providers
    [1] https://copyright.gov/ai/ai_policy_guidance.pdf
    [2] No synthetic audio from open TTS models or "custom voice clones"

Total Dataset Size: A few hundred hours of audio

Total Training Cost: About $1000 for 1000 hours of A100 80GB vRAM

Creative Commons Attribution

The following CC BY audio was part of the dataset used to train Kokoro v1.0.

Audio Data Duration Used License Added to Training Set After
Koniwa tnc <1h CC BY 3.0 v0.19 / 22 Nov 2024
SIWIS <11h CC BY 4.0 v0.19 / 22 Nov 2024

Acknowledgements

  • 🛠️ @yl4579 for architecting StyleTTS 2.
  • 🏆 @Pendrokar for adding Kokoro as a contender in the TTS Spaces Arena.
  • 📊 Thank you to everyone who contributed synthetic training data.
  • ❤️ Special thanks to all compute sponsors.
  • 👾 Discord server: https://discord.gg/QuGxSWBfQy
  • 🪽 Kokoro is a Japanese word that translates to "heart" or "spirit". It is also the name of an AI in the Terminator franchise.
kokoro
Downloads last month
372
GGUF
Model size
88.2M params
Architecture
kokoro
Hardware compatibility
Log In to add your hardware

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for jiangzhuo9357/kokoro-v1-0-gguf