--- license: mit tags: - text-to-speech - voxcpm - gguf - ghana-tts base_model: ghananlpcommunity/ghana-tts-36k --- # ghana-tts-36k — GGUF quantized variants GGUF conversions of [`ghananlpcommunity/ghana-tts-36k`](https://huggingface.co/ghananlpcommunity/ghana-tts-36k) for use with [bluryar/VoxCPM.cpp](https://github.com/bluryar/VoxCPM.cpp). **The `.gguf` files are backend-agnostic** — the same file runs on both a CPU-only and a CUDA build of the inference engine. Only the *binary* you compile differs between CPU and GPU deployment, not the weight file. ## Files | File | Quant | AudioVAE | Size | Notes | |---|---|---|---|---| | `ghana-tts-36k-f32.gguf` | F32 | original | largest | Reference/max-accuracy baseline | | `ghana-tts-36k-f16.gguf` | F16 | mixed | | | | `ghana-tts-36k-f16-audiovae-f16.gguf` | F16 | f16 | | | | `ghana-tts-36k-q8_0.gguf` | Q8_0 | mixed | | **Recommended for CPU** | | `ghana-tts-36k-q8_0-audiovae-f16.gguf` | Q8_0 | f16 | | **Recommended for GPU (CUDA)** | | `ghana-tts-36k-q4_k.gguf` | Q4_K | mixed | | Smallest, more accuracy loss | | `ghana-tts-36k-q4_k-audiovae-f16.gguf` | Q4_K | f16 | smallest of the practical options | Fastest CPU model-only RTF; good compact CUDA option too | ## Which one to use Based on [bluryar/VoxCPM.cpp](https://github.com/bluryar/VoxCPM.cpp)'s own published benchmarks on the same `"voxcpm"` architecture family (not run on this exact checkpoint — validate on your own audio before committing to one in production): **CPU inference:** - Default pick: **`ghana-tts-36k-q8_0.gguf`** — best full-pipeline RTF at this model scale on CPU. - Smaller/faster, slight accuracy trade-off: `ghana-tts-36k-q4_k-audiovae-f16.gguf`. **GPU (CUDA) inference:** - Default pick: **`ghana-tts-36k-q8_0-audiovae-f16.gguf`** — best full-pipeline RTF at this model scale on CUDA. - Smallest CUDA-friendly option: `ghana-tts-36k-q4_k-audiovae-f16.gguf`. ## Inference This repo includes two scripts that build and run [VoxCPM.cpp](https://github.com/bluryar/VoxCPM.cpp) against these weights: - `infer_cpu.sh` — builds a CPU-only `voxcpm_tts` and runs it against `ghana-tts-36k-q8_0.gguf`. - `infer_gpu.sh` — builds a CUDA-enabled `voxcpm_tts` (requires an NVIDIA GPU + CUDA toolkit) and runs it against `ghana-tts-36k-q8_0-audiovae-f16.gguf`. Usage (either script): ```bash bash infer_cpu.sh "Text to synthesize" prompt.wav "Exact transcript of prompt.wav" out.wav bash infer_gpu.sh "Text to synthesize" prompt.wav "Exact transcript of prompt.wav" out.wav ``` Swap in a different `.gguf` from this repo by editing the `MODEL_PATH` variable at the top of either script.