How to use from the
Use from the
VoxCPM library
import soundfile as sf
from voxcpm import VoxCPM

model = VoxCPM.from_pretrained("walusungungulube/ghana-tts-36k-gguf")

wav = model.generate(
    text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.",
    prompt_wav_path=None,      # optional: path to a prompt speech for voice cloning
    prompt_text=None,          # optional: reference text
    cfg_value=2.0,             # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse
    inference_timesteps=10,   # LocDiT inference timesteps, higher for better result, lower for fast speed
    normalize=True,           # enable external TN tool
    denoise=True,             # enable external Denoise tool
    retry_badcase=True,        # enable retrying mode for some bad cases (unstoppable)
    retry_badcase_max_times=3,  # maximum retrying times
    retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech
)

sf.write("output.wav", wav, 16000)
print("saved: output.wav")

ghana-tts-36k β€” GGUF quantized variants

GGUF conversions of ghananlpcommunity/ghana-tts-36k for use with bluryar/VoxCPM.cpp.

The .gguf files are backend-agnostic β€” the same file runs on both a CPU-only and a CUDA build of the inference engine. Only the binary you compile differs between CPU and GPU deployment, not the weight file.

Files

File Quant AudioVAE Size Notes
ghana-tts-36k-f32.gguf F32 original largest Reference/max-accuracy baseline
ghana-tts-36k-f16.gguf F16 mixed
ghana-tts-36k-f16-audiovae-f16.gguf F16 f16
ghana-tts-36k-q8_0.gguf Q8_0 mixed Recommended for CPU
ghana-tts-36k-q8_0-audiovae-f16.gguf Q8_0 f16 Recommended for GPU (CUDA)
ghana-tts-36k-q4_k.gguf Q4_K mixed Smallest, more accuracy loss
ghana-tts-36k-q4_k-audiovae-f16.gguf Q4_K f16 smallest of the practical options Fastest CPU model-only RTF; good compact CUDA option too

Which one to use

Based on bluryar/VoxCPM.cpp's own published benchmarks on the same "voxcpm" architecture family (not run on this exact checkpoint β€” validate on your own audio before committing to one in production):

CPU inference:

  • Default pick: ghana-tts-36k-q8_0.gguf β€” best full-pipeline RTF at this model scale on CPU.
  • Smaller/faster, slight accuracy trade-off: ghana-tts-36k-q4_k-audiovae-f16.gguf.

GPU (CUDA) inference:

  • Default pick: ghana-tts-36k-q8_0-audiovae-f16.gguf β€” best full-pipeline RTF at this model scale on CUDA.
  • Smallest CUDA-friendly option: ghana-tts-36k-q4_k-audiovae-f16.gguf.

Inference

This repo includes two scripts that build and run VoxCPM.cpp against these weights:

  • infer_cpu.sh β€” builds a CPU-only voxcpm_tts and runs it against ghana-tts-36k-q8_0.gguf.
  • infer_gpu.sh β€” builds a CUDA-enabled voxcpm_tts (requires an NVIDIA GPU + CUDA toolkit) and runs it against ghana-tts-36k-q8_0-audiovae-f16.gguf.

Usage (either script):

bash infer_cpu.sh "Text to synthesize" prompt.wav "Exact transcript of prompt.wav" out.wav
bash infer_gpu.sh "Text to synthesize" prompt.wav "Exact transcript of prompt.wav" out.wav

Swap in a different .gguf from this repo by editing the MODEL_PATH variable at the top of either script.

Downloads last month
91
GGUF
Model size
0.7B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for walusungungulube/ghana-tts-36k-gguf

Quantized
(2)
this model