Configuration Parsing Warning:Invalid JSON for config file config.json

Tennda-Waves

XTTS-v2 voice cloning model, fine-tuned via LoRA on coqui/XTTS-v2.

Features

  • Bilingual: Supports Chinese and English speech synthesis
  • Voice Cloning: Input a reference audio to clone any voice
  • Lightweight Fine-tuning: Only 0.88% parameters trained (3.93M / 444.95M)
  • Training Data: AISHELL-3 Chinese male speaker (SSB0273)

Quick Start

from TTS.tts.models.xtts import Xtts
from TTS.tts.configs.xtts_config import XttsConfig
import torchaudio

# Load model
config = XttsConfig()
config.load_json("config.json")
model = Xtts.init_from_config(config)
model.load_checkpoint(
    config,
    checkpoint_path="model.pth",
    vocab_path="vocab.json",
    speaker_file_path=" ",
    eval=True,
)

# Get speaker embeddings
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
    audio_path=["reference.wav"],  # Reference audio
    gpt_cond_len=30,
    gpt_cond_chunk_len=4,
    max_ref_length=60,
)

# Synthesize speech
out = model.inference(
    text="Hello, this is Tennda speaking.",
    language="en",
    gpt_cond_latent=gpt_cond_latent,
    speaker_embedding=speaker_embedding,
    temperature=0.75,
)

# Save audio
torchaudio.save("output.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)

Training Details

  • Base Model: coqui/XTTS-v2 (CPML license)
  • Fine-tuning Method: LoRA (r=8, alpha=32)
  • Target Modules: c_attn, c_proj, c_fc
  • Training Data: AISHELL-3 SSB0273 (100 samples, 7.4 minutes)
  • Training Epochs: 10
  • Optimizer: AdamW (lr=1e-5)

License

This model is fine-tuned from coqui/XTTS-v2, which is licensed under the CPML license (non-commercial research use only).

⚠️ Warning: This model is NOT licensed for commercial use.

Links

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support