Configuration Parsing Warning:Invalid JSON for config file config.json
Tennda-Waves
XTTS-v2 voice cloning model, fine-tuned via LoRA on coqui/XTTS-v2.
Features
- Bilingual: Supports Chinese and English speech synthesis
- Voice Cloning: Input a reference audio to clone any voice
- Lightweight Fine-tuning: Only 0.88% parameters trained (3.93M / 444.95M)
- Training Data: AISHELL-3 Chinese male speaker (SSB0273)
Quick Start
from TTS.tts.models.xtts import Xtts
from TTS.tts.configs.xtts_config import XttsConfig
import torchaudio
# Load model
config = XttsConfig()
config.load_json("config.json")
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_path="model.pth",
vocab_path="vocab.json",
speaker_file_path=" ",
eval=True,
)
# Get speaker embeddings
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
audio_path=["reference.wav"], # Reference audio
gpt_cond_len=30,
gpt_cond_chunk_len=4,
max_ref_length=60,
)
# Synthesize speech
out = model.inference(
text="Hello, this is Tennda speaking.",
language="en",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
temperature=0.75,
)
# Save audio
torchaudio.save("output.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
Training Details
- Base Model: coqui/XTTS-v2 (CPML license)
- Fine-tuning Method: LoRA (r=8, alpha=32)
- Target Modules: c_attn, c_proj, c_fc
- Training Data: AISHELL-3 SSB0273 (100 samples, 7.4 minutes)
- Training Epochs: 10
- Optimizer: AdamW (lr=1e-5)
License
This model is fine-tuned from coqui/XTTS-v2, which is licensed under the CPML license (non-commercial research use only).
⚠️ Warning: This model is NOT licensed for commercial use.
Links
- Base Model: coqui/XTTS-v2
- Training Data: AISHELL-3
- Project Repository: tennda-nano-training
- Downloads last month
- 15