Robust Speaker Cloning artifacts

Large artifacts for the companion GitHub repository. V2 trains a 6-layer robust speaker Transformer and telephone BWE front end while keeping CosyVoice2-0.5B frozen. On 21 paired real CosyVoice generations, ECAPA x-vector similarity was 0.6187 for V2 versus 0.5998 baseline (+0.0189 average). This small evaluation also contains negative HVAC and reverb deltas, so it is not a universal-improvement claim.

CosyVoice, CampPlus, ECAPA and AISHELL retain their upstream licenses.

Downloads last month
104
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support