Fine-tuned ECAPA-TDNN for Vietnamese Speaker Verification

This model is a fine-tuned ECAPA-TDNN speaker embedding model based on speechbrain/spkrec-ecapa-voxceleb.

Training

  • Fine-tuning dataset: Vietnam-Celeb
  • Selected checkpoint: Epoch 10
  • Embedding dimension: 192
  • Sample rate: 16 kHz

Verification Pipeline

Audio → CRDNN VAD → ECAPA-TDNN embedding → L2 normalization → 5-recording enrollment → mean speaker centroid → cosine similarity → threshold decision

Final DEV-calibrated threshold:

0.1566

Speaker-disjoint Evaluation

All-impostor evaluation protocol:

Model DEV EER TEST EER TEST FAR TEST FRR
Pretrained ECAPA 13.30% 11.91% 9.46% 13.47%
Fine-tuned Epoch 10 9.98% 8.42% 9.00% 7.88%

TEST evaluation:

  • 50 unseen speakers
  • 5 enrollment recordings per speaker
  • 698 genuine trials
  • 34,202 impostor trials
  • No speaker overlap between train/DEV/TEST

The TEST threshold was not tuned on TEST. The final threshold was calibrated only on DEV.

Base Model

SpeechBrain: speechbrain/spkrec-ecapa-voxceleb

Limitations

This model is intended for an academic prototype of a speaker-verification system. The reported FAR is not low enough to claim production-grade security.

Performance may vary with microphone quality, background noise, language, recording duration, and speaker characteristics.

Files

  • ecapa_vietnamceleb_epoch10.pt: fine-tuned ECAPA embedding checkpoint
  • config.json: deployment configuration and verification threshold
  • all_impostor_metrics.json: detailed evaluation metrics
  • pretrained_vs_finetuned.csv: baseline comparison
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support