wav2vec2-xlsr-300m-hre-v1

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the hre-audio-dataset8 dataset. It achieves the following results on the evaluation set:

  • Loss: 0.7875
  • Cer Ortho: 51.8126
  • Cer: 51.3707

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0003
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 100
  • num_epochs: 20
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Cer Ortho Cer
3.4833 0.4608 100 3.4517 99.0488 99.0248
3.1952 0.9217 200 3.0203 99.0488 99.0248
2.9279 1.3825 300 2.8936 99.0488 99.0248
2.8948 1.8433 400 2.7970 91.8521 91.9595
2.7727 2.3041 500 2.6553 95.4774 95.4186
2.4391 2.7650 600 2.4019 87.9397 87.6541
2.0669 3.2258 700 1.9134 83.4709 78.6017
1.7675 3.6866 800 1.7641 86.1630 78.7305
1.4705 4.1475 900 1.5949 84.8708 75.3266
1.3456 4.6083 1000 1.3056 78.8047 71.2052
1.3088 5.0691 1100 1.2253 78.9663 69.8068
1.1681 5.5300 1200 1.1573 77.6920 67.7093
0.9910 5.9908 1300 1.1332 74.0129 66.0718
0.9176 6.4516 1400 1.0091 64.1242 61.5823
1.0415 6.9124 1500 0.9127 61.3245 60.7176
0.9014 7.3733 1600 0.9495 60.3553 59.7424
0.8234 7.8341 1700 0.9731 61.4860 60.9384
0.7313 8.2949 1800 0.9365 60.0503 59.2456
0.7660 8.7558 1900 0.9118 59.8349 59.6136
0.6833 9.2166 2000 0.8462 56.5506 56.5593
0.6984 9.6774 2100 0.8768 58.0761 57.3689
0.6272 10.1382 2200 0.9320 58.5068 57.9025
0.6431 10.5991 2300 0.8695 56.4070 56.4489
0.5980 11.0599 2400 0.8595 56.4070 56.5225
0.5701 11.5207 2500 0.9245 57.3044 56.6697
0.6168 11.9816 2600 0.7961 54.0380 53.6707
0.5916 12.4424 2700 0.8187 55.2225 55.0138
0.6809 12.9032 2800 0.8054 54.8636 54.6826
0.5780 13.3641 2900 0.8034 53.9304 53.7626
0.5056 13.8249 3000 0.8062 53.9483 53.8730
0.5218 14.2857 3100 0.8272 54.3611 54.2226
0.5384 14.7465 3200 0.8102 53.5535 52.9899
0.5307 15.2074 3300 0.7969 52.8715 52.2907
0.4648 15.6682 3400 0.7977 53.2843 52.6403
0.4624 16.1290 3500 0.8211 53.5714 52.8795
0.4073 16.5899 3600 0.7983 52.4587 51.7755
0.4627 17.0507 3700 0.7947 52.9612 52.2355
0.4158 17.5115 3800 0.8145 52.1177 51.6651
0.3863 17.9724 3900 0.7885 52.1177 51.7571
0.4024 18.4332 4000 0.8038 51.9921 51.4259
0.3865 18.8940 4100 0.8207 52.0101 51.3523
0.4091 19.3548 4200 0.7871 51.7408 51.1684
0.3589 19.8157 4300 0.7872 51.8665 51.3891
0.3606 20.0 4340 0.7875 51.8126 51.3707

Framework versions

  • Transformers 5.13.1
  • Pytorch 2.11.0+cu128
  • Datasets 2.18.0
  • Tokenizers 0.22.2
Downloads last month
17
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ntviet/wav2vec2-xlsr-300m-hre-v1

Finetuned
(888)
this model