wav2vec2-lv-60-espeak-cv-ft-WCTC-phocab-ds-f8

This model is a fine-tuned version of facebook/wav2vec2-lv-60-espeak-cv-ft on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 11.8064
  • Per: 0.0451

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 3e-05
  • train_batch_size: 16
  • eval_batch_size: 8
  • seed: 42
  • gradient_accumulation_steps: 2
  • total_train_batch_size: 32
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 500
  • num_epochs: 30
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Per
1388.9211 0.7194 400 258.8207 1.0
496.2495 1.4388 800 241.3190 1.0
467.8013 2.1583 1200 209.6756 0.9895
319.2299 2.8777 1600 64.0295 0.2880
171.572 3.5971 2000 38.2525 0.1529
133.5128 4.3165 2400 29.6704 0.1030
114.4326 5.0360 2800 24.8591 0.0853
101.9263 5.7554 3200 21.7826 0.0764
94.2615 6.4748 3600 19.1824 0.0692
90.1317 7.1942 4000 18.4454 0.0708
84.2694 7.9137 4400 17.8917 0.0668
82.0402 8.6331 4800 16.9460 0.0652
79.0095 9.3525 5200 16.8115 0.0652
77.0721 10.0719 5600 16.3140 0.0619
73.6757 10.7914 6000 15.6164 0.0636
70.9067 11.5108 6400 15.7822 0.0611
72.3999 12.2302 6800 14.8963 0.0603
70.6939 12.9496 7200 14.0957 0.0571
69.9356 13.6691 7600 13.9708 0.0523
67.5808 14.3885 8000 13.7943 0.0523
68.9874 15.1079 8400 13.7925 0.0523
63.5566 15.8273 8800 13.4886 0.0531
64.1214 16.5468 9200 13.1847 0.0523
64.8615 17.2662 9600 13.5771 0.0547
62.7331 17.9856 10000 13.7224 0.0531
61.7567 18.7050 10400 13.6806 0.0531
62.0574 19.4245 10800 12.6783 0.0515
59.3861 20.1439 11200 13.1509 0.0507
61.258 20.8633 11600 12.6506 0.0483
58.7836 21.5827 12000 12.7024 0.0491
59.459 22.3022 12400 12.1852 0.0491
58.4748 23.0216 12800 12.5531 0.0475
59.3281 23.7410 13200 12.1136 0.0442
59.496 24.4604 13600 12.0433 0.0467
57.9776 25.1799 14000 12.1108 0.0483
56.2327 25.8993 14400 11.9060 0.0467
58.9476 26.6187 14800 11.8183 0.0467
57.7004 27.3381 15200 11.9450 0.0451
57.3872 28.0576 15600 11.8708 0.0451
54.7268 28.7770 16000 11.9073 0.0467
57.4947 29.4964 16400 11.8064 0.0451

Framework versions

  • Transformers 4.57.6
  • Pytorch 2.9.1+cu128
  • Datasets 4.5.0
  • Tokenizers 0.22.2
Downloads last month
3
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zacdan4801/wav2vec2-lv-60-espeak-cv-ft-WCTC-phocab-ds-f8

Finetuned
(84)
this model