Whisper Large V3 โ€” Luganda NSS v22 (Full Model + SpecAugment)

Stage 2 fine-tune on the CDLI Ugandan Luganda Nonstandard Speech dataset. Built on top of lg_waxal_tune_whisper_large_7 (Stage 1). This is Variant 1 of the Stage 2 experiments โ€” full model with SpecAugment on.

Training Details

  • Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
  • Stage 1 checkpoint: lg_waxal_tune_whisper_large_7/best_model
  • Encoder frozen: No
  • Decoder frozen: No
  • Projection frozen: No
  • SpecAugment: On
  • Learning rate: 5e-5 (polynomial decay)
  • Max steps: 1000
  • Best checkpoint: Step 600
  • Output dir: lg_nss_tune_whisper_large_v22

Results (checkpoint-600)

Severity Baseline WER Fine-tuned WER
Mild 0.89 0.49
Moderate 0.96 0.54
Severe 0.97 0.60
  • Avg WER (normalized): 0.534
  • Meaning preservation: 84.7%
Downloads last month
4
Safetensors
Model size
2B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ElizabethMwangi/whisper-large-v3-luganda-nss-v22

Dataset used to train ElizabethMwangi/whisper-large-v3-luganda-nss-v22