Whisper Large V3 โ Luganda NSS v22 (Full Model + SpecAugment)
Stage 2 fine-tune on the CDLI Ugandan Luganda Nonstandard Speech dataset.
Built on top of lg_waxal_tune_whisper_large_7 (Stage 1).
This is Variant 1 of the Stage 2 experiments โ full model with SpecAugment on.
Training Details
- Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
- Stage 1 checkpoint: lg_waxal_tune_whisper_large_7/best_model
- Encoder frozen: No
- Decoder frozen: No
- Projection frozen: No
- SpecAugment: On
- Learning rate: 5e-5 (polynomial decay)
- Max steps: 1000
- Best checkpoint: Step 600
- Output dir: lg_nss_tune_whisper_large_v22
Results (checkpoint-600)
| Severity |
Baseline WER |
Fine-tuned WER |
| Mild |
0.89 |
0.49 |
| Moderate |
0.96 |
0.54 |
| Severe |
0.97 |
0.60 |
- Avg WER (normalized): 0.534
- Meaning preservation: 84.7%