Whisper Large V3 โ Luganda NSS v23 (Encoder only + SpecAugment)
Stage 2 fine-tune on the CDLI Ugandan Luganda Nonstandard Speech dataset.
Built on top of lg_waxal_tune_whisper_large_7 (Stage 1).
This is Variant 2 of the Stage 2 experiments โ encoder only with SpecAugment on.
Training Details
- Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
- Stage 1 checkpoint: lg_waxal_tune_whisper_large_7/best_model
- Encoder frozen: No (encoder trained only)
- Decoder frozen: Yes
- Projection frozen: Yes
- SpecAugment: On
- Learning rate: 5e-5 (polynomial decay)
- Max steps: 1000
- Best checkpoint: Step 300
- Output dir: lg_nss_tune_whisper_large_v23
Results (checkpoint-300)
| Severity |
Baseline WER |
Fine-tuned WER |
| Mild |
0.89 |
0.49 |
| Moderate |
0.96 |
0.58 |
| Severe |
0.97 |
0.65 |
- Avg WER (normalized): 0.555
- Meaning preservation: 80.4%