Persian speech recognition leaderboard

Which Persian ASR models actually hear better?

This leaderboard compares Persian automatic speech recognition models on two complementary tests. VisualEars6669 is the 6,669-row noisy real-world Persian set. FLEURS-fa Full is the full 4,341-row Persian FLEURS benchmark scored against raw_transcription. Lower WER/CER means fewer transcription mistakes.

Click a column to sort. Triple Threat = 60% Sยณ + 20% WER + 20% CER, with both datasets weighted 50/50.