Quds-v4-sherpa-onnx
๐จ Non Commercial Usage Only
๐ก Fine-tuned specifically for the domain of Islamic lectures and specialized Howzah courses, such as Tafsir, Fiqh, Usul, and Rijal.
๐ FOR NeMo version (includes both rnnt and ctc decoders and word timestamps) send a request in Community
Model Details
- Model Name: Quds-v4-sherpa-onnx
- Model Type: Automatic Speech Recognition (ASR) / Speech-to-Text
- Architecture: NeMo FastConformer Large Hybrid (RNN-T variant)
- Format: ONNX (Open Neural Network Exchange)
- Language: Persian (Farsi)
Model Description
Quds-v4-sherpa-onnx is a highly accurate Persian Automatic Speech Recognition model. It is based on the robust NeMo FastConformer Large Hybrid architecture and is specifically exported in the ONNX format to enable fast, cross-platform inference. This release features an RNN-T (Recurrent Neural Network Transducer) decoder.
Performance & Accuracy
- The model achieves very high accuracy on standard, formal Persian speech.
Training Details
- Training Dataset: The model was trained on a high-quality, non-public (private) dataset.
- Dataset Size: 600 hours of audio data.
Usage
This is sherpa version of model: Usage Docs.
For using without sherpa, check this model.
Limitations
- Dialect Sensitivity: Due to training on standard Persian audio, the model struggles with heavy regional accents, colloquialisms, and non-standard dialects of Persian.
- Audio Quality Sensitivity: As with most ASR models, performance degrades in highly noisy environments.
- Domain Specificity: The model may not be perfectly suited for highly specialized technical or medical domains.
Model tree for hojreh/Quds-v4-sherpa-onnx
Base model
nvidia/stt_fa_fastconformer_hybrid_large