--- library_name: nemo license: cc-by-4.0 language: - ar pipeline_tag: automatic-speech-recognition tags: - automatic-speech-recognition - asr - arabic - arabic-asr - dialectal-arabic - gulf-arabic - fastconformer - ctc - efficient ---
# lemura-arabic-asr **A compact, dialect-aware Arabic speech-recognition model โ€” leaderboard-tier accuracy at ~18ร— fewer parameters.** ![Task](https://img.shields.io/badge/task-ASR-blue) ![Params](https://img.shields.io/badge/params-~115M-f59e0b) ![Architecture](https://img.shields.io/badge/arch-FastConformer--CTC-6f42c1) ![Dialects](https://img.shields.io/badge/dialects-MSA%20%2B%204%20groups-brightgreen) ![Runtime](https://img.shields.io/badge/runtime-CPU%20%C2%B7%20GPU%20%C2%B7%20real--time-blue) [๐Ÿงญ Overview](#-overview) ยท [๐Ÿ“‹ Model Details](#-model-details) ยท [๐Ÿ“Š Benchmarks](#-benchmarks) ยท [โšก Efficiency](#-efficiency) ยท [๐Ÿ’ป Usage](#-usage) ยท [๐ŸŽ™๏ธ Live Demo](https://huggingface.co/spaces/lemuralabs/lemura-arabic-asr-demo)
--- ## ๐Ÿงญ Overview **lemura-arabic-asr** is a compact multi-dialect Arabic ASR model built for accuracy *and* efficiency. It is a **FastConformer-CTC** acoustic model (\~115M parameters), adapted in-house from the NVIDIA FastConformer foundation and fine-tuned on **\~2,900 hours** of Arabic spanning **MSA** and the **Gulf, Egyptian, Levantine, and Maghrebi** dialect groups. **Highlights** - ๐Ÿชถ **Small & fast** โ€” ~115M parameters; runs comfortably on **CPU** and in **real time**, no GPU required. - ๐Ÿ—ฃ๏ธ **Dialect-aware** โ€” trained across five Arabic dialect groups, not MSA-only. - ๐ŸŽง **Robust on real audio** โ€” strongest on broadcast, conversational, and Gulf/MSA speech. - ๐Ÿ”“ **Open & simple** โ€” a single `.nemo` file, loadable in a few lines with NVIDIA NeMo. ## ๐Ÿ“‹ Model Details | | | |---|---| | **Model** | lemura-arabic-asr โ€” compact multi-dialect Arabic ASR | | **Task** | Automatic speech recognition (audio โ†’ text) | | **Architecture** | FastConformer encoder + CTC decoder | | **Parameters** | ~115M | | **Foundation** | Adapted from NVIDIA FastConformer; ~2,900 h Arabic fine-tuning | | **Audio input** | 16 kHz mono (auto-resampled) | | **Languages** | Arabic โ€” MSA + Gulf / Egyptian / Levantine / Maghrebi | | **Runtime** | NVIDIA NeMo โ€” CPU ยท GPU ยท real-time | | **License** | CC-BY-4.0 | ## ๐Ÿ“Š Benchmarks Evaluated on all six **Open Universal Arabic ASR Leaderboard** test sets using the [official leaderboard code](https://github.com/Natural-Language-Processing-Elm/open_universal_arabic_asr_leaderboard) (same normalizer, same WER metric). ### Per-set WER (%) | SADA | Common Voice 18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca | **Average** | |---:|---:|---:|---:|---:|---:|---:| | 37.28 | 9.74 | 7.27 | 23.65 | 14.33 | 58.24 | **25.08** | ### In context (Average WER %, lower is better) | Model | Params | Avg WER | |---|---:|---:| | cohere-transcribe-arabic-07-2026 | ~2.0B | 25.87 | | **lemura-arabic-asr** | **~0.12B** | **25.08\*** | | omniASR_LLM_7B | 7B | 28.32 | | Qwen3-Omni-30B-A3B | 30B | 30.71 | | nvidia-conformer-ctc-large-arabic (lm) | 0.6B | 32.91 | | Qwen3-ASR-1.7B | 1.7B | 33.36 | ## โšก Efficiency The top leaderboard systems are large generative audio-LLMs (2โ€“30B parameters) that need GPUs. lemura-arabic-asr reaches a comparable accuracy tier with a **~115M-parameter** CTC model: | | lemura-arabic-asr | Typical top systems | |---|---|---| | Parameters | **~115M** | 2B โ€“ 30B | | Hardware | **CPU or GPU** | GPU | | Latency | **Real-time** | Seconds / clip | | Footprint | **~0.4 GB** | 4 โ€“ 60 GB | That makes it practical for **on-device, low-cost, and high-throughput** Arabic transcription where the big models are impractical. ## ๐Ÿ’ป Usage ```python import nemo.collections.asr as nemo_asr model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo") print(model.transcribe(["audio.wav"])) # 16 kHz mono ``` Try it live โ€” no install: **[๐ŸŽ™๏ธ lemura-arabic-asr demo](https://huggingface.co/spaces/lemuralabs/lemura-arabic-asr-demo)** ## ๐Ÿ™ Credits - Acoustic foundation: **NVIDIA FastConformer** - Evaluation: the **[Open Universal Arabic ASR Leaderboard](https://github.com/Natural-Language-Processing-Elm/open_universal_arabic_asr_leaderboard)** official code