---
library_name: nemo
license: cc-by-4.0
language:
- ar
pipeline_tag: automatic-speech-recognition
tags:
- automatic-speech-recognition
- asr
- arabic
- arabic-asr
- dialectal-arabic
- gulf-arabic
- fastconformer
- ctc
- efficient
---

# lemura-arabic-asr
**A compact, dialect-aware Arabic speech-recognition model โ leaderboard-tier accuracy at ~18ร fewer parameters.**





[๐งญ Overview](#-overview) ยท [๐ Model Details](#-model-details) ยท [๐ Benchmarks](#-benchmarks) ยท [โก Efficiency](#-efficiency) ยท [๐ป Usage](#-usage) ยท [๐๏ธ Live Demo](https://huggingface.co/spaces/lemuralabs/lemura-arabic-asr-demo)
---
## ๐งญ Overview
**lemura-arabic-asr** is a compact multi-dialect Arabic ASR model built for accuracy *and* efficiency. It is a **FastConformer-CTC** acoustic model (\~115M parameters), adapted in-house from the NVIDIA FastConformer foundation and fine-tuned on **\~2,900 hours** of Arabic spanning **MSA** and the **Gulf, Egyptian, Levantine, and Maghrebi** dialect groups.
**Highlights**
- ๐ชถ **Small & fast** โ ~115M parameters; runs comfortably on **CPU** and in **real time**, no GPU required.
- ๐ฃ๏ธ **Dialect-aware** โ trained across five Arabic dialect groups, not MSA-only.
- ๐ง **Robust on real audio** โ strongest on broadcast, conversational, and Gulf/MSA speech.
- ๐ **Open & simple** โ a single `.nemo` file, loadable in a few lines with NVIDIA NeMo.
## ๐ Model Details
| | |
|---|---|
| **Model** | lemura-arabic-asr โ compact multi-dialect Arabic ASR |
| **Task** | Automatic speech recognition (audio โ text) |
| **Architecture** | FastConformer encoder + CTC decoder |
| **Parameters** | ~115M |
| **Foundation** | Adapted from NVIDIA FastConformer; ~2,900 h Arabic fine-tuning |
| **Audio input** | 16 kHz mono (auto-resampled) |
| **Languages** | Arabic โ MSA + Gulf / Egyptian / Levantine / Maghrebi |
| **Runtime** | NVIDIA NeMo โ CPU ยท GPU ยท real-time |
| **License** | CC-BY-4.0 |
## ๐ Benchmarks
Evaluated on all six **Open Universal Arabic ASR Leaderboard** test sets using the [official leaderboard code](https://github.com/Natural-Language-Processing-Elm/open_universal_arabic_asr_leaderboard) (same normalizer, same WER metric).
### Per-set WER (%)
| SADA | Common Voice 18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca | **Average** |
|---:|---:|---:|---:|---:|---:|---:|
| 37.28 | 9.74 | 7.27 | 23.65 | 14.33 | 58.24 | **25.08** |
### In context (Average WER %, lower is better)
| Model | Params | Avg WER |
|---|---:|---:|
| cohere-transcribe-arabic-07-2026 | ~2.0B | 25.87 |
| **lemura-arabic-asr** | **~0.12B** | **25.08\*** |
| omniASR_LLM_7B | 7B | 28.32 |
| Qwen3-Omni-30B-A3B | 30B | 30.71 |
| nvidia-conformer-ctc-large-arabic (lm) | 0.6B | 32.91 |
| Qwen3-ASR-1.7B | 1.7B | 33.36 |
## โก Efficiency
The top leaderboard systems are large generative audio-LLMs (2โ30B parameters) that need GPUs. lemura-arabic-asr reaches a comparable accuracy tier with a **~115M-parameter** CTC model:
| | lemura-arabic-asr | Typical top systems |
|---|---|---|
| Parameters | **~115M** | 2B โ 30B |
| Hardware | **CPU or GPU** | GPU |
| Latency | **Real-time** | Seconds / clip |
| Footprint | **~0.4 GB** | 4 โ 60 GB |
That makes it practical for **on-device, low-cost, and high-throughput** Arabic transcription where the big models are impractical.
## ๐ป Usage
```python
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"])) # 16 kHz mono
```
Try it live โ no install: **[๐๏ธ lemura-arabic-asr demo](https://huggingface.co/spaces/lemuralabs/lemura-arabic-asr-demo)**
## ๐ Credits
- Acoustic foundation: **NVIDIA FastConformer**
- Evaluation: the **[Open Universal Arabic ASR Leaderboard](https://github.com/Natural-Language-Processing-Elm/open_universal_arabic_asr_leaderboard)** official code