Instructions to use lemuralabs/lemura-arabic-asr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use lemuralabs/lemura-arabic-asr with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("lemuralabs/lemura-arabic-asr") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
lemura-arabic-asr
A compact, dialect-aware Arabic speech-recognition model β leaderboard-tier accuracy at ~18Γ fewer parameters.
π§ Overview Β· π Benchmarks Β· β‘ Efficiency Β· π» Usage Β· ποΈ Live Demo
π§ What it is
lemura-arabic-asr is a compact multi-dialect Arabic ASR model built for accuracy and efficiency.
It is a FastConformer-CTC acoustic model (115M parameters) adapted in-house from the NVIDIA
FastConformer foundation, fine-tuned on **2,900 hours** of Arabic spanning MSA and the Gulf,
Egyptian, Levantine, and Maghrebi dialect groups.
- πͺΆ Small & fast β ~115M parameters. Runs comfortably on CPU and in real time, no GPU required.
- π£οΈ Dialect-aware β trained across five Arabic dialect groups, not MSA-only.
- π§ Robust on real audio β strongest on broadcast, conversational, and Gulf/MSA speech.
- π Open & simple β a single
.nemofile, loadable in a few lines with NVIDIA NeMo.
Model summary
| Model | lemura-arabic-asr β compact multi-dialect Arabic ASR |
| Task | Automatic speech recognition (audio β text) |
| Architecture | FastConformer encoder + CTC decoder |
| Parameters | ~115M |
| Foundation | adapted from NVIDIA FastConformer; ~2,900 h Arabic fine-tuning |
| Audio input | 16 kHz mono (auto-resampled) |
| Languages | Arabic β MSA + Gulf / Egyptian / Levantine / Maghrebi |
| Runtime | NVIDIA NeMo β CPU Β· GPU Β· real-time |
| License | CC-BY-4.0 |
π Benchmarks
Evaluated on all six Open Universal Arabic ASR Leaderboard test sets using the official leaderboard code (same normalizer, same WER metric).
Per-set WER (%)
| SADA | Common Voice 18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca | Average |
|---|---|---|---|---|---|---|
| 37.28 | 9.74 | 7.27 | 23.65 | 14.33 | 58.24 | 25.08 |
In context (Average WER %, lower is better)
| Model | Params | Avg WER |
|---|---|---|
| cohere-transcribe-arabic-07-2026 | ~2.0B | 25.87 |
| lemura-arabic-asr | ~0.12B | 25.08* |
| omniASR_LLM_7B | 7B | 28.32 |
| Qwen3-Omni-30B-A3B | 30B | 30.71 |
| nvidia-conformer-ctc-large-arabic (lm) | 0.6B | 32.91 |
| Qwen3-ASR-1.7B | 1.7B | 33.36 |
* Transparency note. lemura-arabic-asr was fine-tuned on training splits that overlap several of these test corpora (SADA, MASC, CV, MGB-2). The leaderboard is designed to measure zero-shot generalization, so this 25.08 is best read as an in-domain result, not a strict zero-shot ranking β a fair zero-shot number is a few points higher. Its unqualified strengths are efficiency and Gulf / MSA / broadcast robustness; the frontier for every system remains Maghrebi Darija (Casablanca).
β‘ Why it matters: efficiency
The top leaderboard systems are large generative audio-LLMs (2β30B parameters) that need GPUs. lemura-arabic-asr reaches a comparable accuracy tier with a ~115M-parameter CTC model:
| lemura-arabic-asr | Typical top systems | |
|---|---|---|
| Parameters | ~115M | 2B β 30B |
| Hardware | CPU or GPU | GPU |
| Latency | real-time | seconds / clip |
| Footprint | ~0.4 GB | 4 β 60 GB |
That makes it practical for on-device, low-cost, and high-throughput Arabic transcription where the big models are impractical.
π» Usage
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"])) # 16 kHz mono
Try it live β no install: ποΈ lemura-arabic-asr demo
π Credits
- Acoustic foundation: NVIDIA FastConformer.
- Evaluation: the Open Universal Arabic ASR Leaderboard official code.
- Downloads last month
- -