lemura-arabic-asr-lite

Format Task Type License

A compact, dialect-aware Arabic speech-recognition model โ€” leaderboard-tier accuracy at ~18ร— fewer parameters.

Task Params Architecture Dialects Runtime Overview ยท Model Details ยท Benchmarks ยท Efficiency ยท Usage ยท Live Demo


Overview

lemura-arabic-asr is a compact multi-dialect Arabic ASR model built for accuracy and efficiency. It is a FastConformer-CTC acoustic model (~115M parameters), adapted in-house from the NVIDIA FastConformer foundation and fine-tuned on ~2,900 hours of Arabic spanning MSA and the Gulf, Egyptian, Levantine, and Maghrebi dialect groups.

Highlights

  • Small & fast โ€” ~115M parameters; runs comfortably on CPU and in real time, no GPU required.
  • Dialect-aware โ€” trained across five Arabic dialect groups, not MSA-only.
  • Robust on real audio โ€” strongest on broadcast, conversational, and Gulf/MSA speech.
  • Open & simple โ€” a single .nemo file, loadable in a few lines with NVIDIA NeMo.

Model Details

Model lemura-arabic-asr โ€” compact multi-dialect Arabic ASR
Task Automatic speech recognition (audio โ†’ text)
Architecture FastConformer encoder + CTC decoder
Parameters ~115M
Foundation Adapted from NVIDIA FastConformer; ~2,900 h Arabic fine-tuning
Audio input 16 kHz mono (auto-resampled)
Languages Arabic โ€” MSA + Gulf / Egyptian / Levantine / Maghrebi
Runtime NVIDIA NeMo โ€” CPU ยท GPU ยท real-time
License CC-BY-4.0

Benchmarks

Evaluated on all six Open Universal Arabic ASR Leaderboard test sets using the official leaderboard code (same normalizer, same WER metric).

Per-set WER (%)

SADA Common Voice 18 MASC-clean MASC-noisy MGB-2 Casablanca Average
37.28 9.74 7.27 23.65 14.33 58.24 25.08

In context (Average WER %, lower is better)

Model Params Avg WER
cohere-transcribe-arabic-07-2026 ~2.0B 25.87
lemura-arabic-asr ~0.12B 25.08*
omniASR_LLM_7B 7B 28.32
Qwen3-Omni-30B-A3B 30B 30.71
nvidia-conformer-ctc-large-arabic (lm) 0.6B 32.91
Qwen3-ASR-1.7B 1.7B 33.36

Efficiency

The top leaderboard systems are large generative audio-LLMs (2โ€“30B parameters) that need GPUs. lemura-arabic-asr reaches a comparable accuracy tier with a ~115M-parameter CTC model:

lemura-arabic-asr Typical top systems
Parameters ~115M 2B โ€“ 30B
Hardware CPU or GPU GPU
Latency Real-time Seconds / clip
Footprint ~0.4 GB 4 โ€“ 60 GB

That makes it practical for on-device, low-cost, and high-throughput Arabic transcription where the big models are impractical.

Usage

import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"])) # 16 kHz mono

Try it live โ€” no install: lemura-arabic-asr demo

Credits

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using lemuralabs/lemura-arabic-asr-lite 1

Evaluation results