LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
Paper • 2504.02604 • Published • 1
How to use Ghazouaniwala/silma-tts-derja-v3a with F5-TTS:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
Fine-tune of silma-ai/silma-tts on the LinTO Tunisian dataset.
v2 recipe: speaker_mode=multi (single≈AbdelAzizErwi), moderate text
normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False,
30 epochs. Targets pronunciation/vowel/consonant quality by improving the signal
rather than the architecture. Deploy raw or EMA weights per the A/B in §8.
Load with the F5-TTS v1.1.7 / SILMA pipeline: model.pt + vocab.txt + config.yaml.
The vocab.txt here matches this run (SILMA char vocab, or phonemized vocab if piloted).
@misc{linagora2024Linto-tn,
title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect},
author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre},
year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
Base: SILMA TTS (Apache-2.0). Data: LinTO (CC BY 4.0).
Voice cloning requires documented speaker consent. Do not use for deception or impersonation.
Base model
silma-ai/silma-tts
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js