--- license: apache-2.0 language: [ar, fr, en] tags: [text-to-speech, tts, f5-tts, silma-tts, tunisian, derja, arabic] base_model: silma-ai/silma-tts --- # SILMA TTS — Tunisian Derja fine-tune (v2, quality-focused) Fine-tune of [silma-ai/silma-tts](https://huggingface.co/silma-ai/silma-tts) on the [LinTO Tunisian dataset](https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn). **v2 recipe:** speaker_mode=`multi` (single≈`AbdelAzizErwi`), moderate text normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False, 30 epochs. Targets pronunciation/vowel/consonant quality by improving the *signal* rather than the architecture. Deploy **raw** or **EMA** weights per the A/B in §8. ## Usage Load with the F5-TTS v1.1.7 / SILMA pipeline: `model.pt` + `vocab.txt` + `config.yaml`. The `vocab.txt` here matches this run (SILMA char vocab, or phonemized vocab if piloted). ## Attribution (required — CC BY 4.0) ```bibtex @misc{linagora2024Linto-tn, title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect}, author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre}, year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` Base: SILMA TTS (Apache-2.0). Data: LinTO (CC BY 4.0). ## Responsible use Voice cloning requires documented speaker consent. Do not use for deception or impersonation.