SILMA TTS — Tunisian Derja fine-tune (v2, quality-focused)

Fine-tune of silma-ai/silma-tts on the LinTO Tunisian dataset.

v2 recipe: speaker_mode=multi (single≈AbdelAzizErwi), moderate text normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False, 15 epochs. Targets pronunciation/vowel/consonant quality by improving the signal rather than the architecture. Deploy raw or EMA weights per the A/B in §8.

Usage

Load with the F5-TTS v1.1.7 / SILMA pipeline: model.pt + vocab.txt + config.yaml. The vocab.txt here matches this run (SILMA char vocab, or phonemized vocab if piloted).

Attribution (required — CC BY 4.0)

@misc{linagora2024Linto-tn,
  title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect},
  author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre},
  year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL}
}

Base: SILMA TTS (Apache-2.0). Data: LinTO (CC BY 4.0).

Responsible use

Voice cloning requires documented speaker consent. Do not use for deception or impersonation.

Downloads last month
27
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ghazouaniwala/silma-tts-derja-v2-1

Finetuned
(5)
this model

Paper for Ghazouaniwala/silma-tts-derja-v2-1