Ghazouaniwala's picture
Upload folder using huggingface_hub
7bcdfb4 verified
|
Raw
History Blame Contribute Delete
1.43 kB
metadata
license: apache-2.0
language:
  - ar
  - fr
  - en
tags:
  - text-to-speech
  - tts
  - f5-tts
  - silma-tts
  - tunisian
  - derja
  - arabic
base_model: silma-ai/silma-tts

SILMA TTS — Tunisian Derja fine-tune (v2, quality-focused)

Fine-tune of silma-ai/silma-tts on the LinTO Tunisian dataset.

v2 recipe: speaker_mode=multi (single≈AbdelAzizErwi), moderate text normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False, 15 epochs. Targets pronunciation/vowel/consonant quality by improving the signal rather than the architecture. Deploy raw or EMA weights per the A/B in §8.

Usage

Load with the F5-TTS v1.1.7 / SILMA pipeline: model.pt + vocab.txt + config.yaml. The vocab.txt here matches this run (SILMA char vocab, or phonemized vocab if piloted).

Attribution (required — CC BY 4.0)

@misc{linagora2024Linto-tn,
  title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect},
  author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre},
  year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL}
}

Base: SILMA TTS (Apache-2.0). Data: LinTO (CC BY 4.0).

Responsible use

Voice cloning requires documented speaker consent. Do not use for deception or impersonation.