Ghazouaniwala's picture
Upload folder using huggingface_hub
7bcdfb4 verified
|
Raw
History Blame Contribute Delete
1.43 kB
---
license: apache-2.0
language: [ar, fr, en]
tags: [text-to-speech, tts, f5-tts, silma-tts, tunisian, derja, arabic]
base_model: silma-ai/silma-tts
---
# SILMA TTS — Tunisian Derja fine-tune (v2, quality-focused)
Fine-tune of [silma-ai/silma-tts](https://huggingface.co/silma-ai/silma-tts) on the
[LinTO Tunisian dataset](https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn).
**v2 recipe:** speaker_mode=`multi` (single≈`AbdelAzizErwi`), moderate text
normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False,
15 epochs. Targets pronunciation/vowel/consonant quality by improving the *signal*
rather than the architecture. Deploy **raw** or **EMA** weights per the A/B in §8.
## Usage
Load with the F5-TTS v1.1.7 / SILMA pipeline: `model.pt` + `vocab.txt` + `config.yaml`.
The `vocab.txt` here matches this run (SILMA char vocab, or phonemized vocab if piloted).
## Attribution (required — CC BY 4.0)
```bibtex
@misc{linagora2024Linto-tn,
title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect},
author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre},
year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
```
Base: SILMA TTS (Apache-2.0). Data: LinTO (CC BY 4.0).
## Responsible use
Voice cloning requires documented speaker consent. Do not use for deception or impersonation.