Text-to-Speech
Transformers
Safetensors
higgs_multimodal_qwen3
text-generation
speech-generation
voice-agent
expressive-speech
controllable-tts
multilingual-tts
Instructions to use bosonai/higgs-tts-3-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bosonai/higgs-tts-3-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="bosonai/higgs-tts-3-4b")# Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("bosonai/higgs-tts-3-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
How to use Control Tokens?
#13
by Syams - opened
I'm still confused; sometimes emotions or other control tokens don't produce any results.
I have two questions here:
- Does <|emotion:...|> have to be followed by <|style:...|>, <|sfx:...|>?
- What's the benefit of writing "sob" after <|sfx:crying|>? Because in Indonesian, "sob" means "friend." This makes the speech lose its context. Pay attention to the 4th second.
Could you please take a chance to read "README.md", "PROMPTING.md", "AGENTS.md"?
