MLX Speech Models
Collection
Speech AI models for Apple Silicon via MLX. ASR, TTS, VAD, diarization, speaker embedding. โข 98 items โข Updated โข 6
How to use aufklarer/Qwen3-TTS-12Hz-1.7B-Base-MLX-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3-TTS-12Hz-1.7B-Base-MLX-bf16 aufklarer/Qwen3-TTS-12Hz-1.7B-Base-MLX-bf16
Full-precision bf16 (non-quantized) MLX conversion of Qwen/Qwen3-TTS-12Hz-1.7B-Base for Apple Silicon inference โ the highest-quality 1.7B variant.
Used by speech-swift Qwen3TTS module:
let model = try await Qwen3TTSModel.fromPretrained(
modelId: "aufklarer/Qwen3-TTS-12Hz-1.7B-Base-MLX-bf16"
)
let audio = try model.synthesize("Hello, world!")
audio speak "Hello, world!" --model 1.7b -o output.wav
Linear weights)Apple Silicon (M-series), MLX, 1.7B variants:
| Precision | RTF | Peak RAM | Notes |
|---|---|---|---|
| 8-bit | 0.39 | 2.8 GiB | good |
| bf16 | 0.48 | 4.1 GiB | best quality |
Round-trip WER (synthesize โ Qwen3-ASR transcribe โ WER), 15 sentences: 7.05 % (TTS + ASR round-trip; 0 synthesis failures).
| Variant | Precision | Size | Model ID |
|---|---|---|---|
| 0.6B 8-bit | 8-bit | ~1.3 GB | aufklarer/Qwen3-TTS-12Hz-0.6B-Base-MLX-8bit |
| 1.7B 8-bit | 8-bit | ~2.8 GB | aufklarer/Qwen3-TTS-12Hz-1.7B-Base-MLX-8bit |
| 1.7B bf16 | bf16 | ~3.7 GB | aufklarer/Qwen3-TTS-12Hz-1.7B-Base-MLX-bf16 |
Quantized
Base model
Qwen/Qwen3-TTS-12Hz-1.7B-Base