--- license: other license_name: minimax-music3-community-license license_link: LICENSE base_model: MiniMaxAI/MiniMax-Music3 base_model_relation: quantized library_name: mlx-serve tags: - mlx - mlx-serve - music - text-to-music - minimax pipeline_tag: text-to-audio --- # MiniMax Music 3 for mlx-serve (8-bit) [MiniMax-Music3](https://huggingface.co/MiniMaxAI/MiniMax-Music3) converted for [mlx-serve](https://github.com/ddalcu/mlx-serve)'s native Zig + MLX engine. Full songs with sung lyrics at 44.1 kHz stereo, generated locally on Apple Silicon. Quantization: affine 8-bit, group 64, on every real matmul (LLM, depth decoder, DiT, lm_head). Kept dense on purpose: the embedding tables (gather reads), the condition encoder and the whole vocoder (VAE-class precision). Worst per-tensor reconstruction error 1.48% RMS. 13 GB on disk instead of the 57 GB upstream repo (which ships its weights twice). | File | Contents | |---|---| | `language_model.safetensors` | Qwen3 8B global LLM, 36L, vocab 200k | | `rvq_depth_decoder.safetensors` | 0.6B local LLM, 7 residual codebooks | | `transformer.safetensors` | 2.4B flow-matching DiT, 36 blocks | | `condition_encoder.safetensors` | hidden-state mix + resampler, f32 | | `vocoder.safetensors` | Flow-VAE / DAC decoder, f32 | Engine parity against the fp32 reference on these exact weights: prefill cosine 0.9999, condition encoder 0.999999, DiT velocity 0.999, vocoder 1.000000. The autoregressive stage runs about 44 ms per frame on an M-series Mac, so a one-minute song takes roughly a minute of LLM time plus the diffusion pass. ## Run it Download **[MLX Core.app](https://github.com/ddalcu/mlx-serve/releases/latest)**, open the Music tab and pick **MiniMax Music 3**. Style prompt + lyrics in, WAV out. Over HTTP: ```bash mlx-serve --serve curl http://127.0.0.1:11234/v1/audio/music-generations \ -H 'Content-Type: application/json' -o song.wav -d '{ "model": "MiniMax-Music3-MLX-Serve-8bit", "prompt": "upbeat synthwave with driving bass and dreamy pads", "lyrics": "[verse]\nneon lights across the bay\n[chorus]\nwe run all night", "duration_seconds": 60 }' ``` Lyrics are required (the model is lyric-conditioned) and structure tags like `[verse]` or `[chorus]` go on their own lines. `duration_seconds` (1-360) is an upper bound, the model may end the song earlier. ACE-Step style fields (bpm, keyscale, timesignature, vocal_language) do not exist on this model. Rebuild from the upstream repo with [`scripts/convert_music3_weights.py`](https://github.com/ddalcu/mlx-serve/blob/main/scripts/convert_music3_weights.py). Weights are covered by the MiniMax-Music3 Community License (see `LICENSE`, acceptable-use policy included as Exhibit A).