--- title: MiniMax-H3 Reference emoji: 🎭 colorFrom: pink colorTo: purple sdk: gradio sdk_version: 6.20.0 app_file: app.py pinned: true short_description: Unquantized MiniMax-H3 from image, audio, video refs suggested_hardware: zero-a10g --- # MiniMax-H3 — omni-references, unquantized, split across two Spaces Joint video **and** soundtrack out of a single denoising pass, conditioned on an ordered list of image, video and audio references, at **bfloat16 with no quantization anywhere**. This Space is the denoising half of the `ref2va` task: the 61.73 GiB `transformer_ref` partition and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in [`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this Space calls over the gradio API for every request — the same conditioner Space, and the same resident weights, that the keyframe half [`minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3) uses. ## Why split MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at **150 GB of storage**. An unquantized single Space is therefore impossible. Cut the `ref2va` branch of `MiniMaxH3Blocks` at its `text_encoder` step and both halves fit unquantized: | Space | Subfolders | Download | Resident | |---|---|---|---| | [`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) | `text_encoder/` + `tokenizer/` + `processor/` | 66.7 GB | 62.15 GiB bf16 | | this one | `transformer_ref/` + `vae/` + `audio_vae/` | 77.3 GB | 61.73 GiB bf16 + 10.43 GiB float32 | ## References A request carries up to **12** references — at most 9 images, 3 videos and 3 audio clips — **in the order the model reads them**. The order is semantic: it numbers the labels of MiniMax-H3's prompt presentation (``, `