File size: 14,163 Bytes
d11a565
e63a181
186aa49
 
 
d11a565
df357c1
d11a565
ef0bfcc
f2b41a2
cba1972
186aa49
d11a565
 
186aa49
 
 
 
 
 
9da256f
 
 
186aa49
 
 
 
9da256f
 
186aa49
 
 
9da256f
0453841
186aa49
 
 
 
f2b41a2
299d59b
f2b41a2
 
 
 
 
 
 
299d59b
d8b48ff
 
 
942d78c
65a89d4
456aa27
 
 
 
 
d8b48ff
c38b474
 
 
 
2b67062
 
 
 
 
b683692
e2db4ca
 
 
 
b683692
578be89
 
 
a0bf481
578be89
 
 
 
 
 
 
 
 
 
 
 
 
0453841
 
 
 
186aa49
 
9da256f
 
186aa49
 
9da256f
186aa49
 
9da256f
 
186aa49
 
 
 
 
 
9da256f
186aa49
 
 
 
 
 
 
 
9da256f
 
 
 
 
 
186aa49
 
 
5ca1dce
 
 
186aa49
5ca1dce
0453841
5ca1dce
 
 
 
0453841
 
 
5ca1dce
0453841
 
5ca1dce
 
 
 
 
 
 
0453841
186aa49
 
 
 
 
0453841
 
 
 
 
 
 
 
 
5ca1dce
0453841
 
 
 
 
c938c2e
 
 
 
 
 
 
 
 
 
 
 
186aa49
 
 
 
9da256f
 
5ca1dce
186aa49
 
 
65a89d4
d8b48ff
456aa27
c38b474
 
2b67062
b683692
 
186aa49
9c16554
 
bf1199a
5029048
 
 
bf1199a
 
5029048
 
 
9c16554
9da256f
186aa49
9a812e9
 
 
 
186aa49
 
 
9da256f
5029048
 
9da256f
5029048
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
---
title: MiniMax H3
emoji: 🎬
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.23.1
app_file: app.py
pinned: true
hf_oauth: true
short_description: Video generation with a synchronized soundtrack
suggested_hardware: zero-a10g
---

# MiniMax-H3 — unquantized, split across two Spaces

Joint video **and** soundtrack out of a single denoising pass, at **bfloat16 with no quantization anywhere**.

This Space is the denoising half: the 61.73 GiB transformer and the two autoencoders. The 62.14 GiB Qwen3-VL
conditioner runs in
[`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
Space calls over the gradio API for every request. The weights are the public
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3) diffusers checkpoint.

## Why split

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at **150 GB of storage**. An unquantized single
Space is therefore impossible, which is why quantized demos of it run NVFP4 or float8 weights. Cut the
`MiniMaxH3Blocks` sequence at its `text_encoder` step and both halves fit unquantized:

| Space | Subfolders | Download | Resident |
|---|---|---|---|
| [`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) | `text_encoder/` + `tokenizer/` + `processor/` | 66.7 GB | 62.15 GiB bf16 |
| this one | `transformer/` + `vae/` + `audio_vae/` | 77.3 GB | 61.73 GiB bf16 + 10.43 GiB float32 |

Besides the quality argument, unquantized weights are the ones AoTI can export; an NVFP4 checkpoint cannot be
exported at all.

## Workflow frontend (gr.Workflow)

The UI is a [`gr.Workflow`](https://www.gradio.app/docs/gradio/workflow) canvas (`workflow.json`): reference nodes for
prompt, optional first/last keyframes, canvas, duration, steps, seed, prompt upsampling and the LoRA selector feed a
single `generate_video` fn operator, which runs the conditioner call and the `@spaces.GPU` denoise and returns the
muxed MP4 plus the run report and refined prompt to subject nodes. With `hf_oauth: true` the owner can edit and save
the graph; visitors get a runnable read-only canvas, and the pipeline is also exposed through the standard Gradio
API. Keyframe cover-crop / canvas fitting runs server-side in `generate`, so canvas rewires and API callers get the
same treatment.

## 4-step Turbo LoRA

The transformer runs with [`larryvrh/MiniMax-H3-Turbo-Lora`](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
folded into its bf16 weights at startup (`h3_lora.py`), so the default is **6 sampling steps** instead of 28 
(the v4-600 EMA checkpoint, the card's current recommendation; 6–8 steps is where v4 looks its best). The larry fold mirrors the diffusers key
conversion exactly (fused-QKV thirds, the `SwiGLU` gate/value swap, the shared AdaLN row layout); lightx keys are
already diffusers-native. Both happen before the AoTI package is patched in, so compiled blocks carry the update too,
and the low-rank factors stay resident so switching is an in-place unfold/fold through one bf16 rounding.
`H3_LORA` selects the larry file (`off` skips it), `H3_LIGHTX=off` skips lightx, `H3_LORA_DEFAULT` picks which set
starts folded, and `H3_LORA_STRENGTH` scales the larry update (the card's sharpness/artifact dial).

The lightx set also comes as `lightx8` — the 8-step v1.0 diffusers checkpoint
(`minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors`), same PEFT layout and scale, suggested 8 steps.
`H3_LIGHTX8=off` skips it.

A third, non-turbo set is [`fal/MiniMax-H3-Realism-People-LoRA`](https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA)
(`realism` in the dropdown, trigger word `r34l1sm`): realistic people at the full 28 steps. It ships in the same
reference-tree layout as larry under a `diffusion_model.` prefix, so it folds through the same key mapping.
`H3_REALISM=off` skips loading it.

A fourth set is [`joyfox/MiniMax-H3-Turbo`](https://huggingface.co/joyfox/MiniMax-H3-Turbo) (`joyfox`, 4 steps):
another ComfyUI-layout turbo LoRA, additionally covering the MLPs and both output heads (`video_out` -> `proj_out`,
`audio_out` -> `audio_proj_out`), with a per-key `alpha` folded into `lora_B` at load. Its AdaLN deltas are skipped —
the Comfy-Org checkpoint uses an 8-dim modulation input where diffusers uses the full 2688-dim time embedding, so
they have no fold target. `H3_JOYFOX=off` skips loading it.

## AoTI-compiled blocks

With `H3_AOTI=1` the 50 repeated transformer blocks run from a compiled package,
[`multimodalart/minimax-h3-aoti`](https://huggingface.co/multimodalart/minimax-h3-aoti)`:bf16/torch2.11/sm120/dynamic` — a single dynamic-sequence artifact that serves
every canvas, duration and prompt length. It carries no weights (it reads each block's live ones), so patching it in
is startup CPU work and costs no GPU time.

It removes a near-constant ~0.5 s/step — 50 blocks' worth of kernel-launch overhead plus the norm / rotary / AdaLN
epilogues around the matmuls — and cannot touch the matmuls themselves. So it pays best where the block is *not*
compute bound, i.e. on the small canvases:

| canvas (HxW) | eager s/step | AoTI s/step | faster |
|---|---|---|---|
| 768x1344 | 10.20 | 9.73 | +4.6% |
| 640x1152 | 6.46 | 5.88 | +9.1% |
| 544x960 | 4.02 | 3.58 | +11.0% |

At 72.16 GiB of weights on a 95.0 GiB card there is no offloading in the request path at all, which is why an
*unquantized* MiniMax-H3 is also the fastest one measured here: **10.6 s/step** against the 19–21 s/step the 4 bit
Space pays, where the whole cost is the traffic auto-offload has to move.

## How the split is expressed

`MiniMaxH3Blocks` is a `SequentialPipelineBlocks` whose branches are picked per request — and per `workflow=` — from
the inputs:

```
before_encode -> text_encoder -> vae_encoder -> denoise -> after_denoise -> decode
```

where `denoise` is itself `prepare_layout -> prepare_latents -> set_timesteps -> denoise`.

`h3_split_blocks.py` subclasses it with the `text_encoder` step removed. Dropping the step drops the three
components it declares, so `load_components` resolves `transformer` / `vae` / `audio_vae` / the two schedulers out of
the shared `modular_model_index.json` and never fetches the conditioner — and `prompt_embeds` and `text_token_tags`
become ordinary required inputs of the pipeline call:

```py
pipe = MiniMaxH3GeneratorBlocks().init_pipeline("MiniMaxAI/MiniMax-H3")
pipe.load_components(dtype=torch.bfloat16)
state = pipe(prompt_embeds=..., text_token_tags=..., height=768, width=1344, num_frames=124, num_inference_steps=30)
```

The wire format is exactly those two tensors — `(1, num_text_tokens, 5120)` bfloat16 and `(num_text_tokens,)` int64 —
carried as one safetensors file with the resolved `height` / `width` / `num_frames` in its metadata header. A
text-only request is 246 KB of it; one 768x1344 keyframe adds 1016 vision rows and takes it to 10.7 MB.

The keyframe `resize` step runs on **both** halves. It owns no pretrained component (PIL and arithmetic) and it puts
the keyframes onto the target canvas — which the conditioner needs to build its vision blocks and this Space needs to
encode with the video VAE. It is deterministic, and the conditioner returns the plan it resolved so this Space pins
the same canvas rather than re-deriving it. Two things that step no longer does, and that both halves therefore do
themselves: EXIF-transposing a keyframe into upright RGB, and snapping `num_frames` to `17 * n + 5` — the frame count
is resolved by the layout step, which lives on this side of the cut.

## Nothing is paid for with GPU time

The 77.3 GB download and the load happen at **startup**: `import spaces` at module top patches `torch.cuda` before
any GPU is attached, so nothing about the load needs a card. The conditioner round trip is a network call on this
Space's CPU. A `@spaces.GPU` call is therefore only the placement (once) and the denoise loop and the two decoders.

### The 150 GB quota, not the 95 GiB card, is what rules out startup placement

One thing does *not* happen at startup: the move onto the card. `spaces`' startup `torch.pack()` writes every
startup-resident CUDA tensor to a **second copy on disk** and only deletes the downloaded originals afterwards
(`Cleaned 62.13GB of tensor files ... after packing`, which is what keeps the conditioner half comfortable at
66.7 GB). Packing 77.3 GB needs 154.6 GB at once, and this Space is evicted mid-pack:

```
ZeroGPU tensors packing:   0%|          | 0.00/77.3G
OSError: [Errno 28] No space left on device      # os.posix_fallocate, spaces/zero/torch/packing.py
```

Unlinking the shards first does not rescue it. `.to("cuda")` under the startup patch does not release the
memory-mapped safetensors, so nothing is freed — and the pack's own cleanup walks those still-open mappings and
`lstat`s them, so a deleted blob becomes `FileNotFoundError: .../blobs/3d449... (deleted)`.

Placement therefore happens on the **first GPU call**, `PIPE.to("cuda")` at the top of the `@spaces.GPU` function:
about 10 s of PCIe once, then a no-op walk, and the denoise loop runs with everything resident and no offloading at
all. It is the same trick the 4 bit Space uses, for the same reason.

## Generation constraints

Fixed by the checkpoint: 24 fps, a 768 pixel short edge, 5 to 15 s, `num_frames` snapped up to the next `17 * n + 5`,
no CFG and no negative prompt (it is guidance-distilled, so every step is one forward pass).

## Measured

An `rtx-pro-6000` Job — the same silicon as the ZeroGPU pool (RTX PRO 6000 Blackwell, sm120, 95.0 GiB) — running
exactly this blockset over the wire format, 1344x768, 124 frames, 30 steps, bfloat16, cuDNN attention, everything
resident:

| | |
|---|---|
| `load_components` (77.3 GB, warm Xet) | 43 s |
| `.to("cuda")`, once | 10 s |
| resident weights | 72.16 GiB |
| denoise + decode | 317 s, **10.58 s/step** |
| peak allocated / reserved | 78.54 / 85.37 GiB |
| output | h264 1344x768 @ 24 fps, 5.167 s + stereo AAC @ 32 kHz |

And on this Space itself, driven over `gradio_client`. Startup is 93 s — the 77.3 GB download and the load, with
no placement and therefore no pack.

| Request | Conditioner | Denoise + decode | Steady | Round trip |
|---|---|---|---|---|
| text only, 18 tokens | 7 s | 339 s | 10.53 s/step | 353 s |
| one 768x1344 keyframe, 1034 tokens | 9 s | 370 s | 11.39 s/step | 386 s |

The keyframe costs about 8% per step rather than a placement penalty: it puts 1016 vision rows in front of the
prompt *and* 1016 conditioning rows in the packed sequence, and MiniMax-H3 attends over all of it every layer. The
one-time `PIPE.to("cuda")` is inside the first row's 339 s and does not reappear in the second.

## Space variables

| Variable | Default | Meaning |
|---|---|---|
| `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout checkpoint. Public. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
| `H3_PLACEMENT` | `lazy` | `lazy` moves all 72.16 GiB onto the card on the first GPU call and leaves it there; `offload` hands placement to `ComponentsManager.enable_auto_cpu_offload` instead. |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel, 10–20% faster than the SDPA default and needs nothing installed. flash-attention 3 is sm90-only and this pool is sm120. |
| `H3_GPU_DURATION` | `900` | Seconds per request; the pool applies a 1.5 duration factor. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
| `H3_LORA` | `minimax_h3_turbo_v4_step600_ema.safetensors` | Turbo LoRA file folded into the transformer at startup. `off` disables. |
| `H3_LORA_REPO` | `larryvrh/MiniMax-H3-Turbo-Lora` | Hub repo the LoRA is fetched from. |
| `H3_LORA_STRENGTH` | `1.0` | Scales the larry LoRA delta (sharpness/artifact trade-off). |
| `H3_LIGHTX` | `on` | Set to `off` to skip loading the lightx2v 4-step LoRA set. |
| `H3_LIGHTX8` | `on` | Set to `off` to skip loading the lightx2v 8-step v1.0 LoRA set. |
| `H3_REALISM` | `on` | Set to `off` to skip loading the fal realism-people LoRA set. |
| `H3_JOYFOX` | `on` | Set to `off` to skip loading the joyfox 4-step turbo LoRA set. |
| `H3_LORA_DEFAULT` | `larry` | Which loaded LoRA set starts folded (`larry` / `lightx` / `realism` / `joyfox`). |

## Whose GPU quota pays

Two cards are booked per request — this Space's denoise loop and the conditioner's forward — and both are billed to the
requesting user, with nothing here arranging it: `gradio_client` attaches the caller's own `x-ip-token` to every
outgoing call, reading it off gradio's `LocalContext` inside the event listener (`Client.send_data` ->
`add_zero_gpu_headers`), and ZeroGPU charges the booking to whatever that token identifies.

A caller with no token to forward — a `gradio_client` script rather than a browser — leaves the conditioner's booking
attributed to this Space's pod IP and its small shared quota. An unattributed caller may book at most 120 credits at a
time and an `xlarge` booking costs twice its seconds, so the conditioner books the encode (45 s) and a prompt upsample
(60 s) as two separate calls, each within that ceiling.

## Secrets

None are required. Everything this Space downloads is public — the
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3) checkpoint and the
[`multimodalart/minimax-h3-aoti`](https://huggingface.co/multimodalart/minimax-h3-aoti) packages — and the conditioner
is a public Space called on the requesting user's own ZeroGPU token, never on an org token.

## Where diffusers comes from

MiniMax-H3 is modular-only and not in a released `diffusers`, so `requirements.txt` installs it from the canonical
pull request, [huggingface/diffusers#14371](https://github.com/huggingface/diffusers/pull/14371), pinned to the commit
`665f5782` (`refs/pull/14371/head`) rather than to the moving `minimax-h3-refactor` branch.

That PR is a WIP, so it needs re-pinning whenever it updates, and `h3_split_blocks.py` — which subclasses its block
classes to cut the pipeline in two — has to be re-checked against the new head at the same time.