MiniMax H3 clean-latent 2ร— upscaler

This is not a conventional image or video upscaler. It is not intended to make a finished render sharper or better-looking, and a direct decode of its output can look softer than the original. Its purpose is to move a clean MiniMax H3 video latent to a 2ร— larger spatial latent grid very quickly and efficiently, so an H3 workflow can stay in latent space instead of doing a VAE decode โ†’ pixel resize โ†’ VAE re-encode round trip. Use a conventional pixel-space upscaler when the goal is to enhance a finished video.

The model accepts a fully denoised MiniMax H3 video latent and doubles only its spatial latent dimensions. Temporal length is unchanged; when used through the companion ComfyUI node, the audio latent is preserved unchanged.

How it was trained

Training pairs were built from clean H3 latents: each low-resolution latent was decoded with the H3 VAE, enlarged 2ร— in pixel space with Lanczos, then deterministically re-encoded to provide the teacher latent. The lightweight network learned a correction on top of bilinear latent interpolation using latent and decoder-aware reconstruction, SSIM, spatial-consistency and temporal-consistency losses; the H3 generator itself was not trained or modified.

ComfyUI

Use the model with ComfyUI-H3-Latent-Upscaler-Mamad8. Place h3_clean_latent_upscaler_v1_mamad8.safetensors in:

ComfyUI/models/h3_latent_upscalers/

The checkpoint supports clean H3 latents only. Do not apply it to intermediate noisy latents. Continuing H3 generation after the resize requires an explicit re-noising and high-resolution continuation workflow.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support