Spaces:
Running on Zero
title: OmniVAE T2AV
emoji: 🎬
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.24.0
app_file: app.py
short_description: Text to synchronized audio + video with OmniVAE
python_version: '3.12'
startup_duration_timeout: 1h
models:
- OpenMOSS-Team/OmniVAE
tags:
- text-to-video
- text-to-audio
- audio-video
- omnivae
license: apache-2.0
OmniVAE — Text to Audio + Video
Generate a short clip with natively synchronized sound from a single text prompt.
- Model:
OpenMOSS-Team/OmniVAE(Apache-2.0) - Code: OpenMOSS/OmniVAE (Apache-2.0)
OmniVAE is a unified audio-video tokenizer. This Space runs the released
t2av_recon_distill_avclip joint text-to-audio-video model built on top of it: two
Z-Image diffusion-transformer branches (video + audio) coupled by bridge cross-attention,
so the picture and the soundtrack are denoised together rather than dubbed afterwards.
What it produces
The checkpoint's native shape: 121 frames · 256×256 · 24 fps · 48 kHz audio (~5 s), muxed into a single MP4. A default 50-step run takes ~97 s on ZeroGPU (measured: ~1.9 s/step plus ~3 s for text encoding, VAE decoding and muxing).
Faithfulness to the reference implementation
The app is a thin wrapper around the upstream inference path
(generation/infer/t2av/t2av_pipeline.py → load_joint_av_pipeline / generate_one_av)
with the defaults from the release smoke test in generation/docs/inference.md
(validate_checkpoints.sh --cfg 4):
| setting | value |
|---|---|
| mode | joint_av |
| CFG | BridgeDiT dual CFG (NFE = 3) |
| steps | 50 |
| guidance (text + cross-modal, both branches) | 4.0 |
| shape | 121 frames, 256×256, 24 fps, 5.04 s audio |
| prompt wrapping | t2av task prefix on, duration suffix off |
Deployment-only deltas from the upstream tree, all of which leave the numerics untouched:
omnivae_generation/trainer/eval/guided_diffusion.py— the upstream repo imports three small checkpoint/run-dir helpers from an evaluation module it does not publish; they are re-implemented here.omnivae_generation/trainer/vae/_audiotools_compat.py— stand-ins for the twodescript-audiotoolssymbols the DAC audio VAE imports (ml.BaseModel,AudioSignal). The real package pinsprotobuf<3.20; neither symbol carries behaviour on the continuous-VAE inference path.joint_av/__init__.py— the training-onlytorchcodec-backed dataset import is made optional.joint_av/validation.py— the denoising loop is wrapped intqdmso the UI can show progress.torchaudio.saveis routed throughsoundfileinapp.py: torchaudio ≥ 2.10 delegatessavetotorchcodec, whose wheels are bound to a specific FFmpeg ABI. Both write the same 48 kHz PCM WAV.- Components are loaded on CPU and then moved with
.to("cuda"); upstream'sdevice="cuda"path usesdevice_map={"": "cuda:0"}+torch.cuda.set_device, which ZeroGPU does not support.
Example prompts
Taken from the authors' own material: generation/examples/prompts/t2av_valid.jsonl and
the authored prompt list in the released
models/dit/t2av/t2av_recon_distill_avclip/resolved_config.json.