OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 16 days ago • 27
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Paper • 2606.07639 • Published Jun 1 • 5
laion/moss-tts-local-transformer-4.55b-voice-acting Text-to-Speech • 4B • Updated 24 days ago • 3.04k • 1
MOSS Transcribe Collection A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription. • 4 items • Updated about 1 month ago • 13
OpenMOSS-Team/MOSS-Transcribe-preview-2B Automatic Speech Recognition • 2B • Updated Jun 26 • 1.71k • 47