Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Paper • 2605.04128 • Published May 5 • 18
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 14 days ago • 76
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 17 days ago • 138
NeuroCogMap Reveals Cognitive Organization of Large Language Models Paper • 2607.00397 • Published Jul 1 • 12
Scalable Visual Pretraining for Language Intelligence Paper • 2607.09657 • Published 25 days ago • 57
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 25 days ago • 86
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 154
DiffusionBench: On Holistic Evaluation of Diffusion Transformers Paper • 2606.24888 • Published Jun 23 • 11
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning Paper • 2606.24548 • Published Jun 23 • 11
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective Paper • 2605.28819 • Published May 27 • 9
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Paper • 2605.21573 • Published May 20 • 111
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Paper • 2605.21468 • Published May 20 • 51
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 195
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 173
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Paper • 2605.05204 • Published May 6 • 29