WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning
Abstract
WorldToken fuses heterogeneous robot observations into per-timestep world tokens processed by a causal Transformer and diffusion action head, with scaling and temporal-context analyses on RoboCasa and RMBench.
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token. A causal temporal Transformer models the resulting world-token sequence, and a diffusion action head generates action chunks. On 23 RoboCasa tasks, an 85.3M-parameter policy trained from scratch apart from a frozen pretrained CLIP text encoder achieves 59.45% mean closed-loop success using 2,900 generated demonstrations per task. A complete factorial sweep over five dataset sizes, five model sizes, and two training seeds shows consistent gains from additional target-domain data and diminishing returns beyond moderate model size. Under same-checkpoint history truncation, reducing visible history to one or two policy timesteps lowers closed-loop success for all 50 RoboCasa policies. On RMBench Blocks Ranking, reducing visible history from 146 to 8 seconds lowers evaluator success from 95% to 28%, while an exploratory extended rollout sustains the reference swap sequence for over 850 seconds. These results establish the empirical feasibility of the complete WorldToken instantiation and characterize its data-scaling and temporal-context behavior under the tested recipes. They do not establish superiority over alternative sequence organizations or isolate which components of the complete implementation drive the observed performance.
Community
Can a robot read the physical world as a language model reads text? Inspired by this question, we introduce WorldToken, a time-first approach to robotic sequence modeling in which policy timesteps define the top-level temporal sequence. We study its scaling and temporal-context behavior across RoboCasa and RMBench, including controlled history truncation and extended rollouts that sustain ordered behavior for over 850 seconds.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- World Tokens: Enhancing Embodied Policies with Training-Time World Modeling (2026)
- SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation (2026)
- Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation (2026)
- TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning (2026)
- Teaching Tiny VLA Models Where to Look and How to Move (2026)
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026)
- Bridge-WA: Predicting Where and How the World Changes for Robotic Action (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.22591 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper