Hydra-0: Action Flow for Generalist World Modeling and Control
2026-08-18 • Robotics
Robotics
AI summaryⓘ
The authors present Hydra-0, a system that helps robots understand and plan actions by looking at how pixels move in videos, which represents robot movements. This approach works across different robots, tasks, and environments, making it more flexible than older methods. Their system reduces errors in predicting robot and object movements and can even learn from seeing human demonstrations without extra training. Overall, the authors show that using pixel motion as a way to share information helps robots learn and perform better in various situations.
world modelaction flowrobot controlpixel motionzero-shot learningvideo generationrobot embodimentsinverse modellatent featurespolicy evaluation
Authors
Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang
Abstract
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.