Masked Visual Actions for Unified World Modeling

2026-07-21Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionRobotics
AI summary

The authors created a video-based system that helps robots understand and predict how their actions cause things to move in a scene. They use a technique called Masked Visual Actions, which shows only part of the robot's or object's movement in the video, letting the model imagine forward actions or figure out how to move to get a desired result. They trained the model on a small amount of real and simulated video data, and it can accurately predict outcomes in different scenes and for various robot types. This helps robots plan and decide actions better by imagining possible futures and how objects will respond.

video modelsrobot manipulationmasked visual actionsforward dynamicsinverse modelingmodel-based planningvisual controltrajectory predictionrobot embodimentssimulation
Authors
Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang
Abstract
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.