$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
2026-08-06 • Robotics
Robotics
AI summaryⓘ
The authors created a new robot model called ω-0 that helps humanoid robots do tasks requiring moving and using their hands at the same time, like picking something up while walking. Instead of focusing only on the robot's arms or making video predictions, their model predicts whole-body actions directly from current camera views and robot states. They trained it using a large real-world dataset, ω-HOME, with detailed robot motions and sensor data. Tests showed that ω-0 works better than other methods for smoothly doing everyday tasks involving both movement and object handling.
Humanoid robotLoco-manipulationLatent predictive modelWhole-body actionController-compatible latentsEgocentric RGBExocentric depthSMPL motion modelDiffusion-based generationImitation learning
Authors
Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang, Xichen Yuan, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yang, Shanghang Zhang
Abstract
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $ω$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $ω$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $ω$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $ω$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.