PhysV2A: Reachability-Gated and Semantic-Mask-Constrained Feasibility Completion for Video-to-Robot Manipulation
2026-07-10 • Robotics
Robotics
AI summaryⓘ
The authors propose PhysV2A, a system that turns object motions seen in videos into robot movements that a specific robot can actually perform. Unlike usual methods that ignore the robot's body, PhysV2A checks which robot grasps and motions are physically possible and picks the best ones. It also uses a special mask to fine-tune how the robot moves, making sure important parts of the task are done correctly while allowing small, safe adjustments. Tests on real robots showed that PhysV2A works better than other methods at completing tasks and avoids movements that robots can't do.
video-based manipulation6D object motionrobot grasp feasibilityreachability-gated selectionRGB-D perceptiontrajectory optimizationsemantic masksredundancy resolutionmanipulabilityrobot kinematics
Authors
Haohui Huang, Junda Duan, Tao Teng, Chenguang Yang
Abstract
Video-based manipulation provides object-centric motion priors from human demonstrations, generated videos, or RGB-D observations, but such priors are typically embodiment-agnostic and cannot be directly executed by a specific robot. This paper presents \textbf{PhysV2A}, a reachability-gated and semantic-mask-constrained feasibility-completion framework for converting video-derived 6D object motion into robot-executable manipulation trajectories. The key idea is to treat grasp feasibility as trajectory-conditioned rather than local: each RGB-D-generated 6-DoF grasp candidate is rigidly coupled with the recovered object motion to form a grasp-conditioned TCP trajectory hypothesis. PhysV2A then performs hierarchical reachability-gated selection, where infeasible grasp--trajectory pairs are rejected by robot-centric kinematic checks and surviving candidates are ranked by downstream execution suitability. For the selected reachable trajectory, a VLM-assisted and rule-validated S-Mask identifies task-critical and relaxable Cartesian components, enabling semantic-mask-constrained manipulability refinement through redundancy-first optimization and bounded Cartesian relaxation. Real-robot experiments on four tabletop manipulation tasks show that PhysV2A improves task success over representative video-prior and IK-only baselines, reduces kinematic-feasibility failures, and produces better-conditioned trajectories with bounded semantic deviations.