BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

2026-08-31Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors present BLARM, a method that makes 3D animations of objects using just a regular video and a static 3D model. Instead of using complex setups like skeletons or rigs, their method breaks down motion into simple parts that move rigidly and blends them to animate the object smoothly. They use a special attention-based neural network to link video features with the object's shape, creating animations that follow the video's motion closely and stay stable over time. Their training approach helps the system learn clear and compact motion patterns from single-camera videos.

3D mesh animationmonocular videorigid motion componentsskinning weightsspatial-temporal attentiontrajectory reconstructionentropy regularizationcontrastive learningmotion representationlow-dimensional deformation
Authors
Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Haomiao Jiang, Snehasish Mukherjee, Kyle Spence, Mark Stauber, Evangelos Kalogerakis, Yunze Zeng
Abstract
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.