Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

2026-07-10Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionSound
AI summary

The authors address the challenge of creating long and smooth dance videos from music, which current methods struggle with beyond 20 seconds. They introduce a two-step system that first plans key dance poses using the whole music track and then fine-tunes the motion to keep it natural and continuous. Their approach uses smart timing adjustments and a special way to compare motion between frames to improve video quality. Tests show their method can make clear, stable dance videos over one minute long in different styles, guided by both music and text. This work improves how computers generate dance videos that stay consistent and synced over longer times.

diffusion modelskeyframe planningtemporal refinementRoPE embeddingsoptical flowmotion-speed controlvideo synthesislong-range coherenceaudio-conditioned generationdance video generation
Authors
Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang
Abstract
Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.