Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors

2026-07-11Graphics

GraphicsSound
AI summary

The authors developed Learn2Chat, a method to create realistic motion for two people talking by building on existing models that generate motion for one person speaking. Their approach separates the natural movements driven by speech from the social reactions between conversation partners, making it easier to model these interactions. They use a special technique to factor out and predict interactive behaviors over pretrained motion data, allowing more accurate and natural conversational motion. Their experiments show improved performance compared to previous methods, and their approach can work with various existing motion models.

dyadic conversational motionmonologic motion modelsspeech-driven motionsocial interaction modulationmotion factorizationcross-attentionlatent spaceDualTalk benchmarkpretrained modelsinteractive digital humans
Authors
Zikai Huang, Siyue Chen, Xuemiao Xu, Haoxin Yang, Cheng Xu, Yihong Lin, Shengfeng He
Abstract
Dyadic conversational motion generation is essential for realistic interactive digital humans. Existing approaches typically model conversational behaviors within unified dyadic generators. However, such holistic formulations tend to couple self-speech-driven motion with partner-responsive social feedback, leaving the interaction-specific component implicit and underutilizing the speech-motion correspondence already learned by pretrained monologic motion models. We propose Learn2Chat, a unified framework that models dyadic motion as interaction modulation over pretrained monologic motion priors. This design separates intrinsic speech-driven motion from social interaction effects and enables more structured interaction modeling. Specifically, we introduce a Monologic-Anchored Motion Factorization scheme that leverages the semantic motion manifold learned from monologic data to disentangle audio-driven motion dynamics from interaction-induced modulation, yielding clean interaction representations from dyadic sequences. On top of this representation space, a Cross-Attentive Interaction Latent Prediction module maps paired speech signals to interaction latents through cross-branch attention and interaction alignment. During inference, the predicted interaction latents modulate canonical monologic motion to generate coherent and synchronized dyadic behaviors in a data-efficient manner. Extensive experiments on the DualTalk benchmark demonstrate that Learn2Chat achieves state-of-the-art performance across both quantitative metrics and perceptual evaluations. Moreover, the framework is model-agnostic and seamlessly integrates with diverse pretrained monologic motion backbones, highlighting the effectiveness of prior reuse and interaction adaptation for scalable conversational motion generation. More visual results are available on the project page.