A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

2026-07-01Sound

SoundMachine Learning
AI summary

The authors studied two different parts of a speech generation system to see how well they can control emotions in synthetic speech. They found that one part, called the speech language model (SLM), has clear and separate emotion patterns that work well across different speakers. The other part, conditional flow-matching (CFM), mixes up speaker and emotion information, making it less flexible. Combining both parts made emotions stronger but hurt speech quality and control. Their work helps improve how we adjust emotions in computer-generated speech by understanding the underlying data patterns.

text-to-speechemotion controlspeech language modelconditional flow-matchingactivation steeringintrinsic dimensionalitylinear probingspeaker-emotion disentanglementcontrollable speech generation
Authors
Siyi Wang, James Bailey, Ting Dang
Abstract
While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparative study of speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotion-specific subspace with strong speaker--emotion disentanglement, while CFM exhibitspoor cross-speaker generalization due to speaker--emotion entanglement. Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.