[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84604-en":3,"doc-seo-84604-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84604,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models","A comparative study examines how geometric properties of speech-language-model (SLM) and conditional-flow-matching (CFM) modules affect steerable mixed-emotion speech synthesis. Emotion representations are characterized using linear probing and local intrinsic dimensionality (LID), followed by single-site and joint activation steering evaluations on multiple datasets. Results show SLM provides a clean, low-dimensional emotion subspace with strong speaker–emotion disentanglement, whereas CFM generalizes poorly across speakers due to entanglement. Joint steering raises intensity but degrades proportional control and speech quality in-distribution, guiding multi-site steering design.","A Geometric Perspective on Composable Emotion Steering in Text-to-Speech  \nModels  \nSiyi Wang 1 James Bailey 2 Ting Dang 1  \narXiv :2607 .00946v 1 [ cs . SD] 1 Jul 2026  \nAbstract  \nWhile prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood.  \nWe present the first comparative study of speech language model (SLM) and conditional flowmatching (CFM) modules as activation steering sites for mixed-emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotionspecific subspace with strong speaker–emotion disentanglement, while CFM exhibitspoor crossspeaker generalization due to speaker–emotion entanglement.Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.  \n1. Introduction  \nGenerating emotionally controllable speech is essential for applications such as conversational agents, audiobook narration, and assistive communication. Human emotional expression is nuanced, often involving mixed affective cues where multiple emotions coexist within a single utterance (Zhou et al., 2022 ; Cowen & Keltner, 2017), a complexity that current systems generally fail to control effectively. Existing emotion control methods operate through the model’s external interface: label-based approaches can  \n1The University of Melbourne, Australia 2Monash University, Australia. Correspondence to: Siyi Wang \u003C[siyi.wang.4@student.unimelb.edu.au](siyi.wang.4@student.unimelb.edu.au) >.  \nPublished at ICML 2026 Workshop on the Machine Learning for Audio, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nenable explicit emotion conditioning but require costly annotated data and retraining (Cho et al., 2025 ; Gao et al., 2025), while prompt-based methods can describe target emotions but lack precise quantitative control over emotion proportions (Guo et al., 2023 ; Yang et al., 2025) .  \nActivation steering bypasses these limitations by directly injecting learned direction vectors into intermediate activations at inference time, without retraining (Zou et al., 2023 ; Turner et al., 2023) . This paradigm has shown success in LLMs and text-to-image diffusion (Rodriguez et al., 2025 ; Rimsky et al., 2024) . State-of-the-art TTS systems increasingly adopt hybrid architectures combining a speech language model (SLM) with a conditional flow-matching (CFM) decoder (Du et al., 2024 ; Anastassiou et al., 2024 ; Zhou et al., 2026), where the SLM governs high-level prosodic structure and the CFM renders fine-grained acoustics, each a potential site for steering emotional expression. Wang et al. (2026) demonstrate composable mixed-emotion steering via the SLM, while Xie et al. (2025) achieve continuous single-emotion intensity control via CFM. However, no prior work has systematically compared the representation geometry at these two steering sites, examined how geometric properties relate to steering effectiveness, or investigated whether jointly steering both modules yields complementary or interfering effects.  \nWe present the first comparative study of SLM and CFMas activation steering sites for mixed-emotion synthesis. Through linear probing and local intrinsic dimensionality (LID) analysis of both modules’ representation geometry, combined with single-site and joint steering experiments on four datasets, our study reveals three key findings:(i) the SLM encodes emotions in geometrically distinct, lowdimensional subspaces with strong cross-speaker generaliza","cbCaimNXfKrMLOZo","https://ap.wps.com/l/cbCaimNXfKrMLOZo","pdf",2393651,2,1,6,"English","en",105,"# Abstract\n# Introduction\n# Method\n## Geometry Analysis","[{\"question\":\"What is the main goal of the study on composable emotion steering?\",\"answer\":\"The study aims to compare SLM and CFM as activation steering sites for mixed-emotion speech synthesis, focusing on how representation geometry relates to steering effectiveness.\"},{\"question\":\"How are emotion representations analyzed in the paper?\",\"answer\":\"Emotion representations are characterized using linear probing and local intrinsic dimensionality (LID) to study how emotions are organized in the modules’ representation spaces.\"},{\"question\":\"What are the key findings when steering SLM versus CFM, and when steering jointly?\",\"answer\":\"SLM steering offers stronger proportional control with better speaker–emotion disentanglement, while CFM steering yields stronger overall intensity but weaker speaker fidelity and poorer cross-speaker generalization; joint steering increases intensity but degrades proportional control and speech quality on in-distribution data.\"}]",1784197054,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-geometric-perspective-on-composable-emotion-steering-in-text-to-speech-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/a-geometric-perspective-on-composable-emotion-steering-in-text-to-speech-models/84604/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the study on composable emotion steering?","Question",{"text":75,"@type":76},"The study aims to compare SLM and CFM as activation steering sites for mixed-emotion speech synthesis, focusing on how representation geometry relates to steering effectiveness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are emotion representations analyzed in the paper?",{"text":80,"@type":76},"Emotion representations are characterized using linear probing and local intrinsic dimensionality (LID) to study how emotions are organized in the modules’ representation spaces.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the key findings when steering SLM versus CFM, and when steering jointly?",{"text":84,"@type":76},"SLM steering offers stronger proportional control with better speaker–emotion disentanglement, while CFM steering yields stronger overall intensity but weaker speaker fidelity and poorer cross-speaker generalization; joint steering increases intensity but degrades proportional control and speech quality on in-distribution data.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]