[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83099-en":3,"doc-seo-83099-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83099,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","WordVoice Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS","LLM-based text-to-speech (TTS) models deliver strong naturalness but typically use implicit end-to-end generation, which yields coarse-grained control. Precise stylistic interventions and strict temporal alignment—critical for audiobook narration and video dubbing—are limited by the lack of explicit word-level manipulation of acoustic attributes. The bottleneck is intensified by scarce fine-grained annotated data and the difficulty of integrating multi-dimensional control signals into discrete autoregressive generation. WordVoice addresses this with a unified framework using WordVoice-5A and explicit acoustic planning.","WORDVOICE: EXPLICIT AND DECOUPLED MULTI-DIMENSIONAL WORD-LEVEL CONTROL FOR LLM-BASED TTS  \nSihang Nie 1 ,2∗, Jinxin Ji3 ,4, Xiaofen Xing 1†, Deyi Tuo2, Chengbin Jin2, Jialong Mai 1, Xiangmin Xu 1 ,5  \n1 South China University of Technology, 2Huya Inc., 3Tongji University  \n4The Hongkong polytechnic university, 5Foshan University  \n[xfxing@scut.edu.cn](xfxing@scut.edu.cn), [bcshnie@mail.scut.edu.cn](bcshnie@mail.scut.edu.cn)  \narXiv :2607 .06461v1 [ ee ss .AS] 7 Jul 2026  \nABSTRACT  \nWhile recent Large Language Model (LLM)-based Textto-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate wordlevel acoustic attributes remains a critical bottleneck. This limitation is primarily amplified by the severe scarcity of finegrained annotated datasets and the architectural challenge of integrating multi-dimensional control signals into discrete autoregressive generation. To address this, we propose a unified framework for highly precise word-level control. First, we construct WordVoice-5A, a massive 4 .7k-hour bilingual dataset featuring five-dimensional word-level annotations (duration, boundary, energy, pitch and tone) developed through a rigorous linguistically-guided pipeline. Second, we introduce WordVoice to transform the implicit generation process into an explicit, highly controllable paradigm. Specifically, we introduce a bound-token mechanism within the LLM to formulate an explicit “acoustic planning” process, enabling adaptive multi-task prosodic planning and flexible manual intervention. Furthermore, we augment the token-to-waveform stage with a fine-grained acoustic modulation module, bridging the resolution gap to strictly align word-level attributes between highly compressed discrete tokens and continuous waveforms. Extensive experiments demonstrate that WordVoice achieves superior, decoupled control over multiple acoustic dimensions while maintaining competitive zero-shot synthesis stability. The code and audio samples are publicly available at [https:](https:)//[xxh333.github.io/wordvoice-demo/](xxh333.github.io/wordvoice-demo/) .  \nIndex Terms— Text-to-Speech, Large Language Model, Controllable Synthesis, Word-Level Control  \n∗ Work conducted when the author was intern at Huya Inc.†Corresponding author.  \nFig. 1: WordVoice framework. By introducing explicit wordlevel control, WordVoice supports a dual-mode synthesis paradigm. Users can either rely on the model’s autonomous prosodic planning or explicitly manipulate five-dimensional acoustic attributes for specific words to achieve highly expressive and precise stylistic interventions.  \n1. INTRODUCTION  \nRecent advancements in controllable Text-to-Speech (TTS) have enabled expressive style manipulation. Existing approaches primarily fall into two paradigms. The first relies on global embedding control [1, 2, 3], which achieves satisfactory results in specific domains but suffers from coarse control granularity and limited generalization. The second paradigm, instruction-based TTS [4, 5, 6], leverages natural language prompts to enable localized and more generalized control. However, despite their versatility, these models still fall fundamentally short of achieving high-precision, finegrained acoustic control at the character-or word-level.  \nThis granularity bottleneck creates a critical void in realworld applications that demand deterministic prosodic interventions. For instance, in audiobook narration, creators often need to meticulously manipulate the duration, energy (emphasis), or tone (intonation) of specific words to convey nuanced character emotions [7, 8] . Yet, relying solely on textual instructions makes it difficult for current models to reliably enforce such localized, high-pre","cbCaitA2wXpaQgAw","https://ap.wps.com/l/cbCaitA2wXpaQgAw","pdf",1756217,6,1,10,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What control limitation do current LLM-based TTS systems face at word or character level?\",\"answer\":\"They rely on implicit end-to-end generation, which prevents explicit manipulation of fine-grained word-level acoustic attributes, leading to coarse control granularity and weaker temporal precision.\"},{\"question\":\"How does WordVoice improve word-level control for LLM-based TTS?\",\"answer\":\"It introduces an explicit acoustic planning paradigm using a bound-token mechanism in the LLM, enabling adaptive multi-task prosodic planning and manual intervention over multiple acoustic dimensions.\"},{\"question\":\"What dataset is proposed and what annotations does it include?\",\"answer\":\"WordVoice constructs WordVoice-5A, a 4.7k-hour bilingual dataset with five-dimensional word-level annotations: duration, boundary, energy, pitch, and tone, produced via a linguistically guided pipeline.\"}]",1784185241,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"wordvoice-explicit-and-decoupled-multi-dimensional-word-level-control-for-llm-based-tts","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/wordvoice-explicit-and-decoupled-multi-dimensional-word-level-control-for-llm-based-tts/83099/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What control limitation do current LLM-based TTS systems face at word or character level?","Question",{"text":76,"@type":77},"They rely on implicit end-to-end generation, which prevents explicit manipulation of fine-grained word-level acoustic attributes, leading to coarse control granularity and weaker temporal precision.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does WordVoice improve word-level control for LLM-based TTS?",{"text":81,"@type":77},"It introduces an explicit acoustic planning paradigm using a bound-token mechanism in the LLM, enabling adaptive multi-task prosodic planning and manual intervention over multiple acoustic dimensions.",{"name":83,"@type":74,"acceptedAnswer":84},"What dataset is proposed and what annotations does it include?",{"text":85,"@type":77},"WordVoice constructs WordVoice-5A, a 4.7k-hour bilingual dataset with five-dimensional word-level annotations: duration, boundary, energy, pitch, and tone, produced via a linguistically guided pipeline.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]