[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86217-en":3,"doc-seo-86217-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86217,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","BackgroundMellow：面向叙事驱动的电影级多模态连贯音景生成框架","BackgroundMellow proposes an end-to-end multimodal framework for narrative-driven cinematic soundscape generation from long-form text, addressing limitations of existing Text-to-Audio systems in cohesion, temporal alignment, and emotional depth. The master-specialist architecture decomposes stories into multilayer audio cues, generates each sound category with specialized models, and superimposes aligned soundscapes. Built on Tango2 for environmental synthesis and a Cinematic BGM Retriever, it uses an NLP module to predict mixing parameters (start time, duration, relative loudness) from the narrative timeline and evaluates via nearest-neighbor retrieval and semantic-temporal metrics.","BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic  \nSoundscape Generation  \nAjitesh Jamulkar and Aritra Hazra  \nDepartment of Computer Science & Engineering, Indian Institute of Technology Kharagpur, India.  \n13 Jul 2026  \nAbstract—Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with narrative cohesion, temporal alignment, and cinematic emotional depth. We present BackgroundMellow, a framework that treats story-to-audio generation as a precise orchestration and signal processing problem. This framework is enabled without ground-truth through a master-specialist agent architecture that decomposes text into precise and multilayered audio cues, generates each category of sounds with suitable specialist model, and superimposes the soundscapes to create a unified and aligned audio segment. Our pipeline is built over Tango2 latent diffusion model for environmental synthesis alongside a novel Cinematic BGM Retriever mined from professional soundtracks. To automate the sound mixing process, we use an NLP based module that predicts precise audio parameters, like start time, duration, and relative loudness, based on the narrative timeline. We further empirically evaluate and show the efficacy of the proposed framework leveraging nearest-neighbor retrieval against a curated dataset of YouTube  \nadvancements in Text-to-Audio (TTA) generation, Latent Diffusion Models (LDMs) and flow-matching  \narXiv :2607 . 11364v1  \nframeworks, have demonstrated remarkable success in synthesizing high-fidelity acoustic clips. However, a critical frontier remains largely unsolved: the generation of long-form, multitrack narrative soundscapes. Real-world cinematic storytelling is not a monolithic acoustic event; it is a complex, temporally orchestrated mixture of expressive speech, discrete sound effects (SFX), continuous environmental ambience, and contextual background music (BGM) .  \nCurrent state-of-the-art end-to-end models often excel at isolated sound generation but fundamentally struggle with compositional and acoustic complexity. When prompted to generate simultaneous, overlapping events, they frequently suffer from phase cancellation, acoustic muddiness, and temporal drift. For instance, successfully rendering the narrative prompt, “It was softly raining as I walked through the forest, where I heard a dog barking,” requires far more than basic semantic understanding. It demands the cohesive orchestration of a continuous rain ambience, an emotive background score, and a discrete dog bark SFX that triggers at the precise temporal onset of the corresponding spoken word. Maintaining  \nthe structural integrity of these diverse acoustic elements and superimposing them with millisecond precision remains a highly complicated and challenging task.  \nFurthermore, the absence of holistic and rigorous evaluation mechanisms for such multi-stem acoustic narratives presentsa fundamental research bottleneck. Cinematic sound design is inherently subjective; a generative model might synthesize asoundscape that utilizes a different musical genre or richer ambient textures than a human baseline, making it equally valid or even superior. Consequently, establishing a rigid, absolute “ground truth” dataset for multi-track audio is fundamentally flawed, as it penalizes generative diversity and creative variance. Existing objective metrics, such as Frchet Audio Distance (FAD) [16] or global Contrastive Language-Audio Pretraining (CLAP) [15] scores, are heavily optimized for evaluating isolated acoustic events. They are notoriously illequipped to assess the temporal synchronization, hierarchical audio ducking, and semantic harmony required for a multiclass cinematic mix.  \nTo bridge this gap, we propose an end-to-end, multimodal cohesive fram","cbCaitdo7WUhJSFA","https://ap.wps.com/l/cbCaitdo7WUhJSFA","pdf",786172,3,1,7,"English","en",105,"# Abstract\n# Contributions\n## Hierarchical Orchestration Framework\n## Deterministic Temporal Alignment\n## Retrieval-Based Evaluation Metrics\n## Empirical Evaluation and Benchmarking","[{\"question\":\"BackgroundMellow解决了文本到音频生成中的哪些关键问题？\",\"answer\":\"它针对长文本叙事音频生成中的叙事连贯性不足、时间对齐困难以及缺乏电影式情感深度等问题，提出以编排与信号处理为核心的框架来统一多轨音景。\"},{\"question\":\"框架如何实现多音轨的精确时间对齐与叠加？\",\"answer\":\"BackgroundMellow采用混合对齐引擎，用微调的数字声音预测器将SFX、环境氛围与音乐锚定到TTS主轨提取的词级时间戳，并通过参数化叠加生成对齐的音景片段。\"},{\"question\":\"文中提出了哪些评估思路来适配多轨叙事音频？\",\"answer\":\"针对现有评估指标对孤立声事件的优化不足，作者提出基于最近邻检索的评估框架，并引入语义召回与时间交并比（IoU）等指标，用于衡量节奏与同步，同时避免惩罚合理的生成多样性。\"}]",1784209530,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"backgroundmellow-a-multi-modal-cohesive-framework-for-narrative-driven-rich-cinematic-soundscape-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/backgroundmellow-a-multi-modal-cohesive-framework-for-narrative-driven-rich-cinematic-soundscape-generation/86217/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"BackgroundMellow解决了文本到音频生成中的哪些关键问题？","Question",{"text":75,"@type":76},"它针对长文本叙事音频生成中的叙事连贯性不足、时间对齐困难以及缺乏电影式情感深度等问题，提出以编排与信号处理为核心的框架来统一多轨音景。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"框架如何实现多音轨的精确时间对齐与叠加？",{"text":80,"@type":76},"BackgroundMellow采用混合对齐引擎，用微调的数字声音预测器将SFX、环境氛围与音乐锚定到TTS主轨提取的词级时间戳，并通过参数化叠加生成对齐的音景片段。",{"name":82,"@type":73,"acceptedAnswer":83},"文中提出了哪些评估思路来适配多轨叙事音频？",{"text":84,"@type":76},"针对现有评估指标对孤立声事件的优化不足，作者提出基于最近邻检索的评估框架，并引入语义召回与时间交并比（IoU）等指标，用于衡量节奏与同步，同时避免惩罚合理的生成多样性。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]