[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83095-en":3,"doc-seo-83095-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83095,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space","Video-to-audio (V2A) generation synthesizes audio that matches both the semantics and the timing of a silent video, yet many approaches remain computationally heavy due to multi-stage training or rely on video-to-text conversions that miss fine temporal cues. Flowley introduces end-to-end single-stage training with visual features and textual prompts, and a Progressive Softmasked Cross-Attention that embeds synchronization in attention at no extra cost. SoundCap addresses limited benchmark captions by producing detailed, sound-aware annotations. Flowley sets new results on VGGSound and improves zero-shot audio quality.","arXiv :2607 .06405v 1 [ cs .MM] 7 Jul 2026  \nPrecise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space  \nThanh V. T. Tran 1 , Ngoc-Son Nguyen 1 , Luong Tran 1 , Long-Khanh Pham 1 ,  \nPaarth Neekhara2 , Shehzeen Hussain2 , and Van Nguyen 1  \n1 FPT Software AI Center, Vietnam  \n2 NVIDIA Corporation, USA  \n[https://flowley-v2a.github.io](https://flowley-v2a.github.io)  \n{thanhtvt1,sonnn45,luongtk,khanhpl2,[vannth19}@fpt.com](vannth19}@fpt.com)[ ](vannth19}@fpt.com){pneekhara,[shehzeenh}@nvidia.com](shehzeenh}@nvidia.com)  \nAbstract. Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress, many methods still rely on multi-stage training, resulting in high computational costs and long runtimes, or transform visual input into text to leverage pretrained text-to-audio models, sacrificing fine-grained temporal cues. To overcome these limitations, we propose Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts. Crucially, we introduce Progressive Softmasked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers. We further observe that existing V2A benchmarks lack sound-oriented descriptive captions, which can potentially degrade the quality of the synthesized audio. To remedy this, we propose SoundCap, a plug-and-play pipeline for creating detailed, sound-aware captions that guide the model. Remarkably, without integrating any pretrained audio-visual alignment modules, Flowley achieves state-of-the-art performance on VGGSound across multiple metrics. Moreover, by incorporating SoundCap, we further exceed the performance of the strongest existing close-sourced methods in terms of audio quality in the zero-shot setting.  \nKeywords: Video-to-Audio · Cross-attention · Video captioning  \n1 Introduction  \nSound design is the craft of storytelling through sonic composition. A key branch of this field is Foley [46], where sound effects are created in precise synchronization with onscreen action during post-production. These soundscapes are brought to life by a skilled Foley artist working on a purpose-built stage stocked with a wide array of props and materials for generating the required sounds3.  \n3 See how a foley artist performing a scene in our project page.  \n2 Tran et al.  \nRecent advances in generative audio modeling have facilitated text-to-audio (T2A) systems [11, 28, 32], allowing designers to generate sound effects directly from written descriptions. Despite accelerating the search for appropriate sounds, they must still manually adjust timing to align audio with on-screen action. This contrasts sharply with traditional Foley work, where artists naturally craft and sync sound effects in real time by interacting physically with props.  \nTo overcome this limitation, video-to-audio (V2A) generation has gained significant attention, dividing into two major directions. The first converts visual features from silent video frames into text, leveraging pretrained T2A models to synthesize sound [51, 57] . While this approach leverages the strengths of pretrained T2A systems, it inevitably discards fine-grained temporal details, which are crucial for film-production sound effects. The second line of work uses multistage training: early phases learn to extract or align acoustic cues from video frames, either via dedicated regression networks [18, 47, 52] or contrastive objectives [31], and subsequent stages build on these pretrained modules. While effective, these pipelines incur substantial delays and computational overhead. Furthermore, integrating textual descriptions into the generation workflow, beyond their use in T2A systems, remains under-explored, despite the demonstrated benefits of text ","cbCaikFMnTY58jck","https://ap.wps.com/l/cbCaikFMnTY58jck","pdf",41166581,2,1,32,"English","en",105,"# Introduction\n## Video-to-Audio generation\n## Sound design background\n## Proposed Flowley approach\n# Related Works\n## Video-to-Audio generation","[{\"question\":\"What problem does Flowley address in video-to-audio generation?\",\"answer\":\"Flowley targets two key issues: high cost from multi-stage training and loss of fine-grained temporal cues when visual inputs are converted into text for pretrained text-to-audio models.\"},{\"question\":\"How does Progressive Softmasked Cross-Attention (PSCA) help synchronization?\",\"answer\":\"PSCA incorporates audio-visual synchronization directly into the attention mechanism, aligning acoustic and visual information without adding extra computational cost compared to standard attention layers.\"},{\"question\":\"Why is SoundCap introduced, and how does it work?\",\"answer\":\"Existing V2A benchmarks lack sound-oriented descriptive captions, which can harm audio semantic consistency. SoundCap is a plug-and-play captioning pipeline that generates detailed, sound-aware captions using pretrained audio-visual large language models to serve as training guidance.\"}]",1784185201,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"precise-video-to-audio-generation-with-cross-modal-alignment-in-latent-space","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/precise-video-to-audio-generation-with-cross-modal-alignment-in-latent-space/83095/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Flowley address in video-to-audio generation?","Question",{"text":75,"@type":76},"Flowley targets two key issues: high cost from multi-stage training and loss of fine-grained temporal cues when visual inputs are converted into text for pretrained text-to-audio models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Progressive Softmasked Cross-Attention (PSCA) help synchronization?",{"text":80,"@type":76},"PSCA incorporates audio-visual synchronization directly into the attention mechanism, aligning acoustic and visual information without adding extra computational cost compared to standard attention layers.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is SoundCap introduced, and how does it work?",{"text":84,"@type":76},"Existing V2A benchmarks lack sound-oriented descriptive captions, which can harm audio semantic consistency. SoundCap is a plug-and-play captioning pipeline that generates detailed, sound-aware captions using pretrained audio-visual large language models to serve as training guidance.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]