[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84873-en":3,"doc-seo-84873-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84873,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking","Multimodal Video-to-Music Recommendation (VTMR) introduces a two-stage framework that improves how relevant music is matched to a video. Stage 1 aligns multimodal video, music, and LLM-generated text in a joint audio-visual-text embedding space, retrieving semantically compatible candidates with coarse global embeddings. Stage 2 performs temporal reranking using cross-modal attention over video and music sequences, capturing fine-grained timing correspondence. Experiments show improved R@10 and Median Rank over strong baselines, and a human study confirms comparable preference to a commercial system while producing better music quality than a generative baseline.","Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking  \nSeungheon Doh 1 * Minhee Lee 1 Sangmoon Lee 2 Ben Sangbae Chon 2 Juhan Nam 1  \n[https://seungheondoh.github.io/video-to-music-demo](https://seungheondoh.github.io/video-to-music-demo)  \narXiv :2607 .0597 1v 1 [ cs .MM] 7 Jul 2026  \nAbstract  \nWe present VTMR, a two-stage framework for Video-To-Music Recommendation. In Stage 1, VTMR aligns comprehensive video and music signals in a joint audio-visual-text representation space and efficiently retrieves semantically compatible candidates using coarse global embeddings. In Stage 2, it reranks the retrieved candidates by attending to the temporal sequences of both video and music, thereby capturing finegrained temporal correspondence. Evaluated on the video-to-music recommendation task, the multimodal retrieval stage improves R@10 from 14.2 to 15.9 and Median Rank from 75 to 58 over the strongest baseline; the temporal reranker further boosts R@10 to 18.3 and Median Rank to 46, demonstrating complementary gains from richer query encoding and temporal alignment. A human preference study confirms that VTMR is on par with a commercial baseline in overall preference, while outperforming a generative baseline in music quality.  \n1. Introduction  \nMusic is a cornerstone of compelling video content, shaping emotional tone, narrative pacing, and audience engagement across domains from cinematic productions to short-form social media. Yet selecting appropriate background music is far from trivial: it requires not only an understanding of high-level semantic compatibility, such as mood and genre, but also a fine-grained understanding of the temporal correspondence between evolving visual dynamics and musical elements. Identifying music that satisfies both criteria is time-consuming and often yields suboptimal results.  \n1 Graduate School of Culture Technology, KAIST, South Korea 2 Gaudio Lab, Inc, South Korea. ∗Work completed while Seungheon was visiting Gaudio Lab. Correspondence to: Seungheon Doh \u003C[seungheon.doh@gmail.com](seungheon.doh@gmail.com) >.  \nInternational Conference on Machine Learning (ICML) 2026, Learning to Listen: Workshop on Machine Learning for Audio, Copyright 2026 by the author(s) .  \nAutomatic video-to-music (V2M) recommendation has therefore attracted growing interest as a cross-modal music retrieval task (Li & Kumar, 2019 ; Huang et al., 2022 ; Doh et al., 2023b ; Wu et al., 2025), aiming to learn joint embedding spaces between visual content and music (Sur´ıset al., 2022 ; Prtet et al., 2023 ; Wilkins et al., 2023 ; Stewart et al., 2025) . Despite steady progress, current frameworks leave substantial room for improvement across two major dimensions. First, while general-purpose multimodal models (Guzhov et al., 2022 ; Girdhar et al., 2023 ; Zhu et al., 2024) have demonstrated the power of joint audio-visualtext representations, established V2M methods still rely almost exclusively on raw visual features. Even though some previous works (McKee et al., 2023) have utilized both visual and textual modalities for V2M recommendation, existing architectures have yet to fully exploit the multimodal richness of both video and music signals.  \nSecond, all existing V2M methods reduce retrieval to a single global embedding similarity score, compressing entire video and music streams into static vectors. While global embeddings are effective for capturing coarse semantic compatibility, such as overall mood or genre, this single-vector bottleneck fundamentally cannot represent how localized musical events align with specific video moments. The temporal dimension is collapsed by design, making it impossible to distinguish music that is globally compatible from music that is temporally aligned with specific video moments.  \nTo address both limitations, we propose VTMR, a two-stage framework capable of capturing both unified multimodal semantics and temporal dynamics (see Figure 1) . Stage 1 (Sem","cbCaioVPMFeFsC9R","https://ap.wps.com/l/cbCaioVPMFeFsC9R","pdf",870651,2,1,6,"English","en",105,"# Abstract\n# Introduction\n# Methods\n## Multimodal Media Encoding","[{\"question\":\"What is VTMR, and how does it structure video-to-music recommendation?\",\"answer\":\"VTMR uses a two-stage design: semantic retrieval followed by temporal reranking. Stage 1 retrieves candidates using aligned multimodal embeddings, and Stage 2 reranks them by attending to temporal sequences across video and music.\"},{\"question\":\"How does Stage 1 perform retrieval in VTMR?\",\"answer\":\"Stage 1 projects comprehensive video, music, and LLM-generated text into a shared audio-visual-text representation space. It then retrieves top-N candidates efficiently via similarity computed from coarse global embeddings.\"},{\"question\":\"Why does VTMR include temporal reranking, and what does Stage 2 do?\",\"answer\":\"VTMR addresses the limitation of single global embeddings that collapse time and cannot model localized alignment. Stage 2 applies a fine-grained cross-encoder that attends to dense, unpooled temporal sequences of the video and each candidate music track to capture temporal correspondence.\"}]",1784198942,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multimodal-video-to-music-recommendation-via-semantic-retrieval-and-temporal-reranking","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/multimodal-video-to-music-recommendation-via-semantic-retrieval-and-temporal-reranking/84873/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is VTMR, and how does it structure video-to-music recommendation?","Question",{"text":75,"@type":76},"VTMR uses a two-stage design: semantic retrieval followed by temporal reranking. Stage 1 retrieves candidates using aligned multimodal embeddings, and Stage 2 reranks them by attending to temporal sequences across video and music.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Stage 1 perform retrieval in VTMR?",{"text":80,"@type":76},"Stage 1 projects comprehensive video, music, and LLM-generated text into a shared audio-visual-text representation space. It then retrieves top-N candidates efficiently via similarity computed from coarse global embeddings.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does VTMR include temporal reranking, and what does Stage 2 do?",{"text":84,"@type":76},"VTMR addresses the limitation of single global embeddings that collapse time and cannot model localized alignment. Stage 2 applies a fine-grained cross-encoder that attends to dense, unpooled temporal sequences of the video and each candidate music track to capture temporal correspondence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]