[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85881-en":3,"doc-seo-85881-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85881,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Empowering Long-form Omni-modal Understanding with Robust Audio Perception","Recent advances in large-scale multimodal models have improved vision-language performance, yet robust omni-modal understanding remains limited by datasets that lack rich, explicitly aligned auditory cues. This work introduces AVDC (Audio-Visual Decoupled Captions), an automated pipeline producing tripartite captions: visual-only, audio-only, and joint audio-visual to disentangle semantics and model cross-modal interactions. It further proposes AVDC-QA-CoT for chain-of-thought audio-visual reasoning, supported by a two-stage training strategy. Extensive experiments show consistent gains across video captioning, audio-centric analysis, and omni-modal benchmarks.","[ cs .LG] 1 1 Jul 2026  \nEmpowering Long-form Omni-modal Understanding with Robust Audio Perception  \nKaiying Yan 1 , Luoyi Sun2 ,3 , Xiao Zhou 1 , and Weidi Xie 1  \n1 SAI, Shanghai Jiao Tong University, shanghai, China  \n2 Zhejiang University, Hangzhou, Zhejiang, china,  \n3 Shanghai AI Lab, shanghai, China  \nAbstract. Recent advances in large-scale multimodal models have driven remarkable progress in vision-language tasks; however, comprehensive omni-modal understanding remains under-explored, largely due to the scarcity of datasets with rich, explicitly aligned auditory cues. To bridge this gap, we present AVDC (Audio-Visual Decoupled Captions), a large-scale dataset designed to disentangle visual and auditory semantics. Specifically, we propose an automated pipeline that leverages off-the-shelf models to annotate videos with tripartite captions: visual-only (V), audioonly (A), and joint audio-visual (AV) . This decoupled structure explicitly captures both modality-specific nuances and complex cross-modal interactions. Building upon this, we introduce AVDC-QA-CoT, a Chainof-Thought augmented question-answering dataset to foster audio-visual reasoning. To fully exploit these resources, we employ a two-stage training paradigm: omni-modal caption generation pre-training on AVDC, followed by instruction tuning on AVDC-QA-CoT. Extensive experiments across diverse downstream tasks, spanning video captioning, audio-centric analysis, and omni-modal benchmarks, demonstrate consistent and significant performance gains, showing the efficacy of our proposed datasets and training strategy in advancing omni-modal perception. Code and  \nto comof view.  \n2  \n|  |  | (b)\u003Cbr>Question: Which event occurs LAST in the video? Choices:\u003Cbr>A.Crinkling sounds from plastic wrappers\u003Cbr>B.A muted thud as bagged toys are set down\u003Cbr>C.The woman mentions the missing checklist\u003Cbr>D.A hollow clatter from placing the Trash Pack container\u003Cbr>Answer: D | Reasoning:\u003Cbr>Question decomposition: order of events Temporal grounding: Crinkling at 0:02… Visual perception: Trash Pack container… Audio perception: Hollow clatter at 0:14…\u003Cbr>Multimodal reasoning: The clatter (D) aligns with the final visual placement …. |\n| --- | --- | --- | --- |\n|  |  |  |  |\n| (c) \u003Cbr>Q: What changes occurred in the state of the ring-shaped drones in the video?\u003Cbr>A. transitioned from a state of motion to rest and then back to motion. B. changed from a state of rest to motion and then remained in motion.\u003Cbr>C. transitioned from a state of motion to rest and then remained at rest. D. changed from a state of rest to motion and then returned to rest.\u003Cbr>Caption: First, two men are in a room, discussing a drone…The scene then shifts outdoors, where a man in a red hoodie operates a drone with a remote control, performing various flight maneuvers... [MISSING DETAILS]\u003Cbr>Answer: B.\u003Cbr>Caption: Two men are in a workshop, discussing and examining a drone...The video transitions to an outdoor scene where the device is flying,   [MISSING DETAILS] . Finally, the men return to discussing with satisfaction. [HALLUCINATION] .\u003Cbr>Answer: \u003Cthink>...Initially, the drone is hovering...The drone continues to fly around... the man appears to have landed the drone and made it fly...shows transition from motion to rest and then back to motion   \u003C/think>\u003Canswer>B. \u003C/answer>\u003Cbr>Caption: Two men examine a large drone in a workshop with lively background music. Outdoors, a man in red flies the drone performing stunts to cheerful music and occasional laughter until it crashes into a building, and the men discuss and laugh while walking toward it.\u003Cbr>Answer: \u003Cthink> [QUESTION DECOMPOSITION] ... initial state? transition to?…[TEMPORAL GROUNDING]...began at 1:06 and ended at the end. [VISUAL]The drones are seen hovering, flying, landing …[AUDIO] Mechanical sounds align with hovering and flight phases,…. [MULTIMODAL REASONING]...confirms that the drones changed... \u003C/think> \u003Canswer>D. \u003C/answer>\u003Cbr>Qwen2 .5-omni\u003Cbr>\u003Cb","cbCaip9xrnXEU65E","https://ap.wps.com/l/cbCaip9xrnXEU65E","pdf",5308853,4,1,31,"English","en",105,"# Abstract\n# Dataset Construction and Modal Decoupling\n## AVDC: Audio-Visual Decoupled Captions\n## AVDC-QA-CoT: Chain-of-Thought QA","[{\"question\":\"Why is omni-modal understanding still under-explored despite progress in multimodal models?\",\"answer\":\"Because comprehensive datasets with rich, explicitly aligned auditory cues are scarce, which limits learning robust cross-modal reasoning beyond vision-driven cues.\"},{\"question\":\"What is AVDC and how does it address deficiencies in existing datasets?\",\"answer\":\"AVDC is a large-scale audio-visual decoupled caption dataset. It generates aligned captions for audio-only, visual-only, and joint audio-visual views, capturing non-speech events and complex cross-modal interactions.\"},{\"question\":\"How is AVDC-QA-CoT used to improve audio-visual reasoning in models?\",\"answer\":\"AVDC-QA-CoT extends the resources with question-answer pairs augmented by chain-of-thought rationales, enabling instruction tuning that supports reasoning about both auditory and visual evidence.\"}]",1784206915,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"empowering-long-form-omni-modal-understanding-with-robust-audio-perception","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/empowering-long-form-omni-modal-understanding-with-robust-audio-perception/85881/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is omni-modal understanding still under-explored despite progress in multimodal models?","Question",{"text":75,"@type":76},"Because comprehensive datasets with rich, explicitly aligned auditory cues are scarce, which limits learning robust cross-modal reasoning beyond vision-driven cues.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is AVDC and how does it address deficiencies in existing datasets?",{"text":80,"@type":76},"AVDC is a large-scale audio-visual decoupled caption dataset. It generates aligned captions for audio-only, visual-only, and joint audio-visual views, capturing non-speech events and complex cross-modal interactions.",{"name":82,"@type":73,"acceptedAnswer":83},"How is AVDC-QA-CoT used to improve audio-visual reasoning in models?",{"text":84,"@type":76},"AVDC-QA-CoT extends the resources with question-answer pairs augmented by chain-of-thought rationales, enabling instruction tuning that supports reasoning about both auditory and visual evidence.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]