[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84676-en":3,"doc-seo-84676-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84676,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","C3 ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection","Active Speaker Detection identifies whether a visible person is speaking at each video moment. Audio–visual fusion approaches work well on clean data but lose accuracy under real-world corruptions such as background noise, occlusion, and simultaneous degradation of both modalities. The weakness stems from missing explicit consistency constraints that guide robust, semantically aligned cross-modal representations. C3 ASD introduces three complementary constraints: embedding-level inter-modality consistency, sequence-level intra-modality consistency via track-aware contrastive learning, and prediction-level consistency via knowledge distillation.","arXiv :2607 .030 18v2 [ cs .CV] 7 Jul 2026  \nC 3 ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection  \nJin Hong 1 *, Jisoo Park 1 *, and Junseok Kwon 1  \nChung-Ang University, Seoul, Republic of Korea {jindl465,susiehome,[jskwon}@cau.ac.kr](jskwon}@cau.ac.kr)  \n*  \nEqual contribution.  \nFig. 1. Multi-Level Consistency for Robust ASD. In real-world active speaker detection, audio and visual streams are often corrupted by noise and occlusion. Existing models learn modality-clustered representations fragile under such degradations. In contrast, our multi-level consistency framework produces class-aligned, modality-invariant embeddings robust across diverse corruption scenarios.  \nAbstract. Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio–visual fusion methods perform well on clean data, they degrade under realworld corruptions such as background noise, occlusion, or simultaneous modality degradation. We attribute this limitation to the absence of explicit consistency constraints that promote robust, semantically aligned representations across modalities. Without such guidance, models tend to learn fragile modality-specific shortcuts that fail under corrupted conditions. We propose C3 ASD, a multi-level consistency-driven framework with three complementary constraints: embedding-level inter-modality consistency aligns audio-visual representations during speech; sequencelevel intra-modality consistency separates speaking and non-speaking clusters via track-aware contrastive learning; and prediction-level consistency stabilizes fusion through knowledge distillation. Extensive experiments demonstrate significant improvements under diverse audio, visual and joint corruptions, while maintaining competitive performance on clean data.  \nKeywords: Active Speaker Detection · Audio-Visual Consistency · Corruption Robustness  \n2 J. Hong et al.  \n1 Introduction  \nActive Speaker Detection (ASD) determines whether a visible person in a video is speaking at each moment, serving as a fundamental task in video understanding with applications in video conferencing [11], human–robot interaction [20, 35, 37, 41], and multimedia retrieval [7] . Owing to its inherently multimodal nature, ASD relies on both auditory speech signals and visual facial movements, motivating the development of audio–visual fusion methods that exploit complementary information from the two modalities. Recent advances in audio–visual ASD have achieved strong performance using sophisticated fusion architectures, such as temporal convolutional networks, transformers, and attention mechanisms [25, 32, 38, 44, 48] . These methods typically extract modalityspecific features and aggregate them through various fusion strategies. Despite their success on clean data, they often treat audio and visual streams as separate sources to be combined, without explicitly modeling how these modalities should relate to and constrain each other. This oversight becomes particularly problematic in real-world scenarios where one or both modalities may be corrupted by noise, occlusion, or other distortions.  \nIn real-world settings, modality corruption is inevitable. Audio may be contaminated by background noise, music, or overlapping speech [36, 40], while visual inputs can degrade due to motion blur, occlusion, low resolution, or rapid head movement. More challenging are scenarios where both modalities are simultaneously unreliable, such as unconstrained videos in the wild [31] . Under such conditions, existing ASD models often experience significant performance degradation [18,42], indicating that current approaches fail to robustly and complementarily leverage audio and visual information.  \nWe attribute this limitation to how current models learn audio-visual representations. As shown in Fig. 1, existing models learn modality-clustered representations that are fragile under real-world","cbCaisX3qMdDWI5H","https://ap.wps.com/l/cbCaisX3qMdDWI5H","pdf",17231404,2,1,25,"English","en",105,"# Introduction\n## Problem: Robustness Under Multi-Modal Corruption\n## Key Idea: Multi-Level Consistency Constraints\n# Method Overview\n## Inter-Modality Consistency\n## Intra-Modality Consistency\n## Prediction-Level Consistency","[{\"question\":\"What problem does C3 ASD address in active speaker detection?\",\"answer\":\"C3 ASD targets the performance drop of audio–visual active speaker detection under real-world corruption like noise, occlusion, and simultaneous modality degradation.\"},{\"question\":\"Why do existing models degrade when modalities are corrupted?\",\"answer\":\"The degradation is attributed to the lack of explicit consistency constraints, which allows models to learn fragile modality-specific shortcuts instead of robust, semantically aligned representations.\"},{\"question\":\"What are the three consistency constraints proposed by C3 ASD?\",\"answer\":\"C3 ASD uses (1) embedding-level inter-modality consistency during actual speech, (2) sequence-level intra-modality consistency via track-aware contrastive learning, and (3) prediction-level consistency through knowledge distillation to stabilize fusion.\"}]",1784197617,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"c3-asd-multi-level-consistency-driven-representation-learning-for-robust-active-speaker-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/c3-asd-multi-level-consistency-driven-representation-learning-for-robust-active-speaker-detection/84676/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does C3 ASD address in active speaker detection?","Question",{"text":75,"@type":76},"C3 ASD targets the performance drop of audio–visual active speaker detection under real-world corruption like noise, occlusion, and simultaneous modality degradation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do existing models degrade when modalities are corrupted?",{"text":80,"@type":76},"The degradation is attributed to the lack of explicit consistency constraints, which allows models to learn fragile modality-specific shortcuts instead of robust, semantically aligned representations.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the three consistency constraints proposed by C3 ASD?",{"text":84,"@type":76},"C3 ASD uses (1) embedding-level inter-modality consistency during actual speech, (2) sequence-level intra-modality consistency via track-aware contrastive learning, and (3) prediction-level consistency through knowledge distillation to stabilize fusion.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]