[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83249-en":3,"doc-seo-83249-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83249,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment","Action Quality Assessment (AQA) evaluates how accurately and effectively a person performs movement, supporting sports scoring, skill evaluation, and healthcare needs. In real-world settings, unimodal systems often miss subtle quality cues. Existing multi-modal methods face two key issues: heterogeneous modalities create cross-modal misalignment and unstable fusion, and multi-modal labeling is expensive, limiting dataset diversity. DualAlign introduces two-stage adaptive alignment and proposes MM–JDM, a realistic multi-modal benchmark with structured text.","Noname manuscript No.  \n(will be inserted by the editor)  \nTwo-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment  \nKWaanngglei ZhoJiuangu⋅o RLuiizhi CXaiiaohi XLiinninangg Wang  ⋅ Yijian Zheng  ⋅ Liyuan  \nReceived: date / Accepted: date  \n8 Jul 2026  \nAbstract Action Quality Assessment (AQA) aims to evaluate how well a person performs a movement, which is essential in applications such as sports scoring, skill assessment, and healthcare. However, unimodal approaches often struggle to capture subtle cues of movement quality in realworld settings. Although multi-modal inputs provide complementary information, existing methods still face two major challenges: heterogeneous modalities often lead to crossmodal misalignment and unstable fusion, and reliable multi-  \narXiv :2607 .074 38v1 [ cs .CV]  \nmodal annotation is costly, resulting in limited dataset diversity. To address these challenges, we propose DualAlign, a two-stage multi-modal fusion framework with adaptive alignment. The framework first constructs a coherent visual representation by maximizing shared structural information across RGB video, optical flow, and skeleton modalities. Textual semantics are then incorporated after visual stabilization, allowing high-level descriptions to complement rather than distort the underlying visual manifold. To evaluate the framework under realistic multi-modal conditions, we introduce MM–JDM, a movement-quality assessment dataset integrating RGB videos, optical flow, skeleton sequences, and structured text. MM–JDM naturally exhibits  \nKanglei Zhou, Liyuan Wang  \nDept. of Psychological and Cognitive Sciences, Tsinghua University, Beijing, China  \nE-mail: {zhoukanglei, [liyuanwang](liyuanwang}@tsinghua.edu.cn)[}](liyuanwang}@tsinghua.edu.cn)[@tsinghua.edu.cn](liyuanwang}@tsinghua.edu.cn)[ ](liyuanwang}@tsinghua.edu.cn)Ruizhi Cai, Yijian Zheng, Xiaohui Liang (􀀀)  \nState Key Lab of Virtual Reality Technology and Systems, Beihang University, Beijing, China  \nE-mail: {craaaaazy, zhengyijian, liang [xiaohui](xiaohui}@buaa.edu.cn)[}](xiaohui}@buaa.edu.cn)[@buaa.edu.cn](xiaohui}@buaa.edu.cn)[ ](xiaohui}@buaa.edu.cn)Xinning Wang, Jianguo Li (􀀀)  \nDept. of Rheumatology and Immunology, Children’s Hospital, Capital Institute of Pediatrics, Beijing, China  \nE-mail: {xnwangcip, jianguo [li6](li6}@hotmail.com)[}](li6}@hotmail.com)[@hotmail.com](li6}@hotmail.com)[ ](li6}@hotmail.com)Xiaohui Liang  \nZhongguancun Laboratory, Beijing, China  \nmodality noise, class imbalance, and label scarcity, making it a challenging benchmark for studying multi-modal fusion and alignment. Extensive experiments show that DualAlign improves average correlation on MM–JDM by 21.16% over the state-of-the-art methods and achieves gains of 3.53% and 5.95% on the RG and Fis-V benchmarks, respectively. DualAlign also remains robust under missing-modality and label-scarce conditions.  \nKeywords Action Quality Assessment, Multi-Modal Action Quality Assessment, Muscle Weakness Assessment, Multi-Modal Alignment.  \n1 Introduction  \nAction Quality Assessment (AQA) aims to quantify the execution quality and correctness of human movements (Han et al., 2025b ; Xu et al., 2022b ; Dong et al., 2024) . It plays an essential role in sports analysis (Parmar and Tran Morris, 2017 ; Pan et al., 2019 ; Dong et al., 2026), skill assessment (Doughty et al., 2018 ; Gao et al., 2023), and healthcare (Zhou et al., 2023a ; Liu et al., 2021 ; Li et al., 2024a), where consistent and objective evaluation is essential. Most existing AQA systems predominantly rely on unimodal video input (Zhou et al., 2025b ; Xu et al., 2024c ; Xu et al., 2022b ; Ke et al., 2024). However, RGB videos cannot explicitly represent structural cues such as body pose configurations or high-level semantic descriptions of movement quality (Zhou et al., 2026a ; Yin et al., 2026) . For example, two rhythmic gymnastics movements may appear visually similar in RGB frames while differing substantially in ","cbCaiprZjUFkdRUx","https://ap.wps.com/l/cbCaiprZjUFkdRUx","pdf",9313880,2,1,25,"English","en",105,"# Introduction\n## Problem: limitations of unimodal and existing multi-modal AQA\n## Proposed approach: DualAlign two-stage adaptive alignment","[{\"question\":\"What is the goal of Action Quality Assessment (AQA)?\",\"answer\":\"AQA quantifies the execution quality and correctness of human movements. It is used in sports scoring, skill assessment, and healthcare where objective evaluation is required.\"},{\"question\":\"What challenges do existing multi-modal AQA methods face?\",\"answer\":\"They often suffer from cross-modal misalignment and unstable fusion due to heterogeneous modalities, and multi-modal annotation is costly, which restricts dataset scale and modality diversity.\"},{\"question\":\"How does DualAlign address these challenges?\",\"answer\":\"DualAlign first builds a coherent visual representation by maximizing shared structural information across RGB, optical flow, and skeleton. It then incorporates textual semantics after visual stabilization so descriptions complement rather than distort the visual representation.\"}]",1784186245,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"two-stage-multi-modal-fusion-with-adaptive-alignment-for-action-quality-assessment","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/two-stage-multi-modal-fusion-with-adaptive-alignment-for-action-quality-assessment/83249/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the goal of Action Quality Assessment (AQA)?","Question",{"text":75,"@type":76},"AQA quantifies the execution quality and correctness of human movements. It is used in sports scoring, skill assessment, and healthcare where objective evaluation is required.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What challenges do existing multi-modal AQA methods face?",{"text":80,"@type":76},"They often suffer from cross-modal misalignment and unstable fusion due to heterogeneous modalities, and multi-modal annotation is costly, which restricts dataset scale and modality diversity.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DualAlign address these challenges?",{"text":84,"@type":76},"DualAlign first builds a coherent visual representation by maximizing shared structural information across RGB, optical flow, and skeleton. It then incorporates textual semantics after visual stabilization so descriptions complement rather than distort the visual representation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]