[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122140-en":3,"doc-seo-122140-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122140,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Wasserstein Modality Alignment Makes Your Multimodal Transformer More Robust","Multimodal fusion using a multimodal transformer can perform early and late fusion effectively, yet self-attention alone was designed for unimodal token sequences and may harm modality alignment. Empirical findings show that relying only on self-attention makes the model vulnerable to missing modalities and input noise due to over-dependence on a single modality. Wasserstein Modality Alignment (WMA) introduces an implicit Wasserstein-distance based alignment without adding trainable parameters. Experiments across four datasets with 2- and 3-modality settings and early/late fusion demonstrate consistent gains in both performance and robustness over baselines.","Wasserstein Modality Alignment Makes Your Multimodal Transformer More Robust  \nZhuo Zhi  \nDepartment of Electronic and Electrical Engineering University College London ∗  \nYuxuan Sun  \nDepartment of Electronic and Electrical Engineering University College London  \nQiangqiang Wu  \nDepartment of Computer Science City University of Hong Kong  \nZiquan Liu  \nSchool of Electronic Engineering and Computer Science Queen Mary University of London ∗  \nMiguel Rodrigues  \nDepartment of Electronic and Electrical Engineering University College London  \n[zhuo.zhi.21@ucl. ac.uk](zhuo.zhi.21@ucl. ac.uk)  \n[yuxuan. sun.22@ucl. ac.uk](yuxuan. sun.22@ucl. ac.uk)  \n[qiangqwu2@cityu. edu.hk](qiangqwu2@cityu. edu.hk)  \n[ziquan.liu@qmul. ac.uk](ziquan.liu@qmul. ac.uk)  \n[m. rodrigues@ucl. ac.uk](m. rodrigues@ucl. ac.uk)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= 2IkaUZdB62](https: // openreview. net/ forum? id= 2IkaUZdB62)  \nAbstract  \nMultimodal fusion with a multimodal transformer is an effective method for both early and late fusion paradigms. However, in a multimodal transformer, the modality fusion is performed solely through the self-attention mechanism, which is originally designed for unimodal token sequences. To improve the self-attention mechanism for handling multimodal input, a parametric adapter model, like the Q-former in BLIP-2, is often used to align tokens from different modalities. Our empirical study unveils that only using the self-attention layer to perform the modality fusion makes the model less robust to missing modalities and input noise, as the model will overly rely on one certain modality. To improve the robustness of the transformer, our paper proposes an implicit approach based on Wasserstein distance that aligns tokens from different modalities without using any additional trainable parameters.  \nOur empirical study shows that the implicit modality alignment improves the effectiveness of the multimodal Transformer in discriminative tasks, as well as its robustness to input noise and missing modalities. We conduct experiments on four downstream task datasets, including 2-modalities and 3-modalities tasks. We also consider different fusion paradigms, i.e., early and late fusion. The experimental results show that our proposed method has a significant improvement in both performance and robustness over all baselines across all datasets and fusion paradigms.  \n1 Introduction  \nMultimodal machine learning (MML) mimics human perception by integrating multiple modalities such as text, audio, images, video, and sensor data to form a comprehensive understanding of the world. Many  \n∗ Corresponding author  \nmultimodal models have been applied to various tasks like multimodal medical diagnostics Hayat et al.(2022a), sentiment analysis Zadeh et al. (2018) and malicious speech detection Kiela et al. (2020) .  \nAligning heterogeneous data in multimodal learning is crucial since such data often exhibit distinct distributions, representations, and noise levels. Proper alignment enhances the uniform representation of these diverse data types, leading to improved performance and robustness in multimodal tasks Ghahremani Boozandani & Wachinger (2024); Liang et al. (2024); Kim et al. (2020) . To achieve better modality alignment, various strategies are applied in large-scale multimodal models, such as the Q-Former in BLIP-2 Li et al. (2023), contrastive learning in CLIP Radford et al. (2021) and Imagebind Girdhar et al. (2023) .  \nMultimodal fusion based on a one-tower transformer, named as multimodal transformer (MT), by its flexibility and simplicity, are widely used for a variety of multimodal learning tasks Lee et al. (2023); Nagraniet al. (2021); Zhi et al. (2024); Ma et al. (2021) . Although the multimodal transformer can handle multimodal tokens as the input due to the flexibility of self-attention layers, it lacks a mechanism for modality alignment during the fine-tuning process. In other words, it is not opti","cbCaicMry58iK8jn","https://ap.wps.com/l/cbCaicMry58iK8jn","pdf",903502,1,21,"English","en",105,"# Introduction\n# Related Work\n# Method: Wasserstein Modality Alignment (WMA)\n# Experiments\n# Results and Analysis\n# Conclusion","[{\"question\":\"Why does self-attention alone lead to weaker robustness in multimodal transformers?\",\"answer\":\"Self-attention is originally designed for unimodal token sequences, so modality fusion becomes implicit through attention weights. This can cause the model to over-rely on a particular modality, making it less robust to missing modalities and input noise.\"},{\"question\":\"What is Wasserstein Modality Alignment (WMA) in this work?\",\"answer\":\"WMA is an implicit approach that aligns tokens from different modalities using a Wasserstein-distance-based regularization. It does so without introducing additional trainable parameters.\"},{\"question\":\"How is WMA evaluated across tasks and fusion paradigms?\",\"answer\":\"The study runs experiments on four downstream datasets, covering both 2-modality and 3-modality settings. It also compares early and late fusion paradigms, reporting improvements in performance and robustness across baselines.\"}]","Wasserstein Modality Alignment Makes Your Multimodal Transformer More Robust | PDF",1785809030,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wasserstein-modality-alignment-makes-your-multimodal-transformer-more-robust","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/wasserstein-modality-alignment-makes-your-multimodal-transformer-more-robust/122140/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does self-attention alone lead to weaker robustness in multimodal transformers?","Question",{"text":75,"@type":76},"Self-attention is originally designed for unimodal token sequences, so modality fusion becomes implicit through attention weights. This can cause the model to over-rely on a particular modality, making it less robust to missing modalities and input noise.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is Wasserstein Modality Alignment (WMA) in this work?",{"text":80,"@type":76},"WMA is an implicit approach that aligns tokens from different modalities using a Wasserstein-distance-based regularization. It does so without introducing additional trainable parameters.",{"name":82,"@type":73,"acceptedAnswer":83},"How is WMA evaluated across tasks and fusion paradigms?",{"text":84,"@type":76},"The study runs experiments on four downstream datasets, covering both 2-modality and 3-modality settings. It also compares early and late fusion paradigms, reporting improvements in performance and robustness across baselines.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]