[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84781-en":3,"doc-seo-84781-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84781,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Beyond Modality Fusion Deep Ensembles for Multimodal Classification","Multimodal classification often relies on late-fusion, where unimodal networks extract modality-specific features and the system classifies concatenated representations. When modality imbalance is strong, regularization and fusion-focused methods aim to rebalance learning but may still underperform. This work shows that deep ensembles of unimodal networks can classify multimodal data without explicit fusion, consistently outperforming state-of-the-art late-fusion under equal parameter budgets and across additional fusion alternatives. A heuristic selects ensemble model counts per modality to avoid costly search. Synthetic and real benchmarks with controlled modality counts and predictive strength, plus scaling-law fits to bimodal datasets, estimate asymptotic ensemble performance.","Beyond Modality Fusion: Deep Ensembles for Multimodal Classification  \nIlya Burenko  \nScaDS.AI, Technische Universität Dresden Dresden, Germany ilia.burenko@tu-dresden.de  \nDmitry Vetrov  \nConstructor University Bremen, Germany  \narXiv :2607 .050 19v2 [ cs .LG] 7 Jul 2026  \nAbstract  \nIn multimodal classification, late-fusion approaches classify concatenated modalityspecific features extracted by unimodal neural networks. When modality imbalance is pronounced, various regularization techniques have been proposed to balance the learning process and overcome the inferior performance of late-fusion networks. In contrast, this work demonstrates that multimodal data can be effectively classified without any explicit modality fusion, using deep ensembles of unimodal networks.  \nWe systematically compare deep ensembles to late-fusion networks at equal parameter count and show that ensembles consistently outperform state-of-the-art late-fusion methods designed to address modality imbalance. This advantage also holds over intermediate-fusion techniques we evaluated and over hybrid methods that combine unimodal and multimodal predictions. We propose and empirically validate a method for selecting the number of models per modality in an ensemble, avoiding computationally expensive exhaustive search. Under extreme modality imbalance and small ensemble sizes, the heuristic indicates that ensembles of unimodal models trained solely on the stronger modality are preferable; as the ensemble scales up, incorporating models from the weaker modality becomes beneficial. Both predictions align with our empirical findings. To systematically explore the challenges of optimizing multimodal models, we propose a synthetic multimodal framework that allows control over both the number of modalities and their predictive strength; our findings are consistent across synthetic and real-world datasets. Finally, by fitting scaling laws to bimodal datasets, we estimate the asymptotic performance of ensembles.  \n1 Introduction  \nIn training deep neural networks for classification tasks, the primary objective is to extract discriminative features that most effectively explain the output labels. The success of modern unimodal neural networks is attributed to the availability of large datasets and scalable architectures [51], such as Transformers [52, 11, 7] and modality-specific models [48, 55, 34] . These models incorporate modality-focused inductive biases, as seen in convolutional neural networks [28, 34], recurrent neural networks [6, 46], and graph neural networks [39] . Hybrid approaches further integrate multiple architectural advantages [45, 16, 49] .  \nFor multimodal classification, classifiers use multiple input modalities, which may originate from heterogeneous data sources [33] . Combining intermediate features extracted from given modalities can make the output feature vector more informative than in unimodal scenarios. An extensive body of work has studied fusion techniques, from early-and intermediate-fusion to late-fusion approaches [15, 41, 54, 59–61] . We refer to [33] for a comprehensive survey. Typically, the existing fusion approaches for multimodal classification use unimodal neural networks that perform well on the given modalities and optimize the interplay between the unimodal networks.  \nPreprint.  \nRegardless of the type of fusion, access to multiple modalities should improve classification. However, research has demonstrated that naive late-fusion networks, which classify based on concatenated features from unimodal networks, can yield lower accuracy compared to the best-performing unimodal classifier [53] . To address this unexpected outcome, recent studies have examined the problem from both theoretical [20, 36, 62] and practical perspectives [43], seeking optimal methods for combining unimodal features in multimodal classification. Although some progress has been achieved in this respect [22, 26], it is still an open question how to o","cbCaimdLnfZPX5Oq","https://ap.wps.com/l/cbCaimdLnfZPX5Oq","pdf",2640016,3,1,39,"English","en",105,"# Introduction\n## Problem setting: multimodal classification and fusion\n## Proposed approach: heterogeneous and homogeneous deep ensembles\n## Contributions and evaluation plan","[{\"question\":\"What does the paper propose for multimodal classification when modalities are imbalanced?\",\"answer\":\"It proposes using deep ensembles of unimodal networks that perform classification without any explicit modality fusion, and shows they remain effective under pronounced modality imbalance.\"},{\"question\":\"How do deep ensembles compare to late-fusion networks in the experiments?\",\"answer\":\"Deep ensembles consistently outperform state-of-the-art late-fusion methods designed to address modality imbalance, with comparisons performed at equal parameter counts.\"},{\"question\":\"How is the number of models per modality in an ensemble determined?\",\"answer\":\"The paper introduces and validates a heuristic that selects the number of models per modality, avoiding computationally expensive exhaustive search and showing a transition from favoring stronger-modality models to incorporating weaker-modality models as the ensemble grows.\"}]",1784198194,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-modality-fusion-deep-ensembles-for-multimodal-classification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-modality-fusion-deep-ensembles-for-multimodal-classification/84781/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper propose for multimodal classification when modalities are imbalanced?","Question",{"text":75,"@type":76},"It proposes using deep ensembles of unimodal networks that perform classification without any explicit modality fusion, and shows they remain effective under pronounced modality imbalance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do deep ensembles compare to late-fusion networks in the experiments?",{"text":80,"@type":76},"Deep ensembles consistently outperform state-of-the-art late-fusion methods designed to address modality imbalance, with comparisons performed at equal parameter counts.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the number of models per modality in an ensemble determined?",{"text":84,"@type":76},"The paper introduces and validates a heuristic that selects the number of models per modality, avoiding computationally expensive exhaustive search and showing a transition from favoring stronger-modality models to incorporating weaker-modality models as the ensemble grows.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]