[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84386-en":3,"doc-seo-84386-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84386,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition Evidence and Limits","Multimodal emotion and sentiment recognition often uses early fusion or late fusion, each with distinct trade-offs in accuracy, modularity, and cross-modal interaction. The study revisits XAI-guided adaptive fusion (XGAF), a tree-based mixture of unimodal and cross-modal experts whose sample-level expert weights are derived from TreeSHAP attribution magnitudes. When experts have unequal feature dimensionalities, SHAP attribution reduction must be sum-abs rather than mean-abs. On MELD, sum-abs XGAF matches early fusion under multiple face-sequence aggregators and significantly outperforms probability-average late fusion. On CMU-MOSEI, sum-abs XGAF improves weighted-F1 over early and late fusion with statistically supported gains. Ablations show most benefits come from adding cross-modal experts, not complex per-sample routing.","arXiv :2607 .08573v 1 [ cs .AI] 9 Jul 2026  \nSHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits  \nAdis Alihodzic 1* and Selma Skopljakovic Hubljar 1  \n1* Department of Mathematical and Computer Sciences, Faculty of Science, University of  \nSarajevo, Sarajevo, Bosnia and Herzegovina.  \n*Corresponding author(s). E-mail(s): [adis.alihodzic@pmf.unsa.ba](adis.alihodzic@pmf.unsa.ba) ;  \nContributing authors: [selma.skoplajkovic@gmail.com](selma.skoplajkovic@gmail.com) ;  \nAbstract  \nMultimodal emotion and sentiment recognition commonly relies on either early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors. Early fusion is often accurate but monolithic, whereas late fusion is modular but can lose cross-modal interactions. This paper revisits XAI-guided adaptive fusion (XGAF), a tree-based mixture of unimodal and cross-modal experts whose sample-level weights are derived from TreeSHAP attribution magnitudes. The study focuses on a methodological issue that becomes important when experts have unequal feature dimensionalities: reducing feature attributions by mean absolute SHAP values can suppress high-dimensional cross-modal experts, while reducing them by summed absolute SHAP values preserves total attribution mass. On MELD 7-class emotion recognition, the proposed sum-abs reduction closes the gap between the cross-modal expert mixture and early fusion across three face-sequence aggregators, with the Transformer variant reaching 0.5983 weighted-F1 compared with 0.6018 for early fusion and 0.4598 for probability-average late fusion. McNemar testing shows no significant difference between sum-abs XGAF and early fusion on MELD (p = 1 .000), while XGAF remains significantly better than late fusion (p \u003C 0.0001) . On CMU-MOSEI 3-class sentiment recognition, sum-abs XGAF reaches 0.6519 weighted-F1 compared with 0.6485 for early fusion and 0.5696 for late fusion, with a small but statistically significant improvement over early fusion (p = 0 .0452) . An expert-pool ablation indicates that most of the gain comes from adding crossmodal experts, especially the trimodal expert, rather than from rich per-sample routing. A focused three-way ablation further confirms that median-abs SHAP behaves similarly to mean-abs SHAP (0.5682 versus 0.5669 weighted-F1 on MELD), while sum-abs SHAP reaches early-fusion-level performance (0.5957 versus 0.5955) . Diagnostic analysis shows that mean-abs and median-abs weights are nearly uniform, whereas sum-abs weights become strongly concentrated on the trimodal expert. The main contribution is therefore not a new state-of-the-art recognition model, but a transparent empirical study of how SHAP attribution reduction, expert dimensionality, and cross-modal expert design affect modular multimodal fusion.  \nKeywords: multimodal fusion, emotion recognition, sentiment analysis, explainable artificial intelligence, SHAP, XGBoost, MELD, CMU-MOSEI, mixture of experts  \n1 Introduction  \nMultimodal affective computing attempts to infer human emotion, sentiment, or related affective states from several complementary channels, most commonly text, speech, and visual cues. A sentence may look positive in text but sound sarcastic in speech, or a neutral phrase may become emotionally informative when facial expression is considered. This simple observation explains why fusion is not a minor implementation detail but one of the central design choices in multimodal emotion recognition. Benchmark datasets such as MELD [1] and CMU-MOSEI [2] have made it possible to compare modelling strategies under controlled experimental protocols, but the question of how modalities should be fused remains central. Recent surveys emphasise that modern MER systems must be evaluated not only by accuracy but also by fusion design, robustness to missing or noisy modalities, generalisation across users and datasets, and int","cbCaiiWU4YKsjNf6","https://ap.wps.com/l/cbCaiiWU4YKsjNf6","pdf",717810,3,1,12,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the paper address in SHAP-weighted cross-modal expert fusion?\",\"answer\":\"It addresses how to reduce SHAP attributions when experts have unequal feature dimensionalities, since mean-abs reduction can suppress high-dimensional cross-modal experts.\"},{\"question\":\"How does sum-abs SHAP weighting affect performance on MELD and CMU-MOSEI?\",\"answer\":\"On MELD 7-class emotion recognition, sum-abs XGAF reaches early-fusion-level weighted-F1 and significantly outperforms late fusion. On CMU-MOSEI 3-class sentiment recognition, sum-abs XGAF achieves higher weighted-F1 than early fusion with a small but statistically significant improvement.\"},{\"question\":\"What do the ablation and diagnostic analyses suggest about where the gains come from?\",\"answer\":\"Most gains come from adding cross-modal experts, especially the trimodal expert, rather than from rich per-sample routing. Mean-abs and median-abs weights are nearly uniform, while sum-abs weights concentrate strongly on the trimodal expert.\"}]",1784195234,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"shap-weighted-cross-modal-expert-fusion-for-emotion-and-sentiment-recognition-evidence-and-limits","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/shap-weighted-cross-modal-expert-fusion-for-emotion-and-sentiment-recognition-evidence-and-limits/84386/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in SHAP-weighted cross-modal expert fusion?","Question",{"text":75,"@type":76},"It addresses how to reduce SHAP attributions when experts have unequal feature dimensionalities, since mean-abs reduction can suppress high-dimensional cross-modal experts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does sum-abs SHAP weighting affect performance on MELD and CMU-MOSEI?",{"text":80,"@type":76},"On MELD 7-class emotion recognition, sum-abs XGAF reaches early-fusion-level weighted-F1 and significantly outperforms late fusion. On CMU-MOSEI 3-class sentiment recognition, sum-abs XGAF achieves higher weighted-F1 than early fusion with a small but statistically significant improvement.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the ablation and diagnostic analyses suggest about where the gains come from?",{"text":84,"@type":76},"Most gains come from adding cross-modal experts, especially the trimodal expert, rather than from rich per-sample routing. Mean-abs and median-abs weights are nearly uniform, while sum-abs weights concentrate strongly on the trimodal expert.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]