[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82061-en":3,"doc-seo-82061-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82061,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Mixture of Probes Learning from Privileged Modalities in Multimodal LLMs Through Probing","Multimodal Large Language Models (MLLMs) are often trained assuming every modality used at training is available at inference, yet real deployments frequently follow a privileged modality setting where auxiliary modalities exist only during training. Mixture of Probes (MoP) disentangles modality-specific and modality-general signals inside a multimodal LLM to preserve modality-dependent structure while learning transferable representations. A structured probing mechanism extracts organized information from intermediate encoder states rather than only final-layer alignment. MoP Cross-modal Training (MoP-X) adds a probe disentanglement loss to prevent collapse and promote cross-modal learning. Across two domains, eight tasks, and four modalities, MoP consistently outperforms strong baselines, reaching up to 65% relative improvement under modality-only inference.","Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing  \nDominick Reilly 1 ,3 ∗ Qiyu Wu 1† Hiromi Wakaki 1 Srijan Das3 Yuki Mitsufuji 1 ,2  \n1 Sony Group Corporation 2 Sony AI  \n3 University of North Carolina at Charlotte  \n[dreilly1@charlotte.edu](dreilly1@charlotte.edu) [firstname.lastname@sony.com](firstname.lastname@sony.com)  \n[https://github.com/Sony/MoP](https://github.com/Sony/MoP)  \narXiv :2607 .08839v 1 [ cs .CV] 9 Jul 2026  \nFigure 1: (left) MoP improves single-modality inference across five modalities, outperforming both unimodal MLLMs and naive multimodal training. (right) MoP achieves this through a structured probing mechanism with modality-specific and modality-general probes, which disentangle modality representations before integration with the LLM.  \nAbstract  \nMultimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training. Code, model checkpoints, and evaluation protocols will be made available at [https://github.com/Sony/MoP](https://github.com/Sony/MoP).  \n∗ Work done during internship at Sony Group Corporation  \n†Corresponding Author: [qiyu.wu@sony.com](qiyu.wu@sony.com)  \nPreprint.  \n1 Introduction  \nStemming from the success of large language models (LLMs) [39, 4, 38], recent multimodal large language models (MLLMs) have enabled language-based interaction with diverse sensory inputs such as vision, audio, and depth [48, 14, 27, 42] . Yet real-world scenarios often exhibit an asymmetric structure: the sensors that are available for training are not always available for deployment. A model may be trained with rich auxiliary streams, such as egocentric video, depth, or audio, but later be required to operate from a single cheap, reliable, or non-invasive modality. This creates a privileged modality setting, where auxiliary modalities provide supervision during training but are unavailable at inference. For example, in activities of daily living (ADL), training data may include egocentric video from wearable devices [11, 10], depth from specialized sensors [15], and exocentric video from static cameras [9, 6], while deployment is typically restricted to exocentric video alone.  \nExisting MLLMs [49, 19, 36, 2] are not designed to address this setting. They primarily view them","cbCaiqOzfSc9wjWv","https://ap.wps.com/l/cbCaiqOzfSc9wjWv","pdf",1336544,1,16,"English","en",105,"# Abstract\n# Introduction\n## Privileged modality setting\n## Limitations of existing MLLMs\n## Proposed Mixture of Probes (MoP)","[{\"question\":\"What is the privileged modality setting in multimodal LLMs?\",\"answer\":\"It is a scenario where auxiliary modalities are available during training but not accessible during inference. The model must therefore learn from training-only information and still perform using only the deployment modality.\"},{\"question\":\"How does Mixture of Probes (MoP) improve cross-modal knowledge transfer?\",\"answer\":\"MoP uses a structured probing mechanism to disentangle modality-specific and modality-general signals. It extracts organized information from intermediate encoder representations so that transferable knowledge can be learned beyond final-layer alignment.\"},{\"question\":\"What role does MoP Cross-modal Training (MoP-X) play?\",\"answer\":\"MoP-X introduces a probe disentanglement loss designed to prevent probe collapse and encourage cross-modal learning. This supports more effective disentanglement during training.\"}]",1784177923,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mixture-of-probes-learning-from-privileged-modalities-in-multimodal-llms-through-probing","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mixture-of-probes-learning-from-privileged-modalities-in-multimodal-llms-through-probing/82061/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the privileged modality setting in multimodal LLMs?","Question",{"text":75,"@type":76},"It is a scenario where auxiliary modalities are available during training but not accessible during inference. The model must therefore learn from training-only information and still perform using only the deployment modality.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Mixture of Probes (MoP) improve cross-modal knowledge transfer?",{"text":80,"@type":76},"MoP uses a structured probing mechanism to disentangle modality-specific and modality-general signals. It extracts organized information from intermediate encoder representations so that transferable knowledge can be learned beyond final-layer alignment.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does MoP Cross-modal Training (MoP-X) play?",{"text":84,"@type":76},"MoP-X introduces a probe disentanglement loss designed to prevent probe collapse and encourage cross-modal learning. This supports more effective disentanglement during training.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":28,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]