[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85882-en":3,"doc-seo-85882-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85882,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis","Large Multimodal Models (LMMs) achieve strong performance on RGB images, yet reliable generalization to unseen visual modalities is still underexplored. The work frames modalities as different samplings of the same physical world, requiring modality-agnostic scene semantics plus modality-specific adaptability. It introduces VVM-Tuning, which synthesizes appearance-varied images from RGB scenes, aligns appearance with language concepts, and uses modality contexts with instruction tuning for zero-shot adaptation. It further proposes VVM-Bench with six real and synthetic modalities to evaluate both semantic perception and modality understanding, showing consistent gains without in-modality training.","arXiv :2607 . 10308v1 [ cs .CV] 11 Jul 2026  \nGeneralize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis  \nShihao Yuan 1 , Yuanze Li 1 , Ruyi Zhang 1 , Ming Liu 1(􀀌), and Wangmeng Zuo 1   \n{csshihao, sqleopop, [csmliu}@outlook.com](csmliu}@outlook.com),{ruyi.zhang.maggie, [cswmzuo}@gmail.com](cswmzuo}@gmail.com)  \n1 Faculty of Computing, Harbin Institute of Technology, Harbin, China  \nAbstract. Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world.  \nTherefore, effective generalization requires models to possess both modalityagnostic perception of scene semantics and the adaptability to modalityspecific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at [https://github.com/Hunter-Will/VVM-Tuning](https://github.com/Hunter-Will/VVM-Tuning).  \nKeywords: LMMs · Instruction Tuning · Synthetic Data  \n1 Introduction  \nCurrent Large Multimodal Models (LMMs) demonstrate impressive visual performance on RGB images [10]; however, the generalization boundary on other visual modalities (e.g . infrared, depth, etc.) remains largely unexplored yet. Existing non-RGB vision in current LMMs primarily depends on the non-RGB data incorporated during the data scaling process [1,4,33] . This leads to a generalization gap when encountering unseen modalities, as the non-RGB vision is built on in-modality data and limited modality coverage.  \n2 Shihao Yuan, Yuanze Li, et al.  \nIrrespective of visual modality, basic  \nvisual elements of appearance can be recognized generically  \nThermal Image RGB Image  \nSemantic Commonalities  \nBoth images share the  \nsame intrinsic semantics  \nModality context helps the  \nadaptation to unseen modalities  \nShared semantics can be extract  \nin a unified way  \nFig. 1: An overview of our idea. Different visual modalities (Thermal and RGB) are both signal samplings of our physical world, sharing the same underlying semantics (formed by the person and the car) . They also share basic visual concepts, asthe color itself is invariant across modalities. The modality-specific physical meaning is encoded by the rearrangement and remapping of basic visual concepts. Thus, we introduce textual context to complement such knowledge during inference.  \nTo address this gap and explore the generalization boundary of LMMs, we raise the question: Is it possible for LMMs to generalize across versatile visual modalities without training on in-modality data?  \nThe solution to the problem starts from the nature of visual modalities, as shown in Fig. 1, though with different appearances, they are essentially digital signals collected by sensors (thermal and RGB camera in Fig. 1) from the same physical space. Therefore, images from different visual modalities can be viewed ","cbCailEAlTpkXGub","https://ap.wps.com/l/cbCailEAlTpkXGub","pdf",2780947,5,1,40,"English","en",105,"# Abstract\n# Introduction\n# Core Idea (Modality-unaware Perception)\n# Modality Context and Synthesis Framework\n# VVM-Bench Evaluation","[{\"question\":\"Why do LMMs struggle to generalize to unseen visual modalities?\",\"answer\":\"The document states that current non-RGB capability mainly relies on non-RGB data included during scaling, which leads to a generalization gap when encountering modalities not well covered by training data.\"},{\"question\":\"What is the main idea behind VVM-Tuning?\",\"answer\":\"VVM-Tuning trains LMMs using modality synthesis and modality contexts so the model can learn modality-agnostic semantic understanding while also mapping appearance variations back to modality-related attributes.\"},{\"question\":\"How does the paper evaluate the proposed approach?\",\"answer\":\"It introduces VVM-Bench, a benchmark with six real and synthetic modalities, to assess semantic perception and modality understanding; experiments show improvements on real-world and novel synthetic modalities without in-modality training.\"}]",1784206931,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"generalize-lmms-to-versatile-visual-modalities-via-fabricated-modality-synthesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/generalize-lmms-to-versatile-visual-modalities-via-fabricated-modality-synthesis/85882/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do LMMs struggle to generalize to unseen visual modalities?","Question",{"text":76,"@type":77},"The document states that current non-RGB capability mainly relies on non-RGB data included during scaling, which leads to a generalization gap when encountering modalities not well covered by training data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the main idea behind VVM-Tuning?",{"text":81,"@type":77},"VVM-Tuning trains LMMs using modality synthesis and modality contexts so the model can learn modality-agnostic semantic understanding while also mapping appearance variations back to modality-related attributes.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the paper evaluate the proposed approach?",{"text":85,"@type":77},"It introduces VVM-Bench, a benchmark with six real and synthetic modalities, to assess semantic perception and modality understanding; experiments show improvements on real-world and novel synthetic modalities without in-modality training.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]