[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84132-en":3,"doc-seo-84132-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84132,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation","Infrared remote-sensing imagery encodes intensity structure, object–background contrast, and illumination-invariant cues that RGB-trained models often miss. Remote-sensing vision-language resources have largely centered on visible semantics, leaving infrared understanding underexplored. MonoIR-RS introduces a large-scale infrared remote-sensing vision-language dataset and benchmark, generating 600,000 synthesized IR images with 59,032 IR-aware caption records. CLIP-style contrastive adaptation and VLM instruction tuning are fine-tuned across multiple backbones and evaluated on AVIID, showing IR-aware gains and near-zero RGB-color leakage.","arXiv :2607 .06552v 1 [ cs .CV] 7 Jul 2026  \nMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM  \nAdaptation  \nJiaju Han 1 ,3 , Ma Yaqi2 , Yahui Chai 1 , Xuemeng Sun 1 , Xin Li 1 , Qike Zhang 1 , Yingying Zhao 1 , Xiang Chen 1 , Luwei Yang3 , Chengyin Hu 1 , and Jiahuan  \nLong4  \n1 China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China  \n2 Guizhou University, Guiyang, China  \n3 Shenzhen Research Institute of Big Data, Shenzhen, China  \n4 Shanghai Jiao Tong University, Shanghai, China  \nAbstract. Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. We introduce MonoIRRS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains only the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions rewrite supervision around grayscale structure and infrared-style contrast instead of RGB appearance. We show that the synthesized infrared is markedly closer to real thermal imagery than a grayscale conversion  \non the AVIID benchmark. We fine-tune five CLIP backbones and six VLM backbones, and calibrate them against zero-shot behavior: IRaware adaptation lifts CLIP mean recall by up to +12 .8 points (best checkpoint 19.2% on the 9,720-image filtered split) and drives VLM captioning IR-cue coverage to 100% while reducing residual RGB-color leakage to near zero. By isolating the infrared modality from RGB–IR dualmodal learning, MonoIR-RS offers a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.  \nKeywords: Infrared remote sensing · Vision-language learning · CLIP fine-tuning · Visual instruction tuning · Dataset evaluation  \n1 Introduction  \nInfrared remote sensing is important for visual understanding under illumination changes, low-visibility conditions, and intensity contrast not captured by  \n2 J. Han et al.  \nFig. 1: Overview of the MonoIR-RS construction workflow. The shared FusionRS source pool is converted into an infrared image-text corpus, IR-aware text, and evaluation tasks covering retrieval, VLM understanding, and dataset-quality checks.  \nvisible-band imagery, where discriminative evidence includes object-background separation and grayscale response rather than color. This makes infrared imagery valuable for nighttime monitoring, fire observation, and low-visibility analysis, yet difficult to serve with vision-language models trained on RGB web photography. The gap goes beyond a domain shift: infrared and visible sensors register the same scene under different physical principles, so captions written for visible images frequently describe color cues that are absent or misleading in the infrared modality.  \nRecent vision-language models, including CLIP-style contrastive models and instruction-following VLMs, have made remote-sensing retrieval, captioning, and question answering more practical. RemoteCLIP and GeoRSCLIP show that domain-specific supervision improves remote-sensing representations, while GRAFT aligns satellite and ground-level imagery without direct text annotations [18, 23 , 35], and remote-sensing VLMs extend this toward grounded dialogue and geospatial benchmarking [9, 16 , 36] . Infrared-specific efforts such as Infrared-LLaVA and IRGPT have begun to address the modality gap, but they target general surveillance or pedestrian infrared rather than remote sensing and rely on real sensor captures that are scarce an","cbCaipIHrjyH7jVN","https://ap.wps.com/l/cbCaipIHrjyH7jVN","pdf",2840192,6,1,37,"English","en",105,"# Introduction\n## Problem: RGB-centric vision-language models vs infrared evidence\n## Approach: IR-aware data construction and adaptation\n## Experiments and evaluation on AVIID\n## Contributions","[{\"question\":\"What problem does MonoIR-RS address in infrared remote-sensing vision-language learning?\",\"answer\":\"It targets the mismatch between RGB-focused vision-language captions/models and infrared imagery, where thermal evidence and grayscale response differ from visible-band color cues.\"},{\"question\":\"How is the MonoIR-RS dataset constructed?\",\"answer\":\"It synthesizes infrared images from a shared visible remote-sensing source pool and rewrites supervision into IR-aware captions, retaining an infrared-only model-facing modality.\"},{\"question\":\"What improvements are reported from CLIP-style adaptation and VLM tuning?\",\"answer\":\"IR-aware adaptation increases CLIP mean recall by up to +12.8 points on AVIID, and VLM captioning achieves full IR-cue coverage while reducing residual RGB-color leakage close to zero.\"}]",1784193195,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"monoir-rs-infrared-remote-sensing-vision-language-learning-with-clip-and-vlm-adaptation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/monoir-rs-infrared-remote-sensing-vision-language-learning-with-clip-and-vlm-adaptation/84132/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does MonoIR-RS address in infrared remote-sensing vision-language learning?","Question",{"text":76,"@type":77},"It targets the mismatch between RGB-focused vision-language captions/models and infrared imagery, where thermal evidence and grayscale response differ from visible-band color cues.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the MonoIR-RS dataset constructed?",{"text":81,"@type":77},"It synthesizes infrared images from a shared visible remote-sensing source pool and rewrites supervision into IR-aware captions, retaining an infrared-only model-facing modality.",{"name":83,"@type":74,"acceptedAnswer":84},"What improvements are reported from CLIP-style adaptation and VLM tuning?",{"text":85,"@type":77},"IR-aware adaptation increases CLIP mean recall by up to +12.8 points on AVIID, and VLM captioning achieves full IR-cue coverage while reducing residual RGB-color leakage close to zero.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]