[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83572-en":3,"doc-seo-83572-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83572,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization","Indoor robot localization in GPS-denied settings is reformulated as semantic reasoning using a labeled top-down grid map together with a front camera image and a polar LiDAR scan. A fine-tuned Qwen2.5-VL-7B with LoRA is paired with a lightweight regression head that directly predicts continuous pose (x, y, θ) without text generation. Trained on a Gazebo simulation dataset (120,112 samples, 527 scenes), it reaches strong position and direction accuracy and remains robust when maps are incomplete. Cross-modal ablations confirm LiDAR reliability when camera semantics are unavailable.","Where Am I? Semantic Map Grounding via Vision-Language Models  \nfor Multi-Modal Localization  \nSuraj Borate 1 , Aarav Shah2 and Madhu Vadali3  \narXiv :2607 .0 1079v 1 [ cs .RO] 1 Jul 2026  \nAbstract—We address robot localization in GPS-denied indoor environments by reframing it as a semantic reasoning task rather than a geometric estimation problem. Motivated by the way humans localize using object-level cues and a labeled map, we ask: can a vision-language model (VLM), given a front camera image, a polar LiDAR scan, and a top-down semantic grid map, infer the robot’s pose?  \nWe fine-tune Qwen2.5-VL-7B with LoRA and attach a lightweight regression head that predicts continuous pose coordinates (x, y,θ) directly from the model’s final hidden state, bypassing text generation entirely. Training uses a composite position-and-direction loss with curriculum learning on a custom Gazebo simulation dataset (120,112 samples, 527 scenes).  \nOn the in-distribution test set (18,017 samples), the model achieves 98.23% position accuracy (PA), 98.00% direction accuracy (DA), 96.75% full pose accuracy (FPA), a mean position error of 0.11 m, and a mean orientation error of 5.7° at 0.62 s per sample. Position accuracy drops by only 7.2%(absolute) on 7 unseen object categories (90.99%), supporting genuine semantic spatial reasoning rather than appearance memorization. When maps are incomplete, fine-tuning recovers performance to 93.72% PA—demonstrating adaptability tostale or partial map information.  \nTwo ablations highlight cross-modal complementarity. When LiDAR is removed entirely (camera and map only, Exp. 9), PA remains at 95.06%, just 3.2% below the full system. However, when the camera provides no visible objects (wall-facing view, Exp. 6), LiDAR sustains PA at 92.33%—compared to 70.74% with neither LiDAR nor visible objects (Exp. 7)—demonstrating that LiDAR is the primary localization signal precisely when camera semantics are unavailable and acts as a reliable fallback under occlusion or sparse layouts.  \nI. INTRODUCTION AND PROBLEM FORMULATION  \nClassical robot localization-SLAM [2], particle filters [1], and metric visual localization [3], [4]-builds pose estimates by accumulating geometric constraints over time. These methods perform well with dense, consistent maps but degrade under perceptual noise, map incompleteness, or the absence of GPS.  \nHumans localize differently. Entering an unfamiliar room, a person identifies distinctive objects, estimates their distances, consults a labeled floor plan, and triangulates a position; if one cue fails, others compensate. This semantic reasoning over object-level priors is precisely what large VLMs are trained to support. Recent work shows that VLMs  \n*This work is supported by IIT Gandhinagar and Prime Ministers Research Fellowship  \n1 Suraj Borate, PhD Student, IIT Gandhinagar, Gandhinagar, Gujarat, India [surajb@iitgn.ac.in](surajb@iitgn.ac.in)  \n2Aarav Shah, Undergraduate Student, IIT Gandhinagar, Gujarat, India [aarav.shah@iitgn.ac.in](aarav.shah@iitgn.ac.in)  \n3Madhu Vadali, Associate Professor, IIT Gandhinagar, Gujarat, India [madhu.vadali@iitgn.ac.in](madhu.vadali@iitgn.ac.in)  \nsuch as GPT-4V [7], LLaVA [8], and Qwen-VL [15] can describe spatial relationships, reason about viewpoints, and cross-reference visual and textual information. Transformer attention over long token sequences supports implicit multistep inference [11], [12] the same deliberative process underlying human map reading.  \nWe ask: can a VLM, given a labeled semantic map, a camera image, and a LiDAR scan, localize a robot the way a human would? Our hypothesis is that fine-tuning unlocks the VLM’s latent capacity to cross-reference visual observations against map structure and output a continuous pose estimate. Related semantic localization methods use object-level maps [5] or scene graphs [6] but rely on hand-crafted geometric detectors rather than learned reasoning. VLMs have been applied to navigation [9] and rob","cbCaiaLJNGifyVoo","https://ap.wps.com/l/cbCaiaLJNGifyVoo","pdf",30734357,3,1,5,"English","en",105,"# Abstract\n# Introduction and Problem Formulation\n## Problem statement\n## Contributions\n# Method\n## Architecture","[{\"question\":\"How does the method approach robot localization in GPS-denied indoor environments?\",\"answer\":\"It reframes localization as a semantic reasoning task rather than geometric estimation, using object-level cues from a labeled semantic grid map along with a front camera image and a polar LiDAR scan.\"},{\"question\":\"What model and training strategy are used to predict the robot pose?\",\"answer\":\"Qwen2.5-VL-7B is fine-tuned with LoRA and equipped with a PoseHead regression module that outputs continuous pose coordinates (x, y, θ) from the model’s final hidden state, trained with a composite position-and-direction loss using curriculum learning on a custom Gazebo dataset.\"},{\"question\":\"How accurate is the proposed system, and what happens when the map is incomplete?\",\"answer\":\"On the in-distribution test set it achieves 98.23% position accuracy, 98.00% direction accuracy, and 96.75% full pose accuracy, with a mean position error of 0.11 m and orientation error of 5.7°. When maps are incomplete, performance recovers to 93.72% position accuracy after fine-tuning.\"}]",1784188928,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"where-am-i-semantic-map-grounding-via-vision-language-models-for-multi-modal-localization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/where-am-i-semantic-map-grounding-via-vision-language-models-for-multi-modal-localization/83572/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the method approach robot localization in GPS-denied indoor environments?","Question",{"text":75,"@type":76},"It reframes localization as a semantic reasoning task rather than geometric estimation, using object-level cues from a labeled semantic grid map along with a front camera image and a polar LiDAR scan.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What model and training strategy are used to predict the robot pose?",{"text":80,"@type":76},"Qwen2.5-VL-7B is fine-tuned with LoRA and equipped with a PoseHead regression module that outputs continuous pose coordinates (x, y, θ) from the model’s final hidden state, trained with a composite position-and-direction loss using curriculum learning on a custom Gazebo dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"How accurate is the proposed system, and what happens when the map is incomplete?",{"text":84,"@type":76},"On the in-distribution test set it achieves 98.23% position accuracy, 98.00% direction accuracy, and 96.75% full pose accuracy, with a mean position error of 0.11 m and orientation error of 5.7°. When maps are incomplete, performance recovers to 93.72% position accuracy after fine-tuning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]