[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82898-en":3,"doc-seo-82898-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82898,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","LangLoc: Tell Me What You See","LangLoc addresses fine-grained indoor localization from natural language by inferring an observer’s 2D floor position and heading inside a known 3D environment. Lightweight, privacy-preserving language queries avoid bandwidth-heavy camera uploads and the sensitive imagery problem. The three-stage pipeline retrieves the correct scene with a dual-branch GATv2 encoder using CLIP features, estimates pose via visibility-based floor-grid scoring (median error 0.95 m), and resolves ambiguities using a Bayesian dialog module with targeted yes/no questions. A benchmark with 13,000+ pose-indexed descriptions over 1,300+ scans supports the task; code and data will be released.","arXiv :2607 .05077v 1 [ cs .CV] 6 Jul 2026  \nLangLoc: “Tell Me What You See”  \nShaurya Kishore Panwar 1 ,2⋆, Roham Zendehdel Nobari 1 ,2∗, Shirley Feng Yi Lau 1 ,2∗, Abu Bakr Rahman Shaik 1 ,2 ∗, Manuel Günther2, Marc Pollefeys 1 ,3 ,  \nand Daniel Barath 1 ,4  \n1 ETH Zürich, Switzerland  \n2 University of Zürich, Switzerland  \n3 Microsoft  \n4 HUN-REN SZTAKI, Hungary  \nAbstract. We tackle fine-grained indoor localization from natural language: given a free-form description of one’s surroundings, estimate the observer’s 2D position and heading within a known 3D environment.  \nLanguage queries are lightweight, privacy-preserving, and need no camera – yet prior work stops at coarse scene retrieval and cannot resolve an intra-scene pose. We close this gap with LangLoc, a three-stage pipeline that (i) retrieves the correct scene via a dual-branch GATv2 encoder with CLIP semantic features, surpassing the previous best by 8 percentage points in Top-1 recall; (ii) estimates position and heading by scoring a dense floor grid through ray-cast object visibility, reaching a median error of 0.95 m; and (iii) resolves residual ambiguity through a Bayesian dialog module that asks targeted yes/no questions and updatesa pose posterior until the location is pinpointed. To support this task we contribute a benchmark of 13 ,000+ pose-indexed natural-language descriptions over 1 ,300+ indoor 3D scans. Code and data will be released.  \nProject page: [https://rzninvo.github.io/Lang-Loc/](https://rzninvo.github.io/Lang-Loc/) .  \n1 Introduction  \nKnowing where you are is fundamental to almost every location-aware service: indoor navigation, robot assistance, augmented reality, and emergency response all require an accurate pose estimate.  \nThe dominant localization paradigm today is visual: a device captures animage or video stream, uploads it to a server, and receives a pose estimate in return [32, 36] . While effective, this approach carries significant drawbacks. Image transmission is bandwidth-heavy, especially indoors where frequent queries are needed. More critically, it is privacy-invasive: photos of homes, offices, and hospitals inevitably capture sensitive information that users may not wish to share. Finally, capturing a useful image is itself non-trivial – a photo of a blank wall carries little discriminative information, requiring users to know how to frame an informative shot.  \n⋆ Equal contributions.  \n2 S. Panwar et al.  \nLanguage offers a compelling alternative. Telling a system “I’m standing in front of a bookshelf, with a blue sofa on my left and a TV across the room” is natural, fast, and transmits almost no personally identifiable information. A text description is orders of magnitude smaller than an image, requires no camera or special hardware, and mirrors how people naturally communicate their whereabouts to one another in everyday life. This makes language localization natural in camera-prohibited but digitally-twinned settings such as hospitals, labs, and emergency dispatch. Beyond localization, human-to-agent communication also requires grounding free-form verbal goals (e.g.,“go to the bookshelf and find thered book”) into precise 3D poses for robots, drones, and AR assistants.  \nDespite this appeal, language-based localization remains largely unsolved. Existing methods address only coarse scene retrieval – identifying which room in a database a description refers to [14, 23] . Resolving a precise pose within a scene from language is an open problem: many viewpoints within the same room share similar semantics, differing only in subtle geometric or visibility cues that are difficult to capture in plain text.  \nWe present LangLoc, the first pipeline for fine-grained indoor localization from natural language. Given a free-form description and a database of 3D scenes, LangLoc first retrieves the correct scene – surpassing the prior state of the art by 8 percentage points in Top-1 recall – and then estimates a 2D floor position and","cbCaioORJRfGbIDX","https://ap.wps.com/l/cbCaioORJRfGbIDX","pdf",14038893,1,18,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What does LangLoc estimate from a natural-language description?\",\"answer\":\"LangLoc estimates the observer’s 2D floor position and heading within a known 3D indoor environment from a free-form description of surroundings.\"},{\"question\":\"How does LangLoc retrieve the correct scene?\",\"answer\":\"It uses a three-stage pipeline where the first stage retrieves the correct scene using a dual-branch GATv2 encoder with CLIP semantic features, improving Top-1 recall over prior work.\"},{\"question\":\"What happens when the language description is ambiguous?\",\"answer\":\"LangLoc enters an interactive dialog that asks targeted yes/no questions and updates a Bayesian pose posterior until the location is resolved.\"}]",1784183789,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"langloc-tell-me-what-you-see","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/langloc-tell-me-what-you-see/82898/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does LangLoc estimate from a natural-language description?","Question",{"text":75,"@type":76},"LangLoc estimates the observer’s 2D floor position and heading within a known 3D indoor environment from a free-form description of surroundings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LangLoc retrieve the correct scene?",{"text":80,"@type":76},"It uses a three-stage pipeline where the first stage retrieves the correct scene using a dual-branch GATv2 encoder with CLIP semantic features, improving Top-1 recall over prior work.",{"name":82,"@type":73,"acceptedAnswer":83},"What happens when the language description is ambiguous?",{"text":84,"@type":76},"LangLoc enters an interactive dialog that asks targeted yes/no questions and updates a Bayesian pose posterior until the location is resolved.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]