[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85533-en":3,"doc-seo-85533-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85533,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","VehAnchor Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery","Autonomous aerial robots in GPS-denied or communication-degraded environments often lose camera metadata and absolute telemetry, preventing onboard perception from recovering the scene’s absolute metric scale. This limits the safety of LLM/VLM-based embodied planners because experiments show major spatial scale hallucinations, including median area estimation errors over 50%. VehAnchor is a lightweight deterministic geometric perception tool that estimates Ground Sample Distance (GSD) from detected small vehicles, returns a confidence score, and supports agent fallback decisions. On DOTA v1.5 it attains 6.87% median GSD error and, with SAM segmentation, 19.7% median area error with fewer catastrophic failures than VLM baselines.","VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery  \nYifei Chen 1 , Chenqian Le2 , Jiayi Cheng2 , and Xupeng Chen 1 ,∗  \narXiv :2603 .04277v2 [ cs .RO] 10 Jul 2026  \nAbstract—Autonomous aerial robots operating in GPSdenied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene. As LLM/VLM-based planners are increasingly adopted as high-level agents for embodied systems, their ability to reason about physical dimensions becomes safety-critical—yet our experiments show that five state-of-the-art VLMs suffer from spatial scale hallucinations, with median area estimation errors exceeding 50% . We propose VehAnchor, a lightweight, deterministic Geometric Perception Skill designed as a callable tool that any LLM-based agent can invoke to recover Ground Sample Distance (GSD) from ubiquitous environmental anchors: small vehicles detected via oriented bounding boxes, whose modal pixel length is robustly estimated through kernel density estimation and converted to GSD using a pre-calibrated reference length. The tool returns both a GSD estimate anda composite confidence score, enabling the calling agent to autonomously decide whether to trust the measurement or fallback to alternative strategies. On the DOTA v1.5 benchmark, VehAnchor achieves 6.87% median GSD error on 306 images. Integrated with SAM-based segmentation for downstream area measurement, the pipeline yields 19.7% median error on a 100-entry benchmark—with 2.6× lower category dependence and 4× fewer catastrophic failures than the best VLM baseline—demonstrating that equipping agents with deterministic geometric tools is essential for safe autonomous spatial reasoning.  \nI. INTRODUCTION  \nMicro aerial vehicles (MAVs) and unmanned aerial vehicles (UAVs) are increasingly deployed for autonomous tasks—disaster assessment, infrastructure inspection, precision agriculture, and search-and-rescue—where onboard perception must deliver metric-scale spatial understanding in real time. In GPS-denied, communication-degraded, or previously unmapped environments, these platforms routinely lose access to camera metadata and absolute telemetry, leaving monocular imagery as the sole input. Without knowing the Ground Sample Distance (GSD)—the physical size of each pixel—pixel-level measurements cannot be converted to realworld dimensions, and any downstream spatial reasoning becomes unreliable.  \nThe robotics community has increasingly turned to visionlanguage models (VLMs) and large language models (LLMs) as high-level planners and agents for embodied systems [1] . However, our experiments reveal a critical failure mode:  \n1Yifei Chen and Xupeng Chen are with Dimension Gate, Haidian District, Beijing 100083, China.  \n2Chenqian Le and Jiayi Cheng are with New York University, New York, NY, USA.  \n∗ Corresponding author: Xupeng Chen.  \nwhen tasked with estimating physical areas from aerial imagery alone, five state-of-the-art VLMs exhibit median errors of 38–52%, with frequent order-of-magnitude deviations. We term this failure Spatial Scale Hallucination—the systematic inability of VLMs to ground visual observations in metric scale without explicit geometric calibration. This poses a direct safety risk for autonomous UAV operations: a planner that misjudges a landing zone’s dimensions by 50% could attempt a catastrophic landing on an insufficiently sized surface.  \nWe propose VehAnchor—a lightweight, deterministic Geometric Perception Skill and callable tool that LLM-based planners can invoke to recover absolute scale from monocular aerial imagery. Our key observation is that small vehicles are among the most ubiquitous objects in urban and suburban scenes, with physical lengths concentrated around 4– 5 m worldwide. The pipeline detects vehicles via oriented bounding boxes (OBB), robustly estimates their modal pixel length through ","cbCairjBPVIQlDYJ","https://ap.wps.com/l/cbCairjBPVIQlDYJ","pdf",5362461,1,7,"English","en",105,"# Introduction\n## Problem: spatial scale loss in metadata-free aerial perception\n## VehAnchor approach and contributions\n# Related Work\n## GSD and scale estimation","[{\"question\":\"What problem does VehAnchor address in metadata-free aerial imagery?\",\"answer\":\"VehAnchor addresses the inability to recover absolute metric scale when camera metadata and telemetry are unavailable, making pixel measurements unusable for real-world spatial reasoning.\"},{\"question\":\"How does VehAnchor estimate Ground Sample Distance (GSD) from vehicle cues?\",\"answer\":\"It detects small vehicles using oriented bounding boxes, estimates their modal pixel length with kernel density estimation, and converts that length to GSD using a pre-calibrated reference length.\"},{\"question\":\"What evidence shows VLMs struggle with metric scale, and how does VehAnchor compare?\",\"answer\":\"The paper reports median physical area errors of 38–52% for five state-of-the-art VLMs and frequent order-of-magnitude deviations, while VehAnchor achieves 6.87% median GSD error on DOTA v1.5 and 19.7% median area error when integrated with SAM segmentation.\"}]",1784204244,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"vehanchor-metadata-free-metric-scale-recovery-from-vehicle-cues-in-aerial-imagery","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vehanchor-metadata-free-metric-scale-recovery-from-vehicle-cues-in-aerial-imagery/85533/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does VehAnchor address in metadata-free aerial imagery?","Question",{"text":75,"@type":76},"VehAnchor addresses the inability to recover absolute metric scale when camera metadata and telemetry are unavailable, making pixel measurements unusable for real-world spatial reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VehAnchor estimate Ground Sample Distance (GSD) from vehicle cues?",{"text":80,"@type":76},"It detects small vehicles using oriented bounding boxes, estimates their modal pixel length with kernel density estimation, and converts that length to GSD using a pre-calibrated reference length.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows VLMs struggle with metric scale, and how does VehAnchor compare?",{"text":84,"@type":76},"The paper reports median physical area errors of 38–52% for five state-of-the-art VLMs and frequent order-of-magnitude deviations, while VehAnchor achieves 6.87% median GSD error on DOTA v1.5 and 19.7% median area error when integrated with SAM segmentation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]