[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81765-en":3,"doc-seo-81765-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81765,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","What’s Hidden Matters Identifying Planning-Critical Occluded Agents using Vision-Language Models","Autonomous vehicles must navigate environments where planning-critical agents may be occluded from sensors. Existing methods either apply uniform conservatism or infer hidden content without quantifying its effect on trajectory planning. This work bridges perception and planning by using vision-language models to identify and rank the specific hidden agents most critical to the ego vehicle’s plan. A Planning KL-divergence (PKL) metric drives information-theoretic ranking, while an expert VLM generates structured annotations and a new nuScenes benchmark. Experiments show finetuning on PKL-guided data yields major gains, including about 30% improvement over random sampling.","What’s Hidden Matters: Identifying Planning-Critical Occluded Agents  \nusing Vision-Language Models  \nAmirhosein Chahe 1 ,2 Tyler Naes 1 Jovin D’sa 1 Faizan M. Tariq 1 Sangjae Bae 1  \nLifeng Zhou2 David Isele 1  \narXiv :2607 .00283v 1 [ cs .RO] 1 Jul 2026  \nAbstract—Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle’s trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle’s plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of generalpurpose and domain-adapted VLMs, demonstrating that finetuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.  \nI. INTRODUCTION  \nSafe navigation requires that autonomous vehicles reason not only about what their sensors can see but also about what they cannot see. While recent advances in perception have enabled impressive object detection capabilities [1], [2] , a fundamental challenge remains: not all occluded regions pose equal risk to the ego vehicle’s trajectory. Current approaches either treat all occlusions conservatively, leading to overly cautious behavior that disrupts traffic flow, or attempt to generatively complete the scene by predicting the contents of occluded areas, which can produce unreliable and potentially dangerous predictions [3], [4] .  \nThe gap between perception and planning becomes particularly pronounced when dealing with occluded agents. To help bridge this gap, planner-aware metrics like Planning  \n1Honda Research Institute (HRI), San Jose, CA 95134, USA.  \n2Drexel University, Philadelphia, PA 19104, USA.  \nAll work was done while A. Chahe was employed by HRI. Contact: [ac4462@drexel.edu](ac4462@drexel.edu), [sbae@honda-ri.com](sbae@honda-ri.com), [disele@honda-ri.com](disele@honda-ri.com).  \n©2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.  \nKL-divergence (PKL) were introduced to quantify how perception errors impact downstream planning [5] . We propose leveraging this concept for a different task: to identify planning-critical hidden agents whose presence would force a significant change in the ego vehicle’s trajectory. However, effectively identifying these high-risk scenarios requires a deep, contextual understanding that goes beyond raw geometry. Vision-Language Models (VLMs) are well suited to this challenge, bringing semantic grounding and reasoning beyond geometry [6]–[8] . Recen","cbCaijNOEFUVLyig","https://ap.wps.com/l/cbCaijNOEFUVLyig","pdf",16950173,3,1,9,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"How is the dataset and annotation benchmark created, and what is its reported effect?\",\"answer\":\"The approach uses an expert VLM (GPT-5) to generate structured annotations capturing visual evidence and reasoning, then fine-tunes and benchmarks diverse VLMs on the new nuScenes-focused benchmark, improving performance by around 30% versus random sampling and showing strong results for smaller fine-tuned models.\"}]",1784176012,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"whats-hidden-matters-identifying-planning-critical-occluded-agents-using-vision-language-models","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/whats-hidden-matters-identifying-planning-critical-occluded-agents-using-vision-language-models/81765/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is the dataset and annotation benchmark created, and what is its reported effect?","Question",{"text":75,"@type":76},"The approach uses an expert VLM (GPT-5) to generate structured annotations capturing visual evidence and reasoning, then fine-tunes and benchmarks diverse VLMs on the new nuScenes-focused benchmark, improving performance by around 30% versus random sampling and showing strong results for smaller fine-tuned models.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]