[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81502-en":3,"doc-seo-81502-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81502,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Zero-shot 3D General Obstacle Detection via Multimodal Foundation Models and Geometry","General obstacle detection is crucial for autonomous driving, particularly in long-tail settings with rare or previously unseen objects. Supervised and category-dependent methods limit generalization and make scalable development difficult because new obstacle types require retraining. This work introduces a training-free zero-shot pipeline that combines multimodal foundation models with geometric reasoning to detect obstacles as deviations from the road surface, segmented in 2D and localized in 3D via temporal LiDAR aggregation. Experiments report accurate localization up to 100 m, 10–25% recall gains from foundation priors, and scalable autolabeling.","Zero-shot 3D General Obstacle Detection via Multimodal Foundation Models  \nand Geometry  \nTams Matuszka aiMotive  \n[tamas.matuszka@aimotive.com](tamas.matuszka@aimotive.com)  \nPter HajasaiMotive  \n[peter.hajas@aimotive.com](peter.hajas@aimotive.com)  \nDvid SzeghyaiMotive  \n[david.szeghy@aimotive.com](david.szeghy@aimotive.com)  \narXiv :2408 . 12322v2 [ cs .CV] 10 Jul 2026  \nAbstract  \nDetecting general obstacles is critical for autonomous driving, especially in long-tail scenarios with rare or unseen objects. Existing methods rely on supervision or predefined categories, limiting generalization. We propose a training-free approach that combines multimodal foundation models with geometric reasoning for 3D obstacle detection. Our key idea is to detect obstacles as deviations from the road surface, segmented in 2D and localized in 3D via temporal LiDAR aggregation. The pipeline operates in a zero-shot manner without task-specific training. Experiments show accurate localization up to 100 meters and 10– 25% recall gains from foundation model priors, while enabling scalable autolabeling.  \n1. Introduction  \nNumerous 3D perception algorithms in the literature have been successfully applied to autonomous driving with reliable performance. These algorithms are based on different concepts, such as transformers [14], [4], and convolutional neural networks [15], [18], [21], and they perform well on predefined categories defined by established datasets. However, autonomous vehicles operating in real-world scenarios need to be able to recognize categories beyond these predefined classes to ensure safe and reliable deployment. Since general obstacles may fall into long-tail distribution object classes or categories that were not seen during training, traditional supervised learning methods are not suitable for detecting hazardous obstacles on the road surface. Additionally, incorporating a new obstacle category into the training set would require an expensive retraining phase, which hinders the scalable development of perception algorithms as obstacle classes are not known in advance.  \nThere are several methods available for addressing the problem of general obstacle detection, but these approaches have significant limitations. Existing approaches, such as stereo vision, depth estimation, and semantic segmentation,  \nstruggle to reliably detect small, distant, or previously unseen obstacles, especially under real-world conditions.  \nThis raises the question: can general obstacles be detected in 3D without costly training? We address this by combining foundation models with geometric reasoning. Our key insight is that obstacles can be identified as deviations from the road surface, which can be robustly segmented in image space and localized in 3D via temporal aggregation of LiDAR data.  \nFoundation models, such as Grounding DINO [16] and Segment Anything [12], enable promptable and generalpurpose perception through large-scale pretraining. We leverage these models to extract obstacle candidates by focusing on the road surface, while clustering and tracking algorithms are used to achieve consistent 3D localization over time. Unlike open-world detection methods [10], our objective is not to classify unknown objects but to detect their presence, eliminating the need for retraining or oracle labeling. As a result, the proposed pipeline operates ina training-free, zero-shot manner and generalizes to previously unseen obstacle types in real-world driving scenarios.  \nIn summary, this paper makes the following contributions:  \n• A unified approach combining foundation models and geometric reasoning for general obstacle detection.  \n• A training-free (zero-shot), offline method for 3D general obstacle detection enabling large-scale autolabeling without manual supervision.  \n2. Related work  \nOpen World Detection (OWD) aims to identify objects outside predefined categories and incrementally learn them over time [10] . Existing approaches typically ","cbCainGweUjvnlta","https://ap.wps.com/l/cbCainGweUjvnlta","pdf",3492517,2,1,6,"English","en",105,"# Abstract\n# Introduction\n# Related work\n# Hybrid 3D Obstacle Detection\n## Problem Statement\n## Multimodal Foundation Model-based General Obstacle Segmentation","[{\"question\":\"What problem does the paper address in autonomous driving?\",\"answer\":\"The paper targets 3D general obstacle detection, especially for long-tail scenarios involving rare or previously unseen road obstacles.\"},{\"question\":\"How does the proposed method avoid costly training?\",\"answer\":\"It uses a training-free zero-shot pipeline that combines promptable multimodal foundation models with geometric reasoning, operating offline without task-specific retraining.\"},{\"question\":\"What is the key idea for detecting obstacles in 3D?\",\"answer\":\"Obstacles are identified as deviations from the road surface, segmented in 2D and localized in 3D by temporal LiDAR aggregation over time.\"}]",1784173843,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"zero-shot-3d-general-obstacle-detection-via-multimodal-foundation-models-and-geometry","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/zero-shot-3d-general-obstacle-detection-via-multimodal-foundation-models-and-geometry/81502/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in autonomous driving?","Question",{"text":75,"@type":76},"The paper targets 3D general obstacle detection, especially for long-tail scenarios involving rare or previously unseen road obstacles.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method avoid costly training?",{"text":80,"@type":76},"It uses a training-free zero-shot pipeline that combines promptable multimodal foundation models with geometric reasoning, operating offline without task-specific retraining.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key idea for detecting obstacles in 3D?",{"text":84,"@type":76},"Obstacles are identified as deviations from the road surface, segmented in 2D and localized in 3D by temporal LiDAR aggregation over time.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]