[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86393-en":3,"doc-seo-86393-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86393,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Back to Point Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection","Zero-shot (ZS) 3D anomaly detection is pivotal for industrial inspection, enabling defect detection and localization without any target-category anomaly training data. Prior work typically projects 3D point clouds to 2D images and applies pre-trained Vision-Language Models, which discards geometric details and weakens sensitivity to local structural defects. This paper revisits intrinsic 3D representations and proposes BTP, aligning point-cloud and textual embeddings via multi-granularity patch features and geometry-aware descriptors, with joint learning using auxiliary point clouds. Experiments on Real3D-AD and Anomaly-ShapeNet validate superior ZS performance.","Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly  \nDetection  \nKaiqiang Li 1 ,2 Gang Li 1 ,2 * Mingle Zhou 1 ,2 Min Li 1 ,2 Delong Han 1 ,2 Jin Wan 1 ,2 *  \n1 Key Laboratory of Computing Power Network and Information Security,  \nMinistry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China.  \n2 Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing,  \nShandong Fundamental Research Center for Computer Science, Jinan, China.  \n[b1043125006@stu.qlu.edu.cn](b1043125006@stu.qlu.edu.cn) , {lig, zhouml, limin, handl, [wanj](wanj}@qlu.edu.cn)[}](wanj}@qlu.edu.cn)[@qlu.edu.cn](wanj}@qlu.edu.cn)  \narXiv :2603 .21511v3 [ cs .CV] 12 Jul 2026  \nAbstract  \nZero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target-category training data. Existing approaches render 3D point clouds into 2D images and leverage pre-trained Vision-Language Models (VLMs) for anomaly detection. However, such strategies inevitably discard geometric details and exhibit limited sensitivity to local anomalies. In this paper, we revisit intrinsic 3D representations and explore the potential of pre-trained Point-Language Models (PLMs) for ZS 3D anomaly detection. We propose BTP (Back To Point), a novel framework that effectively aligns 3D point cloud and textual embeddings. Specifically, BTP aligns multi-granularity patch features with textual representations for localized anomaly detection, while incorporating geometric descriptors to enhance sensitivity to structural anomalies. Furthermore, we introduce a joint representation learning strategy that leverages auxiliary point cloud data to improve robustness and enrich anomaly semantics. Extensive experiments on Real3D-AD and Anomaly-ShapeNet demonstrate that BTP achieves superior performance in ZS 3D anomaly detection. Code will be available at [https://github.com/wistful-](https://github.com/wistful-)[ ](https://github.com/wistful-)8029/BTP-3DAD.  \n1. Introduction  \n3D anomaly detection is pivotal to industrial quality inspection, as it safeguards the safety, functionality, and reliability of products [2, 7, 10, 15, 16, 21, 38, 42, 47, 48, 55] . Existing unsupervised methods [2, 9, 15, 21, 22, 55] mainly rely on memory-bank retrieval or reconstruction-based paradigms,  \n*Co-corresponding authors.  \n(a) Existing VLM-based approaches  \n(b) Our proposed PLM-based approach  \nFigure 1 . Comparison between VLM-based and PLM-based (ours) zero-shot 3D anomaly detection. (a): VLM-based approaches rely on multi-view rendering and back-projection, making their anomaly localization performance sensitive to the number and angles of rendered views. (b): Our proposed PLM-based approach directly processes point clouds, avoiding such view-dependent limitations and achieving more accurate 3D anomaly localization.  \nyet they struggle to generalize to unseen anomaly types. Moreover, practical deployment is often impeded by privacy restrictions and data acquisition limitations, which make it challenging to collect sufficiently diverse training samples in real-world scenarios [1, 19, 25, 26, 39, 50, 54, 57] . Consequently, zero-shot (ZS) 3D anomaly detection has emerged as a promising paradigm, enabling defect de-  \ntection and localization without anomalous samples, and effectively addressing the challenges of data scarcity and unseen anomaly generalization [12, 28, 43, 53] .  \nRecent advances in large-scale Vision-Language Models (VLMs) have significantly advanced ZS learning, resulting in remarkable progress in 2D anomaly detection. Mainstream approaches [3, 4, 6, 17, 27–29, 32, 44, 52] leverage pre-trained VLMs (e.g., CLIP [33]) to align textual descriptions of normal and abnormal conditions with image features, thereby enabling anomaly detection without requiring any anomalou","cbCaitqp866xBD8M","https://ap.wps.com/l/cbCaitqp866xBD8M","pdf",1456377,2,1,11,"English","en",105,"# Introduction\n## Background: VLM-based zero-shot 3D anomaly detection\n## Motivation: limitations of 2D projection and view dependence\n## Point-Language Models and the proposed BTP approach","[{\"question\":\"What is the goal of zero-shot (ZS) 3D anomaly detection in this work?\",\"answer\":\"It detects and localizes defects without requiring any anomalous samples from the target category for training, addressing data scarcity and unseen anomaly generalization.\"},{\"question\":\"Why do VLM-based methods underperform for fine-grained 3D anomalies?\",\"answer\":\"They render point clouds into 2D images, which discards rich 3D geometric cues and leads to limited sensitivity to local structural anomalies, besides being sensitive to view selection.\"},{\"question\":\"How does BTP improve zero-shot 3D anomaly detection compared with prior approaches?\",\"answer\":\"BTP directly encodes 3D point clouds and aligns them with text using multi-granularity patch features and geometry-aware descriptors, plus joint representation learning with auxiliary point cloud data for robustness and enriched anomaly semantics.\"}]",1784211468,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"back-to-point-exploring-point-language-models-for-zero-shot-3d-anomaly-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/back-to-point-exploring-point-language-models-for-zero-shot-3d-anomaly-detection/86393/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the goal of zero-shot (ZS) 3D anomaly detection in this work?","Question",{"text":75,"@type":76},"It detects and localizes defects without requiring any anomalous samples from the target category for training, addressing data scarcity and unseen anomaly generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do VLM-based methods underperform for fine-grained 3D anomalies?",{"text":80,"@type":76},"They render point clouds into 2D images, which discards rich 3D geometric cues and leads to limited sensitivity to local structural anomalies, besides being sensitive to view selection.",{"name":82,"@type":73,"acceptedAnswer":83},"How does BTP improve zero-shot 3D anomaly detection compared with prior approaches?",{"text":84,"@type":76},"BTP directly encodes 3D point clouds and aligns them with text using multi-granularity patch features and geometry-aware descriptors, plus joint representation learning with auxiliary point cloud data for robustness and enriched anomaly semantics.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]