[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86025-en":3,"doc-seo-86025-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86025,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","TOLiD Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation","Cross-modal distillation from vision foundation models (VFMs) to LiDAR backbones has emerged as a self-supervised pretraining strategy that reduces dependence on dense point-wise annotation for 3D scene understanding. Existing pipelines often freeze the VFM and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to handle both modality and cross-architecture gaps. TOLiD addresses these gaps by coupling a LiDAR backbone with a student ViT teacher, using token-level distillation with visibility masking and token lifting to per-point representations. Evaluations on multiple datasets show improved transfer using frozen backbones and lightweight heads.","TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation  \nSutharsan Mahendran 1 ,2 , Darshana Priyasad 1 , Kaushik Roy2 , Tharindu Fernando 1 , Sridha Sridharan 1 , Clinton Fookes 1 , Peyman Moghadam 1 ,2  \narXiv :2607 . 10762v1 [ cs .CV] 12 Jul 2026  \nAbstract—Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged asa self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding. However, existing distillation pipelines typically treat the VFM as a frozen feature source and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. We propose TOLiD, a self-supervised pretraining method for LiDAR representation learning that addresses this gap by coupling a LiDAR backbone with a student Vision Transformer (ViT) initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations. TOLiD converts the set of point features within each image patch frustum into a token using Frustum Pooling followed by Frustum Attention, and performs token-level distillation with visibility masking. For LiDAR-only deployment, we lift token features back to per-point representations using masked bilinear sampling to avoid patches that have limited LiDAR points. We extensively evaluate TOLiD on five heterogeneous LiDAR datasets and four cross-sensor adaptation pairs, demonstrating improved transfer with frozen backbones and lightweight heads.  \nI. INTRODUCTION  \n3D semantic segmentation is a core capability for autonomous robots and self-driving cars, enabling 3D scene understanding for navigation, mapping, and safe interaction [1]–[5] . In many robotics systems, 3D scene understanding relies primarily on LiDAR, as it provides metric geometric measurements that are largely invariant to illumination and support reliable long-range perception.  \nLiDAR-based learning models must operate across heterogeneous sensor suites that differ in beam count, field of view, point density, and scanning pattern, and these variations induce cross-sensor domain shifts when models are transferred between platforms. Addressing such sensor gaps by collecting new annotations is often impractical, since dense point-wise labelling of LiDAR scans is labourintensive and costly [1] .  \nWhile 2D perception tasks have benefited substantially from Vision Foundation Models (VFMs) pre-trained on large-scale datasets [6], [7], progress for LiDAR-based tasks has been comparatively constrained by the absence of suitable 3D foundation models and by the modality gap that  \n1 SAIVT Group, School of Electrical Engineering and Robotics, Queensland University of Technology, Australia. E-mails: {s2 .mahendren, s.sridharan, c.fookes, [peyman.moghadam}@qut.edu.au](peyman.moghadam}@qut.edu.au)  \n2CSIRO Robotics, CSIRO, Australia. E-mails:{sutharsan .mahendran, kaushik .roy, [peyman.moghadam}@csiro.au](peyman.moghadam}@csiro.au)  \narises when utilizing VFMs for LiDAR representations [8],[9] . To address these limitations, recent work has leveraged VFMs as teachers and distilled their representations directly to LiDAR backbones as a self-supervised pretraining step, followed by a finetuning stage that requires substantially less labeled 3D data. These VFM-to-LiDAR cross-modal distillation methods typically follow a two-stage pipeline. In the pretraining stage, synchronized camera–LiDAR pairs are used to establish 2D–3D correspondences: LiDAR points [10],[11](or superpoints [12], [13]) are projected into the image plane and associated with dense VFM features (patch tokens or intermediate feature embeddings) . Then LiDAR encoder is trained to predict or align to these VFM embeddings, often through feature regression [11], [14], or contrastive obje","cbCaisTIotbx9Lqy","https://ap.wps.com/l/cbCaisTIotbx9Lqy","pdf",2398623,4,1,"English","en",105,"# Introduction\n## Cross-modal distillation from VFMs to LiDAR\n## Architectural gap and limitations\n## Proposed method: TOLiD\n## Token representation via Frustum Pooling and Frustum Attention","[{\"question\":\"What problem does TOLiD aim to solve in VFM-to-LiDAR distillation?\",\"answer\":\"TOLiD targets the fundamental limitation where current pipelines must bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. This mismatch can cause information loss and unstable transfer.\"},{\"question\":\"How does TOLiD convert LiDAR points within an image patch frustum into tokens?\",\"answer\":\"TOLiD represents each image patch by its corresponding 3D frustum and converts the set of point features in that frustum into a single token. It uses Frustum Pooling for stable aggregation and Frustum Attention to emphasize the most informative points.\"},{\"question\":\"How are token features lifted back for LiDAR-only deployment?\",\"answer\":\"For LiDAR-only use, TOLiD lifts token features back to per-point representations using masked bilinear sampling. This masking avoids patches with limited LiDAR points.\"}]",1784207884,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"tolid-bridging-the-architecture-gap-in-vision-foundation-model-to-lidar-pretraining-via-token-lifting-for-distillation","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/tolid-bridging-the-architecture-gap-in-vision-foundation-model-to-lidar-pretraining-via-token-lifting-for-distillation/86025/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does TOLiD aim to solve in VFM-to-LiDAR distillation?","Question",{"text":74,"@type":75},"TOLiD targets the fundamental limitation where current pipelines must bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. This mismatch can cause information loss and unstable transfer.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does TOLiD convert LiDAR points within an image patch frustum into tokens?",{"text":79,"@type":75},"TOLiD represents each image patch by its corresponding 3D frustum and converts the set of point features in that frustum into a single token. It uses Frustum Pooling for stable aggregation and Frustum Attention to emphasize the most informative points.",{"name":81,"@type":72,"acceptedAnswer":82},"How are token features lifted back for LiDAR-only deployment?",{"text":83,"@type":75},"For LiDAR-only use, TOLiD lifts token features back to per-point representations using masked bilinear sampling. This masking avoids patches with limited LiDAR points.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]