[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83931-en":3,"doc-seo-83931-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83931,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Vision Pretraining for Dense Spatial Perception","Dense spatial perception is crucial for physical intelligence, yet modern visual foundation models often emphasize semantic invariance while underperforming on fine-grained spatial understanding. This work proposes a boundary-centric self-supervised pretraining approach: masked boundary modeling that learns sub-pixel boundary representations and uses boundary-bearing tokens as masked targets for dense visual token learning. Scaling the method yields LingBot-Vision, improving diverse downstream tasks with DINOv3 as baseline and advancing LingBot-Depth for depth completion.","arXiv :2607 .05247v 1 [ cs .CV] 6 Jul 2026  \nVision Pretraining for Dense Spatial Perception  \nZelin Fu∗ Bin Tan∗ Changjiang Sun Shaohui Liu Kecheng Zheng Yinghao Xu Xing Zhu Yujun Shen Nan Xue†  \n∗ Equal contributions †Project Lead  \nDense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representationsand subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.  \nWebsite: [https://technology.robbyant.com/lingbot-vision](https://technology.robbyant.com/lingbot-vision)  \nGithub: [https://github.com/robbyant/lingbot-vision](https://github.com/robbyant/lingbot-vision)  \nCheckpoints: [https://huggingface.co/collections/robbyant/lingbot-vision](https://huggingface.co/collections/robbyant/lingbot-vision)  \n1 Introduction  \nVisual intelligence is not only about recognizing what is in an image, but also about recovering the dense spatial structure that makes the physical world measurable, navigable, and actionable. Across segmentation, depth, motion, and scene layout, this structure is repeatedly revealed by shapes and boundaries: object masks delineate entities, depth discontinuities expose geometry, motion silhouettes separate dynamic regions, and occlusion contours organize visible surfaces. Yet many modern visual foundation models are driven by objectives that favor semantic invariance, cross-modal alignment, or appearance reconstruction, and remain comparatively weak at the fine-grained spatial understanding that downstream physical intelligence depends on.  \nA central reason is how shapes and boundaries are treated. They are usually regarded as outputs of perception, recovered by task-specific heads under dedicated supervision. Such an output-centric view ties them to annotations that are expensive, ambiguous, and often unavailable, making them difficult to exploit in large-scale pretraining. As a result, foundation models rarely use shapes and boundaries as native learning signals, relying instead on view-invariant self-distillation [9, 28, 39], cross-modal alignment [32, 61], or masked image reconstruction [19] . We argue that this is a missed opportunity: boundaries and shape discontinuities are not merely outputs to be predicted, but fundamental organizing signals for learning dense representations.  \nIn this work, we therefore study how randomly initialized Vision Transformers can learn shape-and boundary-aware representations from raw images alone, without human annotations, external edge detectors, or pretrained backbones. Our study builds on dense, over-parameterized field representations of boundary geometry [58, 59], which encode vectorized boundaries through pixel-wise, distance-transform-like fields. Such representations have been extensively used in supervised line-segment and wireframe detection, where they turn sparse geometric st","cbCaipUHLXAy2AD6","https://ap.wps.com/l/cbCaipUHLXAy2AD6","pdf",12324500,3,1,31,"English","en",105,"# Introduction\n## Boundary-centric view of spatial structure\n## Shape and boundary signals in pretraining\n## Boundary-Forcing Masked Modeling","[{\"question\":\"Why do modern vision foundation models often struggle with dense spatial understanding?\",\"answer\":\"They are frequently trained with objectives that promote semantic invariance, cross-modal alignment, or appearance reconstruction, while treating shapes and boundaries as outputs predicted by task-specific heads under supervision.\"},{\"question\":\"What is masked boundary modeling in this work?\",\"answer\":\"It is a self-supervised paradigm where a model learns sub-pixel boundary representations and then uses the discovered boundary-bearing tokens as masked targets to train dense visual token learning.\"},{\"question\":\"How does LingBot-Vision improve downstream tasks such as depth completion?\",\"answer\":\"By scaling the framework, LingBot-Vision demonstrates strong transfer across diverse vision tasks with DINOv3 as a baseline, and it progresses from LingBot-Depth to LingBot-Depth 2.0 for depth completion, leading to enhanced depth estimation.\"}]",1784191519,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vision-pretraining-for-dense-spatial-perception","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/vision-pretraining-for-dense-spatial-perception/83931/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do modern vision foundation models often struggle with dense spatial understanding?","Question",{"text":75,"@type":76},"They are frequently trained with objectives that promote semantic invariance, cross-modal alignment, or appearance reconstruction, while treating shapes and boundaries as outputs predicted by task-specific heads under supervision.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is masked boundary modeling in this work?",{"text":80,"@type":76},"It is a self-supervised paradigm where a model learns sub-pixel boundary representations and then uses the discovered boundary-bearing tokens as masked targets to train dense visual token learning.",{"name":82,"@type":73,"acceptedAnswer":83},"How does LingBot-Vision improve downstream tasks such as depth completion?",{"text":84,"@type":76},"By scaling the framework, LingBot-Vision demonstrates strong transfer across diverse vision tasks with DINOv3 as a baseline, and it progresses from LingBot-Depth to LingBot-Depth 2.0 for depth completion, leading to enhanced depth estimation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]