[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82100-en":3,"doc-seo-82100-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82100,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images","Figure-ground organization in the human visual system depends on shape-based cues such as surroundedness, convexity, and symmetry. While these cues are often studied using artificial stimuli, their behavior in natural scenes and how they arise from natural image statistics remain unclear. This study evaluates shape-based figure-ground organization in Vision Transformers by training linear probes on intermediate patch representations across 25 supervised and self-supervised models, using both natural and cue-isolated stimuli.","arXiv :2607 .08932v 1 [ cs .CV] 9 Jul 2026  \nVision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images  \nMatthias Tangemann 1 ,2 Benjamin Lo3 ,4 Zygmunt Pizlo5 Kaleem Siddiqi3 ,4  \nDirk B. Walther 1 Sven Dickinson 1 ,2  \n1University of Toronto 2Vector Institute 3McGill University 4MILA 5UC Irvine  \n[mtangemann@cs.toronto.edu](mtangemann@cs.toronto.edu)  \nAbstract  \nFigure-ground organization in the human visual system relies on several shapebased cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTsas a compelling model system for studying the computational mechanisms of perceptual organization.  \nCode and data is available at [https://github.com/mtangemann/mlvbench](https://github.com/mtangemann/mlvbench).  \n1 Introduction  \nFigure-ground organization is one of the most fundamental processes in human visual perception. Research over the past century has identified several shape-based cues that drive this process, including surroundedness, convexity, and symmetry [52] . These cues have been extensively characterized using controlled, artificial stimuli (e.g., [41]), yet we lack a precise understanding of how they contribute to the perceptual organization of natural scenes. It has been hypothesized that many visual cues are rooted in the statistics of natural scenes [26, 21], but the mechanisms that enable learning general cues from experience remain unknown.  \nComputational models that rely on the same figure-ground cues as humans could offer tractable experimental access to these open questions. Such models could serve a role analogous to model organisms in neuroscience: they allow for controlled experimentation that is difficult or impossible inhuman observers, while generating hypotheses that can subsequently be tested in humans. Vision Transformers (ViTs) are strong candidates for such model systems. Their self-attention mechanism enables global processing across the image, and several studies have demonstrated that structured  \nPreprint.  \nFigure 1: We extract patch representations from frozen, pre-trained ViTs and fit linear probes to predict figure-ground assignment. We evaluate probes on both natural images and controlled stimuli that isolate individual shape cues while semantic information, texture, and region size are uninformative. We exclude patches that span both foreground and background (shown in black) .  \nscene representations emerge in their intermediate layers [8, 36, 1] . These studies have established that ViTs develop rich segmentation capabilities, but the question of what cues underlie these capabilities remains open. In particular, it is unclear whether ViTs rely on generalizable, shape-based cues as humans do, or whether they primarily rely on sema","cbCaimfKtjvQcjGH","https://ap.wps.com/l/cbCaimfKtjvQcjGH","pdf",5907280,1,26,"English","en",105,"# Abstract\n# Introduction\n## Figure-ground organization and prior work\n## Vision Transformers as model systems\n# Method and evaluation approach\n## Linear probes on patch representations\n## Natural vs cue-isolated synthetic stimuli\n# Results\n## Surroundedness and convexity encoding\n## Symmetry and texture dependence\n## Layer-wise analysis across training objectives\n# Contributions","[{\"question\":\"What figure-ground cues are investigated in this work?\",\"answer\":\"The study focuses on shape-based cues: surroundedness, convexity, and symmetry, examining how they influence figure-ground organization in natural images and controlled stimuli.\"},{\"question\":\"How do the authors test whether ViTs encode these cues?\",\"answer\":\"They extract intermediate patch representations from frozen, pre-trained ViTs and train linear probes to predict whether each patch belongs to foreground or background.\"},{\"question\":\"What do the results show about different cues and generalization?\",\"answer\":\"Vision Transformers robustly encode surroundedness and convexity, and probes trained on natural images often generalize zero-shot to artificial stimuli that isolate these cues; symmetry results are mixed and depend on texture.\"}]",1784178212,66,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"vision-transformers-learn-gestalt-like-figure-ground-cues-from-natural-images","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vision-transformers-learn-gestalt-like-figure-ground-cues-from-natural-images/82100/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What figure-ground cues are investigated in this work?","Question",{"text":75,"@type":76},"The study focuses on shape-based cues: surroundedness, convexity, and symmetry, examining how they influence figure-ground organization in natural images and controlled stimuli.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors test whether ViTs encode these cues?",{"text":80,"@type":76},"They extract intermediate patch representations from frozen, pre-trained ViTs and train linear probes to predict whether each patch belongs to foreground or background.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about different cues and generalization?",{"text":84,"@type":76},"Vision Transformers robustly encode surroundedness and convexity, and probes trained on natural images often generalize zero-shot to artificial stimuli that isolate these cues; symmetry results are mixed and depend on texture.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]