[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85078-en":3,"doc-seo-85078-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85078,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Hierarchical Slot Attention for Multi-granularity Scene Decomposition","Slot attention decomposes visual scenes into latent slots for object-centric learning, but existing approaches typically produce flat slot sets at one granularity and rely on appearance rather than semantics. Humans, however, interpret scenes through semantic hierarchies—foreground/background, object categories, and individual instances. Hierarchical semantic structure cannot arise from appearance alone. Hierarchical Slot Attention (HSA) learns three levels jointly from a single model using only 10% labeled data and hierarchical alignment loss, with improved ARI on COCO and Pascal VOC.","HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition  \nNeelu Madan 1 ,2 , Rongzhen Zhao3 , Andreas Mogelmose 1 ,2 , Juho Kannala3 Joni Pajarinen3 Graham W. Taylor4 ,5 , Thomas B. Moeslund 1 ,2  \n1Aalborg University, Denmark 2Pioneer Centre for AI, Denmark, 3Aalto University, Finland,  \n4University of Guelph, Canada, 5Vector Institute, Canada  \narXiv :2607 .08249v 1 [ cs .CV] 9 Jul 2026  \nAbstract  \nSlot attention is a powerful framework for objectcentric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods share two critical limitations: they decompose scenes into a flat set of slots at a single granularity, and this decomposition is based on appearance rather than semantics. Yet humans understand scenes through semantic hierarchies: separating foreground from background, recognizing object categories, and identifying individual instances. Crucially, such semantic hierarchies cannot emerge without supervision, because category names are human constructs, not visual patterns. We propose Hierarchical Slot Attention (HSA), which learns multi-granularity semantic scene decomposition from a single model. HSA decomposes scenes at three levels: holistic (foreground/background), semantic (object categories), and panoptic (individual instances) . Using only 10% labeled data, combined with hierarchical alignment loss, HSA learns all three levels jointly. We further introduce grouping purity and containment to measure whether the hierarchy is encoded in representation space, not just output masks. Experiments on COCO and PASCAL VOC demonstrate that HSA outperforms the strongest flat baseline by up to +41.5 ARI at holistic, +14.6 at semantic, and +10.4 at panoptic level on COCO, with even larger gains on Pascal VOC, while requiring a single model instead of three. Code will be made available upon acceptance.  \n1. Introduction  \nObject-centric learning [3, 9, 14, 15, 31] seeks to decompose visual scenes into meaningful entity representations [43] . Among recent approaches, slot attention [31, 39] has demonstrated remarkable success by treating objects as latent “slots” that compete to explain image features through iterative attention. Owing to its simplicity, slot attention has become a prominent approach for unsupervised object discovery [15] and scene decomposition [10] . Recent  \nextensions [16, 21, 33, 39] leverage self-supervised visual features such as DINOv2 [34], which encode rich patchlevel representations where spatial regions with similar appearance cluster naturally in feature space [4, 32], substantially improving slot attention’s applicability to real-world scenes.  \nYet humans understand scenes not only through appearance but also through semantic meaning: a region is not just a set of pixels, it is a car, a person, orforeground. This semantic understanding is inherently hierarchical and cannot emerge from appearance alone. Humans naturally organize visual information at multiple levels of abstraction: distinguishing background from foreground (holistic level), recognizing object categories (semantic level), and identifying individual instances within those categories (panoptic level) [43] . Symbolic cognition theory [20, 46] and concept grounding methods [36] root this connection between perception and naming, pointing towards a minimal grounding signal as the mechanism for understanding semantics. However, existing slot-based methods decompose scenes at only a single granularity, and this decomposition is spatial, finding regions and parts through appearance cues alone. We refer to these methods as flat baselines in later sections. Although some works attempt hierarchical scene decomposition [17, 25, 27], all existing unsupervised approaches [17, 25] still decompose scenes through spatially-guided appearance. We argue that semantic hierarchy cannot emerge without supervision, because category names are human constructs, not visual pa","cbCaijniVU2LDv24","https://ap.wps.com/l/cbCaijniVU2LDv24","pdf",43850592,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What limitations do existing slot attention methods have for scene decomposition?\",\"answer\":\"They usually decompose scenes into a flat set of slots at a single granularity and ground decomposition primarily in appearance, not semantic meaning.\"},{\"question\":\"How does HSA achieve multi-granularity semantic scene decomposition?\",\"answer\":\"HSA learns three decomposition levels—holistic (foreground/background), semantic (object categories), and panoptic (individual instances)—using three parallel slot attention modules on shared features plus minimal supervision and hierarchical alignment.\"},{\"question\":\"What performance improvements does HSA report and on which datasets?\",\"answer\":\"Experiments on COCO and PASCAL VOC show HSA outperforms the strongest flat baseline with ARI gains up to +41.5 (holistic), +14.6 (semantic), and +10.4 (panoptic) on COCO, along with larger gains on Pascal VOC while using a single model.\"}]",1784200907,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"hierarchical-slot-attention-for-multi-granularity-scene-decomposition","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hierarchical-slot-attention-for-multi-granularity-scene-decomposition/85078/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What limitations do existing slot attention methods have for scene decomposition?","Question",{"text":74,"@type":75},"They usually decompose scenes into a flat set of slots at a single granularity and ground decomposition primarily in appearance, not semantic meaning.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does HSA achieve multi-granularity semantic scene decomposition?",{"text":79,"@type":75},"HSA learns three decomposition levels—holistic (foreground/background), semantic (object categories), and panoptic (individual instances)—using three parallel slot attention modules on shared features plus minimal supervision and hierarchical alignment.",{"name":81,"@type":72,"acceptedAnswer":82},"What performance improvements does HSA report and on which datasets?",{"text":83,"@type":75},"Experiments on COCO and PASCAL VOC show HSA outperforms the strongest flat baseline with ARI gains up to +41.5 (holistic), +14.6 (semantic), and +10.4 (panoptic) on COCO, along with larger gains on Pascal VOC while using a single model.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]