[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86168-en":3,"doc-seo-86168-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86168,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Slot RAE Streamlining Object-Centric Learning via Direct Representation Auto-Encoders","Slot-RAE addresses the challenge of deploying object-centric models for scene understanding while avoiding complex, pipeline-heavy designs for decomposition and high-fidelity generation. It introduces a unified framework that performs feature-space diffusion directly inside a frozen visual foundation model’s continuous semantic representation (e.g., DINOv3). A DiT decoder and a Representation Alignment (REPA) head train the generative core from scratch, removing VAE bottlenecks and reliance on large text-to-image priors. Experiments on COCO show state-of-the-art object discovery, reconstruction fidelity, and robust zero-shot compositionality with improved speed and efficiency.","Slot-RAE: Streamlining Object-Centric Learning via Direct Representation Auto-Encoders  \nAlexandre Chapin Ecole Centrale de Lyon, LIRIS 69130, Ecully, France  \n[alexandre.chapin@ec-lyon.fr](alexandre.chapin@ec-lyon.fr)  \nEmmanuel Dellandrea Ecole Centrale de Lyon, LIRIS 69130, Ecully, France  \n[emmanuel.dellandrea@ec-lyon.fr](emmanuel.dellandrea@ec-lyon.fr)  \nLiming Chen  \nEcole Centrale de Lyon, LIRIS 69130, Ecully, France  \n[liming.chen@ec-lyon.fr](liming.chen@ec-lyon.fr)  \narXiv :2607 . 11196v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nDeploying object-centric models for real-world scene understanding typically requires complex pipelines to achieve both robust scene decomposition and high-fidelity generation. Recent diffusion-based approaches have improved visual quality, but they almost universally rely on heavy, pretrained generative priors (e.g., Stable Diffusion) and external VAE latent spaces. In this paper, we propose SlotRAE, a much simpler, fully integrated framework that operates directly within the continuous semantic feature space of visual foundation models (e.g., DINOv3). Slot-RAE employs a feature-space diffusion process using a Diffusion Transformer (DiT) decoder and a Representation Alignment (REPA) head. Unlike existing diffusion-based objectcentric methods that rely heavily on subsidized text-toimage priors, the generative core of Slot-RAE (Slot Attention and the DiT) is trained from scratch within the frozen VFM feature space. This eliminates the need for VAE bottlenecks and task-agnostic generative pre-training. Experiments on the COCO dataset demonstrate that despite its architectural simplicity, Slot-RAE achieves state-of-the-art results. It delivers comparable unsupervised object discovery, higher-fidelity image reconstruction, and robust zero-shot compositionality, all while being significantly faster and more computationally efficient than existing object-centric latent diffusion models.  \n1. Introduction  \nObject-centric learning (OCL) and vision foundation models (VFMs) pursue complementary objectives. The for-  \nmer seeks structured, compositional representations that decompose a visual scene into distinct, interpretable entities [8, 15], while the latter extracts rich, dense semantic feature spaces from web-scale data [4, 18, 22] . Recently, the boundaries between these two fields have increasingly blurred: advanced object-centric frameworks now utilize pre-trained VFMs as target representations for reconstruction [5, 12, 21], while foundation models themselves have been shown to exhibit emergent, object-level spatial organization without explicit training for scene decomposition.  \nTo scale OCL to complex, real-world images, the community has progressively transitioned from lightweight spatial broadcast decoders [28] to highly expressive generative architectures [12, 23, 24, 32] . In particular, recent diffusionbased object-centric methods, such as Slot-Diffusion [30], Latent Slot Diffusion [10], GLASS [26] and CODA [17], have achieved unprecedented visual fidelity, enabling advanced object-level manipulation and zero-shot compositional generation on complex scenes.  \nAt the same time, Representation Auto-Encoders (RAEs) [25, 34] have demonstrated that the dense feature spaces of self-supervised foundation models are sufficiently expressive to support high-fidelity generative modeling directly, bypassing traditional VAE compression bottlenecks entirely. While deterministic frameworks like DINOSAUR [21] have successfully established that VFM feature spaces are an excellent substrate for unsupervised object discovery, they lack the capacity for probabilistic synthesis, image manipulation, and compositional generation. Conversely, existing generative diffusion-based OCL methods remain tightly bound to VAE latents and external generative priors. This clear dichotomy prompts a fundamental question: Can generative object-centric diffusion modeling  \nObject Slots S  \nFigure 1 . Overview of the Slot-RAE Archi","cbCaiikCcOxfkaoD","https://ap.wps.com/l/cbCaiikCcOxfkaoD","pdf",4015042,2,1,13,"English","en",105,"# Introduction\n## Object-centric learning and vision foundation models\n## Diffusion-based object-centric approaches\n## Representation Auto-Encoders and the core question\n## Slot-RAE: unified object discovery and generation","[{\"question\":\"What problem does Slot-RAE aim to solve in object-centric scene understanding?\",\"answer\":\"Slot-RAE targets the difficulty of achieving both robust scene decomposition and high-fidelity generation without complex pipelines and compression bottlenecks.\"},{\"question\":\"How does Slot-RAE differ from diffusion-based object-centric methods that rely on VAE latents?\",\"answer\":\"Slot-RAE performs diffusion directly in the frozen visual foundation model feature space, eliminating VAE latent bottlenecks and reducing dependence on external generative priors.\"},{\"question\":\"What evidence is provided that Slot-RAE works effectively on real datasets?\",\"answer\":\"Experiments on the COCO dataset report state-of-the-art results, including comparable unsupervised object discovery, higher-fidelity reconstruction, and robust zero-shot compositionality, while remaining faster and more computationally efficient than prior latent diffusion baselines.\"}]",1784209072,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"slot-rae-streamlining-object-centric-learning-via-direct-representation-auto-encoders","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/slot-rae-streamlining-object-centric-learning-via-direct-representation-auto-encoders/86168/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Slot-RAE aim to solve in object-centric scene understanding?","Question",{"text":75,"@type":76},"Slot-RAE targets the difficulty of achieving both robust scene decomposition and high-fidelity generation without complex pipelines and compression bottlenecks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Slot-RAE differ from diffusion-based object-centric methods that rely on VAE latents?",{"text":80,"@type":76},"Slot-RAE performs diffusion directly in the frozen visual foundation model feature space, eliminating VAE latent bottlenecks and reducing dependence on external generative priors.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is provided that Slot-RAE works effectively on real datasets?",{"text":84,"@type":76},"Experiments on the COCO dataset report state-of-the-art results, including comparable unsupervised object discovery, higher-fidelity reconstruction, and robust zero-shot compositionality, while remaining faster and more computationally efficient than prior latent diffusion baselines.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]