[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134154-en":3,"doc-seo-134154-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},134154,962085564549,"Aditya","https://ap-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","RELATE - Physically Plau sible Multi-Object Scene Synthesis Using Structured Latent Spaces - NeurIPS 2020 Paper","RELATE presents an end-to-end model for generating physically plausible scenes and videos containing multiple interacting objects. Trained on raw, unlabeled data, it combines an object-centric GAN with an explicit mechanism to capture correlations between individual objects, yielding a physically interpretable parameterization. The correlation modeling is shown to be necessary for learning disentangled object positions and identity, enabling realistic scene editing. RELATE further extends naturally to dynamic settings, producing high-fidelity videos and outperforming prior object-centric generation on both synthetic and real-world datasets.","RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent Spaces  \nSébastien Ehrhardt 1 􀀃 Oliver Groth1 􀀃 Áron Monszpart2,3 Martin Engelcke1  \nIngmar Posner1 Niloy J. Mitra2,4 Andrea Vedaldi1  \n1Department of Engineering Science, University of Oxford  \n2Department of Computer Science, University College London  \n3Niantic, 4 Adobe Research {hyenal,[ogroth}@robots.ox.ac.uk](ogroth}@robots.ox.ac.uk)  \nAbstract  \nWe present RELATE, a model that learns to generate physically plausible scenes and videos of multiple interacting objects. Similar to other generative approaches, RELATE is trained end-to-end on raw, unlabeled data. RELATE combines an object-centric GAN formulation with a model that explicitly accounts for correlations between individual objects. This allows the model to generate realistic scenes and videos from a physically-interpretable parameterization. Furthermore, we show that modeling the object correlation is necessary to learn to disentangle object positions and identity. We ﬁnd that RELATE is also amenable to physically realistic scene editing and that it signiﬁcantly outperforms prior art in object-centric scene generation in both synthetic (CLEVR, ShapeStacks) and real-world data (cars) .  \nIn addition, in contrast to state-of-the-art methods in object-centric generative modeling, RELATE also extends naturally to dynamic scenes and generates videos of high visual ﬁdelity. Source code, datasets and more results are available at [http://geometry.cs.ucl.ac.uk/projects/2020/relate/](http://geometry.cs.ucl.ac.uk/projects/2020/relate/) .  \n1 Introduction  \nWe consider the problem of learning to generate plausible images of scenes starting from parameters that are physically interpretable. Furthermore, we wish to learn such a capability from raw images alone, without any manual or external supervision. Image generation is often approached via Generative Adversarial Networks (GAN) [10] . These models learn to map noise vectors, used as a source of randomness, to image samples. While the resulting images are realistic, the random vectors that parameterize them are not interpretable. To address this issue, authors have recently proposed to structure the latent space of deep generative models, giving it a partial physical interpretability [28, 29, 36] . For example, HoloGAN [28] samples volumes and cameras to generate 2D images of 3D objects, and BlockGAN [29] creates scenes by composing multiple objects. The resulting GANshave been shown to learn concepts such as viewpoint and object disentangling from raw images.  \nBlockGAN is of particular interest because, via its relatively strong architectural biases, it provides interpretable parameters for the scene, incorporating concepts such as position and orientation. However, BlockGAN comes with a signiﬁcant limitation in that it assumes that objects are mutually independent. This approximation is acceptable only when objects interact weakly, but it is badly violated for medium to densely packed scenes, or for scenes such as stacking wooden blocks or cars following a path, where the (object) correlation is strong.  \n􀀃 indicates equal contribution  \n34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.  \nFigure 1: Image generation using RELATE. Individual scene components, such as background and foreground objects are represented by appearance z0 and pairs of appearance and pose vectors (zi ; 􀀒i); i 2 f1; : : : ; Kg, respectively. The key spatial relationship module 􀀀 adjusts the initial independent pose samples ^􀀒i to be physically plausible (e.g., non-intersecting) to produce 􀀒i. The structured scene tensor W is ﬁnally transformed by the the generator network G to produce an image ^I. RELATE is trained end-to-end in a GAN setup (D denotes the discriminator) on real unlabelled images.  \nRecent work in object-centric generative modeling has attempted to speciﬁcally address this by capturing correlations in latent s","cbCaiaiDSAJG79oV","https://ap.wps.com/l/cbCaiaiDSAJG79oV","pdf",3879051,3,1,12,"English","en",105,"# Introduction\n## Problem setting and motivation\n## Background: structured latent spaces and BlockGAN\n## Key limitation: object independence assumption\n# RELATE approach and contributions\n## Modeling object correlations in latent space\n## Benefits: disentanglement and physical plausibility\n## Applications and evaluations","[{\"question\":\"What does RELATE generate and how is it trained?\",\"answer\":\"RELATE generates physically plausible scenes and videos with multiple interacting objects. It is trained end-to-end in a GAN setup using raw, unlabeled images without manual supervision.\"},{\"question\":\"Why is object correlation modeling important in RELATE?\",\"answer\":\"The paper shows that explicitly modeling correlations is necessary to learn disentangled object positions and identity. Without this component, results may look realistic but fail to align parameters with physically plausible object associations.\"},{\"question\":\"How does RELATE handle dynamic scenes and what evidence is reported?\",\"answer\":\"RELATE extends naturally from static object-centric generation to dynamic settings. The paper reports generation of high-visual-fidelity videos and demonstrates that it outperforms prior work on object-centric scene generation benchmarks.\"}]","RELATE - Physically Plau sible Multi-Object Scene Synthesis Using Structured Latent Spaces - NeurIPS 2020 Paper | PDF",1787235252,30,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"relate-physically-plausible-multi-object-scene-synthesis-using-structured-latent-spaces-neurips-2020-paper","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/relate-physically-plausible-multi-object-scene-synthesis-using-structured-latent-spaces-neurips-2020-paper/134154/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does RELATE generate and how is it trained?","Question",{"text":76,"@type":77},"RELATE generates physically plausible scenes and videos with multiple interacting objects. It is trained end-to-end in a GAN setup using raw, unlabeled images without manual supervision.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why is object correlation modeling important in RELATE?",{"text":81,"@type":77},"The paper shows that explicitly modeling correlations is necessary to learn disentangled object positions and identity. Without this component, results may look realistic but fail to align parameters with physically plausible object associations.",{"name":83,"@type":74,"acceptedAnswer":84},"How does RELATE handle dynamic scenes and what evidence is reported?",{"text":85,"@type":77},"RELATE extends naturally from static object-centric generation to dynamic settings. The paper reports generation of high-visual-fidelity videos and demonstrates that it outperforms prior work on object-centric scene generation benchmarks.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":30,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]