[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123866-en":3,"doc-seo-123866-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123866,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Learning Object-Centric Representations - Thesis","Whenever an agent interacts with its environment, it must account for objects present in that environment, yet many machine learning approaches treat objects only implicitly or rely on engineered pipelines built on object detection. This thesis investigates supervised and unsupervised learning of object-centric representations from vision. It emphasizes end-to-end learning where object information is extracted directly from images and represented with per-object vector variables. Three new methods are introduced for tracking, generative modeling, and unsupervised part- and object-discovery.","Learning Object-Centric Representations  \nAdam Roman Kosiorek  \nWolfson College University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy Michaelmas 2019  \nAcknowledgements  \nI would like to thank  \n• My mum for encouraging my creativity and answering my endless questions when I was growing up.  \n• My dad for pushing me to work hard, teaching me about business and negotiation, and mentoring me in my carrer so far. It was him who planted the idea of doing a PhD in my mind.  \n• My grandparents for raising me, always being there for me, providing a secure base, and showing me what it means to be loved.  \n• My dear childhood friend Rafał Worobiew for taking care of me at a crucial moment, and showing me an alternative path.  \n• Prof. Barbara Siemi ˛atkowska of the Warsaw Univeristy of Technology for teaching an intro to computer vision course, which is where my journey to AI started.  \n• Andrzej Ruta, at the time a team lead at Samsung R&D in Warsaw, who hired the clueless student that I was as an intern in his computer vision research team.  \n• Patrick van der Smagt of the Technical University of Munich (at the time) for teaching my first formal ML course.  \n• Martin Engelcke who, during a chance encounter, convinced me to apply to Oxford and specifically to his supervisor Ingmar Posner’s group.  \n• My supervisor Ingmar Posner who welcomed me to his group at Oxford, initiated me into the land of independent research, but was also able to tolerate my sometimes harsh remarks and moves, and ultimately allowed me to move  \naway.  \n• My supervisor Yee Whye Teh who allowed me to join his group in my 2nd year of PhD.  \n• Danilo Rezende for hosting me as an intern at DeepMind.  \n• Geoff Hinton for inviting me for an internship after a brief discussion at a conference.  \n• My friends and colleagues who turned this PhD into an unforgettable journey: Alex Bewley, Julie Dequire, Neil Dhir, Martin Engelcke, Fabian Fuchs, Adam Goli ´nski, Oliver Groth, Sandy Huang, Hyunjik Kim, Tuan Anh Le, Milan Muso, Tom Rainforth, S, tefan S˘aftescu, Hillary Shakespeare, Jimmy Shi, Georgi Tinchev, Markus Wulfmeier, and many others.  \nThe funding for this thesis was provided by the University of Oxford and a scholarship funded by Google Deepmind. I extend my thanks to Ingmar Posner for helping secure the scholarship.  \nAbstract  \nWhenever an agent interacts with its environment, it has to take into account and interact with any objects present in this environment. And yet, the majority of machine learning solutions either treat objects only implicitly or employ highly-engineered solutions that account for objects through object detection algorithms. In this thesis, we explore supervised and unsupervised methods for learning object-centric representations from vision. We focus on end-to-end learning, where information about objects can be extracted directly from images, and where every object can be separately described by a single vector-valued variable. Specifically, we present three novel methods:  \n• Hart and mohart, which track single-and multiple-objects in video, respectively, by using rnns with a hierarchy of differentiable attention mechanisms. These algorithms learn to anticipate future appearance changes and movement of tracking objects, thereby learning representations that describe every tracked object separately.  \n• Sqair, a vae-based generative model of moving objects, which explicitly models disappearance and appearance of new objects in the scene. It models every object with a separate latent variable, and disentangles appearance, position and scale of each object. Posterior inference in this model allows for unsupervised object detection and tracking.  \n• Scae, an unsupervised autoencoder with in-built knowledge of two-dimensional geometry and object-part decomposition, which is based on capsule networks. It learns to discover parts present in an image, and group those parts into objects. Each object is modelled by a ","cbCaiea8VfNW1Ies","https://ap.wps.com/l/cbCaiea8VfNW1Ies","pdf",20250192,1,158,"English","en",105,"# Introduction\n## Works Omitted from the Thesis and Developed in the Course of This DPhil\n# Background\n## Relational Reasoning\n## Object Detection and Tracking\n## Generative Modelling and Representation Learning\n# Hierarchical Attentive Recurrent Tracking\n## Introduction\n## Related Work\n## Hierarchical Attention\n## Loss\n## Experiments\n## Discussion\n## Conclusion\n# End-to-end Recurrent Multi-Object Tracking and Trajectory Prediction with Relational Reasoning\n## Introduction\n## Related Work\n## Recurrent Multi-Object Tracking with Self-Attention\n## Validation on Simulated Data\n## Relational Reasoning in Real-World Tracking\n## Conclusion","[{\"question\":\"What problem does the thesis address in machine learning for object understanding?\",\"answer\":\"It focuses on how agents should interact with objects in their environment, while existing solutions often treat objects implicitly or depend on engineered object detection pipelines.\"},{\"question\":\"What is the thesis’s main approach to learning object-centric representations?\",\"answer\":\"It explores supervised and unsupervised end-to-end methods where object information is extracted directly from images and each object is represented by its own vector-valued variable.\"},{\"question\":\"What three novel methods does the thesis introduce?\",\"answer\":\"It presents Hart/mohart for hierarchical attentive recurrent tracking, Sqair for generative moving-object modeling with latent variables per object, and Scae for unsupervised autoencoding using 2D geometry and object-part decomposition with capsule networks.\"}]","Learning Object-Centric Representations - Thesis | PDF",1785818967,398,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-object-centric-representations-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-object-centric-representations-thesis/123866/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address in machine learning for object understanding?","Question",{"text":75,"@type":76},"It focuses on how agents should interact with objects in their environment, while existing solutions often treat objects implicitly or depend on engineered object detection pipelines.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the thesis’s main approach to learning object-centric representations?",{"text":80,"@type":76},"It explores supervised and unsupervised end-to-end methods where object information is extracted directly from images and each object is represented by its own vector-valued variable.",{"name":82,"@type":73,"acceptedAnswer":83},"What three novel methods does the thesis introduce?",{"text":84,"@type":76},"It presents Hart/mohart for hierarchical attentive recurrent tracking, Sqair for generative moving-object modeling with latent variables per object, and Scae for unsupervised autoencoding using 2D geometry and object-part decomposition with capsule networks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]