[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118107-en":3,"doc-seo-118107-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118107,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Variational Autoencoders for Supervision, Calibration and Multimodal Learning - Thesis","Learning representations of data is a central goal in machine learning for efficient downstream use such as classification and object detection. Interpretability is also essential, enabling fine-grained intervention and reasoning over input characteristics. This thesis uses Variational Autoencoders to incorporate label information into latent learning, build shared representations for multimodal data, and calibrate predictions from neural classifiers. It develops methods for structured latent spaces with available labels, handles missing labels to reduce data requirements, and learns symmetric mutual supervision for cross-modal generation. Finally, it produces reliable confidence estimates via fast post-hoc calibration.","Variational Autoencoders for Supervision, Calibration and Multimodal Learning  \nThomas W. Joy St Catherine’s College University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nTrinity 2022  \nAcknowledgements  \nI would like to extend my deepest gratitude to my supervisor Phil and my cosupervisor Sid. Their consistent and unwavering support has been invaluable throughout my PhD, I would like to thank them for believing in me and providing me with unbridled opportunities to undertake research, allowing me to develop both professionally and personally. I would also like to thank Puneet, for his inspirational mentorship and guidance throughout my PhD.  \nI would also like to extend my gratitude to Pawan, who supervised me throughout my Masters degree and set me on a path towards a PhD. His support and patience provided me with the confidence and vision to pursue a path in research.  \nI would also like to thank my undergraduate tutors, Byron Byrne and David Gillespie for their academic and administrative support during my applications. Furthermore, I would also to thank the staff at St Catherines College for being the unsung heroes and heroines of the university. I would also like to particularly thank Joanna Zapisek for her continued support and reassurance throughout the process.  \nFrom a personal point of view, I wold like to thank my family for believing in meand providing essential emotional support. I would also like to thank my friends for their empathy during the difficult times and covering the Uber when I was financially insolvent. Finally, to anyone else who was part of the past four years, no matter how insignificant, I would like to say thank you for being a part of this capricious but immensely gratifying period of my life.  \nAbstract  \nLearning representations of data has long been a desirable goal in machine learning. Constructing such representations enables downstream tasks such as classification or object detection to be preformed efficiently. Furthermore, it is desirable to have these representations be constructed in such a way so they are interpretable, which allows for fine grained intervention and reasoning on characteristics of the input. Other tasks may include, cross-generation between modalities, or calibrating predictions such that their confidence matches their accuracy. An effective way to learn representations is through a Variational Autoencoder (VAE), which performs variational inference on the latent variables of the observable input. In this thesis we show how the VAE, can be utilised to: incorporate label information into the learning process; learn shared-representations of multimodal data; and calibrate predictions of existing neural classifiers.  \nData sources are often accompanied by additional label information, which may indicate the presence of a characteristic in the input. A question naturally arises as to whether the additional label information can be used to structure the representation such that it provides a notion of interoperability about the characteristic; such as “to what extent is the person smiling?”. The first contribution of this thesis is to address the aforementioned problem and propose a method which successfully uses label information to structure the latent space. Furthermore, this allows us to perform additional tasks such as fine grained interventions; classification; and conditional generations. Moreover, we are also successfully able to handle the case when label information is missing, drastically reducing the data burden when training these models.  \nRather than being presented with labels, we sometimes instead observe another unstructured observation of the same object, e .g. a caption of an image. In this scenario, the objective changes slightly to one where the model is able to learn shared-representations of data, allowing it to perform cross-generations between modalities. The second contribution of this theses addresses this problem. ","cbCaiuuuRdhybXFc","https://ap.wps.com/l/cbCaiuuuRdhybXFc","pdf",37894642,1,179,"English","en",105,"# 1 Introduction\n## 1.1 Preamble\n## 1.2 Variational Autoencoder\n## 1.3 Variational Autoencoders for Supervision, Calibration and Multimodal Learning\n## 1.3.1 Capturing Label Characteristics in VAEs\n## 1.3.2 Multi-Modal learning through Mutual Supervision\n## 1.3.3 Sample-dependent Temperature Scaling for Improved Calibration\n# 2 Literature Review and Background\n## 2.1 Capturing Label Characteristics in VAEs\n## 2.2 Learning Multimodal VAEs through Mutual Supervision\n## 2.3 Sample-dependent Temperature Scaling for Improved Calibration","[{\"question\":\"How does this thesis use variational autoencoders to improve learning representations?\",\"answer\":\"It employs Variational Autoencoders to perform variational inference on latent variables and to structure representations for downstream tasks. The work tailors VAE usage for supervision, multimodal representation learning, and prediction calibration.\"},{\"question\":\"What is the first main contribution related to supervision and labels?\",\"answer\":\"The thesis proposes a method that uses available label information to structure the latent space so it supports further tasks such as fine-grained interventions and conditional generation. It also addresses missing-label settings to reduce training data burden.\"},{\"question\":\"How does the thesis handle multimodal learning when labels are not available?\",\"answer\":\"When labels are not provided, it learns shared representations by introducing mutual supervision between modalities and a bi-directional objective that enforces symmetry. This enables cross-generation and robustness when some modalities are missing during training.\"}]","Variational Autoencoders for Supervision, Calibration and Multimodal Learning - Thesis | PDF",1785681649,451,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"variational-autoencoders-for-supervision-calibration-and-multimodal-learning-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/variational-autoencoders-for-supervision-calibration-and-multimodal-learning-thesis/118107/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does this thesis use variational autoencoders to improve learning representations?","Question",{"text":75,"@type":76},"It employs Variational Autoencoders to perform variational inference on latent variables and to structure representations for downstream tasks. The work tailors VAE usage for supervision, multimodal representation learning, and prediction calibration.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the first main contribution related to supervision and labels?",{"text":80,"@type":76},"The thesis proposes a method that uses available label information to structure the latent space so it supports further tasks such as fine-grained interventions and conditional generation. It also addresses missing-label settings to reduce training data burden.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis handle multimodal learning when labels are not available?",{"text":84,"@type":76},"When labels are not provided, it learns shared representations by introducing mutual supervision between modalities and a bi-directional objective that enforces symmetry. This enables cross-generation and robustness when some modalities are missing during training.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]