[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118101-en":3,"doc-seo-118101-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118101,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","End-to-End Learning - Audio-Visual Human-Centric Video Understanding","Machine learning has advanced rapidly through deep neural networks that learn semantic representations end-to-end from large labeled datasets. End-to-end learning depends on usable input data, target outputs, and an objective function aligning predictions to targets. This thesis addresses challenges in formatting and scaling these components. It develops Smooth-AP to optimize a differentiable approximation of Average Precision for image retrieval, and a self-supervised strategy for end-to-end image editing when target data is scarce. It also uses audio-visual signals for human-centric video understanding, including transfer from speaker verification and methods that fuse identity-discriminating modalities for scene understanding, labelling, and identity clustering.","End-to-End Learning, and  \nAudio-Visual Human-Centric Video Understanding  \nAndrew Brown  \nLinacre College University of Oxford Supervisor: Professor Andrew Zisserman A thesis presented for the degree of Doctor of Philosophy  \nMichaelmas 2022  \nAcknowledgements  \nI am very grateful for the guidance and support of many important people throughout this journey.  \nFirst, to my supervisor Andrew Zisserman. Your creative, exciting approach to problem-solving has been inspiring, and has shaped me into the researcher that I am today. It has been a privilege to work with you for the last ﬁve years. Thank you for your guidance and for seeing the potential in a Matlabwielding ﬂedgling all of those years ago.  \nThank you to the collaborators that I have been lucky enough to work with. Weidi Xie, who taught me the value and importance of solving hard problems. Thank you for guiding me through my ﬁrst research paper, for providing clarity with your vast knowledge of the ﬁeld, and for answering my hundreds of questions without hesitation, even as a new and tired parent. To Andrea Vedaldi, for being so generous with his insights and time. Thank you for showing me the importance of detail, and for the iconic COVID-19 memories of having equations explained to me on virtual whiteboards written on from Venetian public transport. To Arsha Nagrani, who has been an invaluable mentor to me throughout my time at VGG and guided me up the many steep learning curves. To Ernesto Coto, thank you for your unwavering patience in helping me with the many engineering mysteries encountered on this journey, such as compiling Caffe. To all my collaborators: Vicky Kalogeiton, Joon Son Chung, Max Bain, Jaesung Huh, Samuel Albanie, Yang Liu, Bruno Korbar, Cheng-Yang Fu, Omkar Parkhi, and Tamara Berg. Thank you sharing your ideas and time with me, and for making research meetings so exciting and inspiring.  \nTo all the members of the VGG family: Tengda Han, Chuhan Zhang, Yuki Asano, Olivia Wiles, Sagar Vaze, Shangzhe Wu, Mandela Patrick, Gul Varol, Lili Momeni, Rhydian Windsor, Daffy Afouras, Tomas Jakab, Xu Ji, Christian Rupprecht, Erika Lu, and João Henriques. Thank you for creating such a welcoming and supportive environment, for the lunches at St. Anne’s, and the good evenings at the Royal Oak. I feel very privileged to have learnt from all of you, and I am excited to see you all grow tobe giants in the research ﬁeld in the coming years.  \nTo my family (Jo, Phil, Will, Katie, and Candice), thank you for supporting me when I raised the prospect of spending yet another few years in education. To my friends in London who were still there  \nfor me after I disappeared off the face of the earth for some time. You all mean the world to me and Iam so lucky to have you all in my life.  \nFinally, to Maya. Thank you for your endless encouragement, support, and understanding. Thank you for believing in my dream before I did, and for subsequently telling a young and lost undergraduate what was coming if he instead decided to go down a career-path in the City. Thank you for pushing me to follow this dream, and for encouraging me every time I fell along the way. I could not have done this without her, and I am so excited for our next journey.  \nThis is dedicated to her.  \nAbstract  \nThe ﬁeld of machine learning has seen tremendous progress in the last decade, largely due to the advent of deep neural networks. When trained on large-scale labelled datasets, these machine learning algorithms can learn powerful semantic representations directly from the input data, end-to-end. End-to-end learning requires the availability of three core components: useful input data, target outputs, and an objective function for measuring how well the model’s predictions match the target outputs. In this thesis, we explore and overcome a series of challenges as related to assembling these three components in the sufﬁcient format and scale for end-to-end learning.  \nThe ﬁrst key idea presented in th","cbCaimcUbSFSTuZf","https://ap.wps.com/l/cbCaimcUbSFSTuZf","pdf",30894141,1,152,"English","en",105,"# Introduction and Background\n## Motivation\n## Key Ideas and Contributions\n## Enabling","[{\"question\":\"什么是端到端学习，并且它需要哪些关键组成部分？\",\"answer\":\"端到端学习通过深度网络直接从输入数据学习语义表示。它需要可用的输入数据、目标输出以及衡量预测与目标匹配程度的目标函数。\"},{\"question\":\"Smooth-AP在图像检索任务中解决了什么难点？\",\"answer\":\"图像检索常用的基于排序的Average Precision度量不可微。论文提出Smooth-AP，通过优化Average Precision的平滑近似来实现端到端训练。\"},{\"question\":\"论文如何在缺少足够规模目标数据的情况下进行图像编辑的端到端学习？\",\"answer\":\"论文提出自监督方法，通过对现成图像数据进行增强来模拟目标数据，从而在数据不足的情况下获得相对既有工作的收益。\"}]","End-to-End Learning - Audio-Visual Human-Centric Video Understanding | PDF",1785681600,383,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"end-to-end-learning-audio-visual-human-centric-video-understanding","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/end-to-end-learning-audio-visual-human-centric-video-understanding/118101/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"什么是端到端学习，并且它需要哪些关键组成部分？","Question",{"text":75,"@type":76},"端到端学习通过深度网络直接从输入数据学习语义表示。它需要可用的输入数据、目标输出以及衡量预测与目标匹配程度的目标函数。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Smooth-AP在图像检索任务中解决了什么难点？",{"text":80,"@type":76},"图像检索常用的基于排序的Average Precision度量不可微。论文提出Smooth-AP，通过优化Average Precision的平滑近似来实现端到端训练。",{"name":82,"@type":73,"acceptedAnswer":83},"论文如何在缺少足够规模目标数据的情况下进行图像编辑的端到端学习？",{"text":84,"@type":76},"论文提出自监督方法，通过对现成图像数据进行增强来模拟目标数据，从而在数据不足的情况下获得相对既有工作的收益。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]