[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128448-en":3,"doc-seo-128448-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128448,962085662650,"Jiven","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Self-supervised Video Representation Learning - Doctor of Philosophy Thesis","Videos provide rich data for training computer vision models, yet exhaustive manual annotation is impractical at scale. This thesis learns strong video representations efficiently using self-supervised learning, relying on data signals rather than human labels. The work focuses on three areas: short-term video learning, efficient representation learning, and long-term video learning. Results show predictive future frames, cross-modal RGB-optical flow mutual teaching, prompt tuning with vision-language models, patch dropping acceleration, and temporal alignment from weak video-text correspondence.","Self-supervised  \nVideo Representation Learning  \nTengda Han  \nLady Margaret Hall University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nTrinity 2022  \nAbstract  \nVideos are an appealing source of data to train computer vision models. There exist almost inﬁnite supplies of videos online, but exhaustive manual annotation is infeasible. The goal of this thesis is to learn strong video representations eﬃciently via self-supervised learning: a method that learns from the data rather than human annotations.  \nThe thesis is structured around three themes: (1) self-supervised learning for short-term videos, (2) eﬃcient video representation learning, and (3) selfsupervised learning for long-term videos.  \nFor short-term videos lasting only a few seconds, we show that predicting the video in the future is a strong learning signal at a large scale. We further show that strong video representations can be learned by taking two complementary modalities, namely RGB and optical ﬂow, and using them to teach each other.  \nFor eﬃcient video representation learning, we show that large-scale pre-trained vision-language models can be eﬀectively adapted via a prompt tuning technique. We also show that dropping image patches can accelerate the ﬁnetuning of classiﬁcation tasks and pre-training of video-language models.  \nFor long-term videos that last longer than a few minutes, we show that temporal alignment networks can be trained from the weak visual-textual correspondence within instructional videos. The resulting networks can automatically clean up the natural videos for eﬀective vision-language training. In addition, we show that movie description models can be trained by leveraging the pre-trained visionlanguage models.  \nKeywords – video understanding, deep learning, self-supervision, eﬃcient learning  \nThis thesis is submitted to the Department of Engineering Science, The University of Oxford, in fulﬁlment of the requirements for the degree of Doctor of Philosophy. This thesis is entirely my own work, and except where otherwise stated, describes my own research.  \nTengda Han, October 2022 .  \nAcknowledgement  \nFirst and foremost to my supervisor Andrew Zisserman, thank you for your guidance, advice and trust along this journey. I feel extremely fortunate to have worked with you. You are my role model in so many aspects. Similarly to Weidi Xie, my wonderful co-supervisor and long-term collaborator, thank you for your intuition and in-depth guidance. The support from you two makes this experience much more enjoyable.  \nTo all the VGG members: thank you for making VGG such a lovely group, of which I feel privileged to be a member. In particular to Andrew Brown, Shangzhe Wu and Chuhan Zhang for the great times we spent together. To Max Bain, Gül Varol and Arsha Nagrani for the great collaboration. To Yuki Asano for insightful discussions and wine tasting in Porto. Especially to Triantafyllos Afouras, who bought me so much beer in the pubs over the years, thank you for your great company. I also thank Robert McCraith, Hala Lamdouar, Liliane Momeni, Charig Yang and other VGG members for their support and camaraderie. Importantly, I thank Ashish Thandavan for answering my endless server questions and Jenny Hu for crucial group management eﬀorts.  \nI also thank my friends outside of VGG: Yang Zhang, Zirui Wang and Yizhe Wu, for the laughter on the football courts and foosball tables. Berkeley friends Allan Jabriand Sheng-Yu Wang, for the wonderful memories on the river and with a guitar during my ﬁrst year. Richard Hartley, for kindly introducing me to VGG and Oxford at the beginning. Jean-Baptiste Alayrac, Karen Simonyan and the entire Flamingo team at Deepmind, for their kindness and patience during my internship.  \nThank Deepmind for generously funding my DPhil programme. Also, I thank my college Lady Margaret Hall for the pianos, library, lively gardens and the River Cherwell.  \nFinally, I thank my parents for their ","cbCaibo4K6xjUrnK","https://ap.wps.com/l/cbCaibo4K6xjUrnK","pdf",14686637,1,204,"English","en",105,"# Introduction and Background\n## Motivation and Key Ideas\n## Thesis Outline and Contributions\n## Publications\n# Self-supervised Learning for Short-term Videos\n## Video Representation Learning by Dense Predictive Coding\n## Introduction\n## Related Work\n## Dense Predictive Coding (DPC)\n## Experiments and Analysis\n## Comparison with State-of-the-art Methods\n## Conclusion","[{\"question\":\"What is the main goal of this thesis?\",\"answer\":\"To learn strong video representations efficiently using self-supervised learning, avoiding reliance on exhaustive human annotations.\"},{\"question\":\"How does the thesis approach learning from short-term videos?\",\"answer\":\"It shows that predicting future video content provides a strong learning signal and demonstrates mutual teaching between RGB and optical flow modalities.\"},{\"question\":\"How are long-term videos handled for vision-language training?\",\"answer\":\"Temporal alignment networks are trained from weak visual-textual correspondence in instructional videos to clean natural videos and enable effective vision-language training.\"}]","Self-supervised Video Representation Learning - Doctor of Philosophy Thesis | PDF",1786001112,514,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"self-supervised-video-representation-learning-doctor-of-philosophy-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/self-supervised-video-representation-learning-doctor-of-philosophy-thesis/128448/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main goal of this thesis?","Question",{"text":76,"@type":77},"To learn strong video representations efficiently using self-supervised learning, avoiding reliance on exhaustive human annotations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis approach learning from short-term videos?",{"text":81,"@type":77},"It shows that predicting future video content provides a strong learning signal and demonstrates mutual teaching between RGB and optical flow modalities.",{"name":83,"@type":74,"acceptedAnswer":84},"How are long-term videos handled for vision-language training?",{"text":85,"@type":77},"Temporal alignment networks are trained from weak visual-textual correspondence in instructional videos to clean natural videos and enable effective vision-language training.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]