[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125830-en":3,"doc-seo-125830-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125830,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Cross-Modal Learning from Visual Information for Activity Recognition on Inertial Sensors - 博士论文致谢与摘要","The lack of large-scale, labeled datasets limits robust, generalized human activity recognition (HAR) from wearable inertial sensors. Inertial data collection is costly and annotation is time-consuming and error-prone, making public datasets small in subjects, classes, recorded hours, and environmental diversity. Models trained on these datasets fail to cover real-world activity variations across populations, slowing wearable sensing progress. This thesis leverages the visual modality as a source domain for cross-modal learning to transfer data and knowledge into inertial HAR, enabling scalable learning and improved generalization.","Cross-Modal Learning from Visual Information for Activity Recognition on Inertial Sensors  \n􀀀  \nCatherine Tong  \nUniversity of Oxford  \nA thesis submitted for the degree of  \nDoctor of Philosophy  \nTrinity Term 2021-2022  \nAcknowledgements  \nMy gratitude for help in writing this thesis runs wide and deep.  \nAs ever, my thanks go ﬁrst to my supervisor Nic Lane. He brought his wisdom and experience to my support through all important moments during my studies. I thank him especially for taking a leap of faith in a Physics student who knew very little about the world of Computer Science.  \nAt the very beginning ofthe DPhil programme, I took up an internship at Nokia Bell Labs in Cambridge. It turned out to an invaluable experience that sparked my interest in Ubicomp. I owe thanks to Fahim Kawsar, Akhil Mathur, Sourav Bhattacharya, Alberto Gil Ramos, and Vincent Tseng, who ﬁrst introduced me to the ﬁeld. I also would like to thank Danielle Belgrave who was my manager during an internship at Microsoft Research Cambridge, for whose advice I am particular grateful. At Oxford, I spent a considerable amount of time at the Big Data Institute working with Aiden Doherty and his group, namely Shing Chan, Hang Yuan, Rosemary Walmsley and Andrew Creagh, who always made me feel welcomed.  \nI am also grateful to my friends and family who have been a constant source of support and advice during the past few years, particularly: Milad Alizadeh, Javier Fernandez-Marques, Shyam Tailor, Edgar Liberis, Shuyu Lin, Ronnie Clark, Bo Yang, Yuge Shi, Hyeokhyen Kwon, Harish Haresamudram, Chongyang Wang, Yan Gao, Xinchi Qiu, Priscilla Wong, Emma Rocheteau, Elden Tse, Dongge Han and Lily Liu.  \nI also thank my mother, without whom my journey in the UK would not have even started.  \nFinally, I would like to thank William, who has to endure the ups and downs of a PhD without earning one himself. He has powered me through different challenges and this thesis would not have been possible without him.  \nAbstract  \nThe lack of large-scale, labeled datasets impedes progress in developing robust and generalized predictive models for human activity recognition (HAR) from wearable inertial sensor data. Labeled data is scarce as sensor data collection is expensive, and their annotation is time-consuming and error-prone. As a result, public inertial HAR datasets are small in terms of number of subjects, activity classes, hours of recorded data, and variation in recorded environments. Machine learning models, developed using these small datasets, are effectively blind to the diverse expressions of activities performed by wide-ranging populations in the real world, and progress in wearable inertial sensing is held back by this bottleneck for activity understanding.  \nBut just as Internet-scale text, image and audio data have pushed their respective pattern recognition ﬁelds to systems reliable enough for everyday use, easy access to large quantities of data can push forward the ﬁeld of inertial HAR, and by extension wearable sensing. To this end, this thesis pioneers the idea of exploiting the visual modality as a source domain for cross-modal learning, such that data and knowledge can be transferred across to beneﬁt the target domain of inertial HAR.  \nThis thesis makes three contributions to inertial HAR through cross-modal approaches. First, to overcome the barrier of expensive inertial data collection and annotation, we contribute a novel pipeline that automatically extracts virtual accelerometer data from videos of human activities, which are readily annotated and accessible in large quantities. Second, we propose acquiring transferable representations about activities, from HAR models trained using large quantities of visual data to enrich the development of inertial HAR models. Finally, the third contribution exposes HAR models to the challenging setting of zero-shot learning; we propose mechanisms that leverage cross-modal correspondence to enable inference on pr","cbCaigZwXubTgkIG","https://ap.wps.com/l/cbCaigZwXubTgkIG","pdf",7626691,1,148,"English","en",105,"# Acknowledgements\n# Abstract\n# Contents\n## 1 Introduction\n### 1.1 Research Questions and Contributions\n### 1.2 Publications\n## 2 Background\n### 2.1 Activity Recognition Using Wearable Inertial Sensors\n### 2.2 Learning Across Modalities\n### 2.3 Summary\n## 3 Generating Virtual Inertial Data from Videos\n### 3.1 Introduction\n### 3.2 IMUTube\n### 3.3 Experiments\n### 3.4 Discussion\n### 3.5 Related Work\n### 3.6 Summary","[{\"question\":\"Why are large-scale labeled inertial datasets important for activity recognition?\",\"answer\":\"Large-scale labeled data enables robust and generalized predictive models. Without it, inertial HAR models trained on small public datasets struggle to handle diverse real-world activity expressions.\"},{\"question\":\"How does the thesis propose to obtain inertial information without expensive inertial collection?\",\"answer\":\"It introduces a pipeline that automatically extracts virtual accelerometer data from videos of human activities, leveraging readily annotated and abundant visual data.\"},{\"question\":\"What does the thesis mean by zero-shot learning in inertial HAR?\",\"answer\":\"It studies mechanisms that use cross-modal correspondence so HAR models can infer on previously unseen activity classes.\"}]","Cross-Modal Learning from Visual Information for Activity Recognition on Inertial Sensors - 博士论文致谢与摘要 | PDF",1785901461,373,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cross-modal-learning-from-visual-information-for-activity-recognition-on-inertial-sensors-phd-thesis-acknowledgements-and-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/cross-modal-learning-from-visual-information-for-activity-recognition-on-inertial-sensors-phd-thesis-acknowledgements-and-abstract/125830/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are large-scale labeled inertial datasets important for activity recognition?","Question",{"text":75,"@type":76},"Large-scale labeled data enables robust and generalized predictive models. Without it, inertial HAR models trained on small public datasets struggle to handle diverse real-world activity expressions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis propose to obtain inertial information without expensive inertial collection?",{"text":80,"@type":76},"It introduces a pipeline that automatically extracts virtual accelerometer data from videos of human activities, leveraging readily annotated and abundant visual data.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the thesis mean by zero-shot learning in inertial HAR?",{"text":84,"@type":76},"It studies mechanisms that use cross-modal correspondence so HAR models can infer on previously unseen activity classes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]