[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121868-en":3,"doc-seo-121868-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121868,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Leveraging Automated Machine Learning - Edge Computing for Video Understanding - A Thesis","Computer Vision is growing rapidly due to deep learning advances in classification, action recognition, segmentation, and object detection. Video-based action recognition is central to video understanding and supports security and behavior analysis, yet achieving strong performance demands heavy engineering for module selection and hyperparameter tuning. This thesis presents AutoVideo, an AutoML framework for automated video action recognition that uses a modular pipeline, a broad set of primitives, data-driven tuners, and an easy GUI. It also introduces BED for deploying object detection on edge devices using the MAX78000 DNN accelerator, integrating on-device inference with a camera and display while addressing low-latency, low-power, and privacy constraints.","LEVERAGING AUTOMATED MACHINE LEARNING, EDGE COMPUTING  \nFOR VIDEO UNDERSTANDING  \nA Thesis  \nby  \nZAID PERVAIZ BHAT  \nSubmitted to the Graduate and Professional School of Texas A&M University  \nin partial fulfillment of the requirements for the degree of  \nMASTER OF SCIENCE  \nChair of Committee,  \nCo-Chair of Committee, Committee Member, Head of Department,  \nXia “Ben” Hu James Caverlee Xiaoning Qian Scott Schaefer  \nDecember 2022  \nMajor Subject: Computer Science  \nCopyright 2022 Zaid Pervaiz Bhat  \nABSTRACT  \nComputer Vision is witnessing unprecedented growth over the past few years mainly because of the applications of deep learning methods to computer vision tasks like classification, action recognition, segmentation, and object detection. Video-based action recognition is an important task for video understanding with broad applications in security and behavior analysis. However, developing an effective action recognition solution often requires extensive engineering efforts in building and testing different combinations ofthe modules and optimizing for the best set oftheir hyperparameters. The recent advancements in computer vision has shown its vast applicability across several real-world problems. However, developing an optimal end-end machine learning pipeline requires considerable knowledge in computer vision and significant engineering efforts by the developers. To address these problems, in this paper, we present AutoVideo, an AutoML framework for automated video action recognition. AutoVideo aims to tackle these problems by 1) being a highly modular and extendable infrastructure following the standard pipeline language, 2) having an exhaustive list of primitives for pipeline construction, 3) including data-driven tuners to save the efforts of pipeline tuning, and 4) integrating an easy-to-use Graphical User Interface (GUI) .  \nAnother major problem with computer vision applications is the deployment of these machine learning models to edge devices for real world applications, especially because these usually require low latency, low power or data privacy. This requires significant research and engineering efforts due to the computational and memory limitations of edge devices. To tackle this problem, we also present BED, an object detection system for edge devices practiced on the MAX78000 DNN accelerator. To demonstrate real world applicability, we integrate on-device  \nDNN inference with a camera and a screen for image acquisition and output exhibition respectively.  \nAutoVideo is released at GitHub-AutoVideo-GitHub under MIT license with a demo video hosted at Demo Video-AutoVideo while BED is released at Github-BED_main-GitHub with a demo video at Demo Video-BED.  \nACKNOWLEDGMENTS  \nThere are several people I would like to acknowledge for this thesis work who not only opened the door of Artificial Intelligence to me but also made this 2-year Masters program at TAMU a memorable experience, one that I will always cherish.  \nI would first like to truly appreciate my advisor, Dr. Xia (Ben) Hu, for his support, guidance, and encouragement throughout the course of my Masters. He believed in me and gave me the opportunity to work on novel research at TAMU to build on ideas focussed on making a difference. His insightful discussions in the weekly and bi-weekly lab meetings have had a huge impact on the way I view research and development as the discussions motivate me to discover and tackle novel problems.  \nI would also like to thank my committee members, Dr. James Caverlee and Dr. Xiaoning Qian for their guidance and valuable comments on the research.  \nFurthermore, I would like to thank all members of the DATA lab, especially Daochen Zha, Guanchu Wang, Yi-Wei, Henry Lai, and Zhimeng Zhang Jiang. It was a wonderful experience and so much fun working with such a great group of people. The DATA Lab at TAMU maintained a healthy, inclusive and friendly environment. Finally, my deepest gratitude goes to my family for the","cbCaiaDezK7zSkiK","https://ap.wps.com/l/cbCaiaDezK7zSkiK","pdf",15123590,1,87,"English","en",105,"# Abstract\n# Acknowledgments\n# Contributors and Funding Sources\n# Nomenclature\n# Table of Contents\n# List of Figures\n# List of Tables","[{\"question\":\"What challenges does the thesis address in video understanding and action recognition?\",\"answer\":\"Developing effective video action recognition often requires extensive engineering to combine modules and optimize hyperparameters. Building an optimal end-to-end machine learning pipeline also demands substantial computer vision knowledge and effort from developers.\"},{\"question\":\"What is AutoVideo and how does it automate video action recognition?\",\"answer\":\"AutoVideo is an AutoML framework designed for automated video action recognition. It provides a modular, extendable pipeline, includes a comprehensive set of primitives, adds data-driven tuners to reduce pipeline tuning effort, and offers an easy-to-use GUI.\"},{\"question\":\"How does BED enable object detection on edge devices?\",\"answer\":\"BED targets edge deployment by running object detection on the MAX78000 DNN accelerator. The system integrates on-device DNN inference with a camera for image acquisition and a screen for output exhibition to support real-world use under latency, power, and privacy constraints.\"}]","Leveraging Automated Machine Learning - Edge Computing for Video Understanding - A Thesis | PDF",1785807344,219,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"leveraging-automated-machine-learning-edge-computing-for-video-understanding-a-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/leveraging-automated-machine-learning-edge-computing-for-video-understanding-a-thesis/121868/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What challenges does the thesis address in video understanding and action recognition?","Question",{"text":75,"@type":76},"Developing effective video action recognition often requires extensive engineering to combine modules and optimize hyperparameters. Building an optimal end-to-end machine learning pipeline also demands substantial computer vision knowledge and effort from developers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is AutoVideo and how does it automate video action recognition?",{"text":80,"@type":76},"AutoVideo is an AutoML framework designed for automated video action recognition. It provides a modular, extendable pipeline, includes a comprehensive set of primitives, adds data-driven tuners to reduce pipeline tuning effort, and offers an easy-to-use GUI.",{"name":82,"@type":73,"acceptedAnswer":83},"How does BED enable object detection on edge devices?",{"text":84,"@type":76},"BED targets edge deployment by running object detection on the MAX78000 DNN accelerator. The system integrates on-device DNN inference with a camera for image acquisition and a screen for output exhibition to support real-world use under latency, power, and privacy constraints.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]