[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120242-en":3,"doc-seo-120242-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120242,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Generalization in Machine Learning through Information-Theoretic Lens - Dissertation Abstract","This thesis investigates generalization theory in machine learning using an information-theoretic framework. It derives novel information-theoretic generalization bounds by analyzing models trained with stochastic gradient descent (SGD), via an auxiliary weight process and an approximation of SGD using stochastic differential equations (SDE). The results characterize phenomena such as epoch-wise double descent of gradient dispersion under noisy labels, and motivate new regularization methods including dynamic gradient clipping and Gaussian model perturbation. The framework extends beyond SGD to black-box algorithms, and yields bounds for unsupervised domain adaptation (UDA).","Generalization in Machine Learning through Information-Theoretic Lens  \nby  \nZiqiao Wang  \nThesis submitted to the University of Ottawa in partial fulfillment of the requirements for the degree of Doctorate in Philosophy  \nin  \nElectrical and Computer Engineering  \nSchool of Electrical Engineering and Computer Science Faculty of Engineering  \nUniversity of Ottawa  \n© Ziqiao Wang, Ottawa, Canada, 2024  \nExamining Committee Membership  \nThe following served on the Examining Committee for this thesis. The decision of the Examining Committee is by majority vote.  \nSupervisor: Yongyi Mao  \nProfessor  \nSchool of Electrical Engineering and Computer Science University of Ottawa  \nInternal Member: Maia Fraser  \nAssociate Professor  \nDepartment of Mathematics and Statistics  \nUniversity of Ottawa  \nInternal Member: Tommaso R. Cesari  \nAssistant Professor  \nSchool of Electrical Engineering and Computer Science University of Ottawa  \nCarleton Member: James R. Green  \nProfessor  \nDepartment of Systems and Computer Engineering  \nCarleton University  \nExternal Member: Ashish Khisti  \nProfessor  \nDepartment of Electrical and Computer Engineering University of Toronto  \nAuthor’s Declaration  \nI hereby declare that I am the sole author of this thesis. This is a true copy of the thesis, including any required final revisions, as accepted by my examiners.  \nI understand that my thesis may be made electronically available to the public.  \nAbstract  \nIn this thesis, we utilize an information-theoretic framework to investigate generalization theory in machine learning, a critical area of research today. Specifically, we develop novel information-theoretic generalization bounds for machine learning algorithms. First, we apply information-theoretic analysis for models trained using stochastic gradient descent (SGD) . We do so by invoking an auxiliary weight process and by approximating SGD using stochastic differential equations (SDE) . Our analysis reveals intriguing phenomena such as epoch-wise double descent of gradient dispersion when trained with noisy labels. We also use our bounds to design new regularization techniques, including dynamic gradient clipping and Gaussian model perturbation, that can improve generalization performance. Furthermore, our framework is not limited to SGD-based algorithms; we also derive new information-theoretic bounds for any black-box learning algorithm, which are tighter than previous results based on the same settings. In addition, we apply our analysis to unsupervised domain adaptation (UDA), obtaining generalization bounds for two notions of the generalization error. Our algorithm-dependent bounds enable us to design new regularization techniques that can boost the performance of domain adaptation algorithms. Finally, we combine the stability-based generalization analysis with our information-theoretic analysis to derive novel generalization bounds, which explain the generalization in cases where previous information-theoretic bounds have fallen short.  \nTo Mom, Dad, and Xinyue  \nAcknowledgements  \nThe past several years of my PhD career have been a wonderful journey. This journey has been supported tremendously from people I have encountered. Before I delve into the technical content of my thesis, I would like to thank all of them.  \nFirst and foremost, I would like to thank my amazing advisor, Yongyi Mao, for being the best type of PhD adviser any student could wish for. He is a brilliant researcher and always is full of enthusiasm for research. His way of thinking, in-depth knowledge, passion, and optimism guided me throughout my PhD journey. He has cultivated in me a taste for fundamental research problems, a gift that I believe I could never have acquired without his mentorship. I vividly recall one of our early conversations, during my time as a master’s student, when we discussed artificial intelligence and machine learning. He remarked,“Communication is human encoding and human decoding; AI is God encodi","cbCaisvrCMiDxHYM","https://ap.wps.com/l/cbCaisvrCMiDxHYM","pdf",46089007,1,243,"English","en",105,"# Abstract\n## Information-theoretic framework\n## SGD analysis and bounds\n## Double descent under noisy labels\n## Regularization methods\n## Black-box generalization bounds\n## Unsupervised domain adaptation (UDA)\n## Stability plus information-theoretic analysis","[{\"question\":\"What is the main contribution of the thesis regarding generalization in machine learning?\",\"answer\":\"The thesis develops a unified information-theoretic framework and derives novel information-theoretic generalization bounds for machine learning algorithms.\"},{\"question\":\"How does the thesis analyze SGD in its information-theoretic approach?\",\"answer\":\"It uses an auxiliary weight process and approximates SGD with stochastic differential equations (SDE) to obtain the generalization bounds.\"},{\"question\":\"What new regularization techniques does the thesis propose, and what are their goals?\",\"answer\":\"It proposes dynamic gradient clipping and Gaussian model perturbation to improve generalization performance, including when adapting across domains.\"}]","Generalization in Machine Learning through Information-Theoretic Lens - Dissertation Abstract | PDF",1785728949,612,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"generalization-in-machine-learning-through-information-theoretic-lens-dissertation-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/generalization-in-machine-learning-through-information-theoretic-lens-dissertation-abstract/120242/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main contribution of the thesis regarding generalization in machine learning?","Question",{"text":75,"@type":76},"The thesis develops a unified information-theoretic framework and derives novel information-theoretic generalization bounds for machine learning algorithms.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis analyze SGD in its information-theoretic approach?",{"text":80,"@type":76},"It uses an auxiliary weight process and approximates SGD with stochastic differential equations (SDE) to obtain the generalization bounds.",{"name":82,"@type":73,"acceptedAnswer":83},"What new regularization techniques does the thesis propose, and what are their goals?",{"text":84,"@type":76},"It proposes dynamic gradient clipping and Gaussian model perturbation to improve generalization performance, including when adapting across domains.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]