[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118927-en":3,"doc-seo-118927-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118927,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Principled and Efficient Bilevel Optimization for Machine Learning - Doctor of Philosophy Dissertation","Automatic differentiation (AD) underpins most modern machine learning libraries by enabling efficient derivative computation from programs. Gradient-based learning often relies on exact gradients, but in meta-learning or hyperparameter optimization exact computation can be infeasible or too costly, motivating bilevel optimization. This work develops efficient gradient-based bilevel algorithms, proving convergence rates for approximating the hypergradient using smooth objectives and lower-level fixed points of contraction maps. Results cover deterministic and stochastic settings and include an efficient implementation.","Principled and Efficient Bilevel Optimization for Machine Learning  \nRiccardo Grazzi  \nA dissertation submitted in partial fulfillment of the requirements for the degree of  \nDoctor of Philosophy  \nof  \nUniversity College London.  \nDepartment of Computer Science  \nUniversity College London  \nOctober 2, 2023  \n2  \nI, Riccardo Grazzi, confirm that the work presented in this thesis is my own. Where information has been derived from other sources, I confirm that this has been indicated in the work.  \nAbstract  \nAutomatic differentiation (AD) is a core element of most modern machine learning libraries that allows to efficiently compute derivatives of a function from the corresponding program. Thanks to AD, machine learning practitioners have tackled increasingly complex learning models, such as deep neural networks with up to hundreds of billions of parameters, which are learned using the derivative (or gradient) of a loss function with respect to those parameters. While in most cases gradientscan be computed exactly and relatively cheaply, in others the exact computation is either impossible or too expensive and AD must be used in combination with approximation methods. Some of these challenging scenarios arising for example in meta-learning or hyperparameter optimization, can be framed as bilevel optimization problems, where the goal is to minimize an objective function that is evaluated by first solving another optimization problem, the lower-level problem. In this work, we study efficient gradient-based bilevel optimization algorithms for machine learning problems. In particular, we establish convergence rates for some simple approaches to approximate the gradient of the bilevel objective, namely the hypergradient, when the objective is smooth and the lower-level problem consists in finding the fixed point of a contraction map. Leveraging such results, we also prove that the projected inexact hypergradient method achieves a (near) optimal rate of convergence. We establish these results for both the deterministic and stochastic settings. Additionally, we provide an efficient implementation of the methods studied and perform several numerical experiments on hyperparameter optimization, meta-learning, datapoisoning and equilibrium models, which show that our theoretical results are good indicators of the performance in practice.  \nImpact Statement  \nGradient-based bilevel optimization methods have recently started to be more popular in machine learning research, while they are still rarely employed in industrial applications. The bilevel formulation allows to cover a wide variety of problems, but the complexity and high time and memory cost of gradient-based bilevel methods make them often less appealing than simpler and cheaper alternatives based on heuristics or tailored to specific applications. For these reasons, we foresee that this work will affect almost exclusively the academic world, at least until either a suitable industrial application is found or bilevel methods become easier to set up and/or cheaper to run.  \nThis thesis can be seen as an effort to build a quantitative theoretical foundation for gradient-based bilevel optimization methods. The scale of modern bilevel optimization problems combined with the complexity of gradient-based methods used to solve them makes it difficult and expensive to perform comprehensive experimental evaluations, and our analysis provides an alternative way to compare different bilevel methods in terms of iteration and sample complexity. This may be useful to researchers and practitioners for designing methods which are provably better than established ones, but also to better understand strength and weaknesses of each method and when it is suitable to apply.  \nThe experimental part of this work validates the theory on different small and medium scale problems in a variety of scenarios, and is accompanied by open-source code designed to be (and that has been) used by researchers to deve","cbCaiu0Qji2R4SoO","https://ap.wps.com/l/cbCaiu0Qji2R4SoO","pdf",9971450,1,196,"English","en",105,"# Abstract\n## Impact Statement\n## Acknowledgements\n## Dissertation Metadata","[{\"question\":\"What problem does the dissertation address in machine learning?\",\"answer\":\"It studies efficient gradient-based bilevel optimization methods for machine learning tasks, especially when exact gradients are hard or expensive to compute.\"},{\"question\":\"How does the work approximate the bilevel objective gradient?\",\"answer\":\"It focuses on approximating the hypergradient, analyzing convergence when the bilevel objective is smooth and the lower-level problem is the fixed point of a contraction map.\"},{\"question\":\"What evidence supports the theoretical results?\",\"answer\":\"The dissertation includes an efficient implementation and numerical experiments across hyperparameter optimization, meta-learning, data poisoning, and equilibrium models to validate the theory in practice.\"}]","Principled and Efficient Bilevel Optimization for Machine Learning - Doctor of Philosophy Dissertation | PDF",1785720991,494,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"principled-and-efficient-bilevel-optimization-for-machine-learning-doctor-of-philosophy-dissertation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/principled-and-efficient-bilevel-optimization-for-machine-learning-doctor-of-philosophy-dissertation/118927/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the dissertation address in machine learning?","Question",{"text":76,"@type":77},"It studies efficient gradient-based bilevel optimization methods for machine learning tasks, especially when exact gradients are hard or expensive to compute.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the work approximate the bilevel objective gradient?",{"text":81,"@type":77},"It focuses on approximating the hypergradient, analyzing convergence when the bilevel objective is smooth and the lower-level problem is the fixed point of a contraction map.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence supports the theoretical results?",{"text":85,"@type":77},"The dissertation includes an efficient implementation and numerical experiments across hyperparameter optimization, meta-learning, data poisoning, and equilibrium models to validate the theory in practice.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]