[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117296-en":3,"doc-seo-117296-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},117296,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","iDARTS - Differentiable Architecture Search with Stochastic Implicit Gradients","Differentiable Architecture Search (DARTS) drives neural architecture search through an efficient gradient-based bi-level optimization that alternates between inner model weights and outer architecture parameters in a weight-sharing supernet. A central scalability and quality issue is how to differentiate through the inner-loop optimization. This work derives DARTS hypergradient computation using the implicit function theorem, making it depend only on the inner-loop solution and not on the optimization path. It further introduces a stochastic hypergradient approximation, with theory and experiments showing convergence to stationary points and architectures that outperform baseline methods.","iDARTS: Differentiable Architecture Search with Stochastic Implicit Gradients  \nMiao Zhang 1 2 Steven Su 2 Shirui Pan 1 Xiaojun Chang 1 Ehsan Abbasnejad 3 Reza Haffari 1  \nAbstract  \nDifferentiable ARchiTecture Search (DARTS) has recently become the mainstream of neural architecture search (NAS) due to its efﬁciency and simplicity. With a gradient-based bi-level optimization, DARTS alternately optimizes the inner model weights and the outer architecture parameter in a weight-sharing supernet. A key challenge to the scalability and quality of the learned architectures is the need for differentiating through the inner-loop optimisation. While much has been discussed about several potentially fatal factors in DARTS, the architecture gradient, a.k.a. hypergradient, has received less attention. In this paper, we tackle the hypergradient computation in DARTS based on the implicit function theorem, making it only depends on the obtained solution to the innerloop optimization and agnostic to the optimization path. To further reduce the computational requirements, we formulate a stochastic hypergradient approximation for differentiable NAS, and theoretically show that the architecture optimization with the proposed method, named iDARTS, is expected to converge to a stationary point. Comprehensive experiments on two NAS benchmark search spaces and the common NAS search space verify the effectiveness of our proposed method.  \nIt leads to architectures outperforming, with large margins, those learned by the baseline methods.  \n1. Introduction  \nNeural Architecture Search (NAS) is an efﬁcient and effective method on automating the process of neural network design, with achieving remarkable success on image recognition (Tan & Le, 2019 ; Li et al., 2021b ; 2020), language modeling (Jiang et al., 2019), and other deep learning ap-  \n1Faculty of Information Technology, Monash University, Australia 2Faculty of Engineering and Information Technology, University of Technology Sydney, Australia 3Australian Institute for Machine Learning, University of Adelaide, Australia. Correspondence to: Shirui Pan \u003C[Shirui.Pan@monash.edu](Shirui.Pan@monash.edu) > .  \nProceedings of the 38 th International Conference on Machine Learning, PMLR 139, 2021 . Copyright 2021 by the author(s) .  \nplications (Ren et al., 2020 ; Cheng et al., 2020 ; Chen et al., 2019b ; Hu et al., 2021 ; Zhu et al., 2021 ; Ren et al., 2021) . The early NAS frameworks are devised via reinforcement learning (RL) (Pham et al., 2018) or evolutionary algorithm (EA) (Real et al., 2019) to directly search on the discrete space. To further improve the efﬁciency, a recently proposed Differentiable ARchiTecture Search (DARTS) (Liu et al., 2019) adopts the continuous relaxation to convert the operation selection problem into the continuous magnitude optimization for a set of candidate operations. By enabling the gradient descent for the architecture optimization, DARTS signiﬁcantly reduces the search cost to several GPU hours (Liu et al., 2019 ; Xu et al., 2020 ; Dong & Yang, 2019a) .  \nDespite its efﬁciency, more current works observe that DARTS is somewhat unreliable (Zela et al., 2020a ; Chen & Hsieh, 2020 ; Li & Talwalkar, 2019 ; Sciuto et al., 2019 ; Zhang et al., 2020c ; Li et al., 2021a ; Zhang et al., 2020b) since it does not consistently yield excellent solutions, performing even worse than random search in some cases. Zela et al. (2020a) attribute the failure of DARTS to its supernet training, with empirically observing that the instability of DARTS is highly correlated to the dominant eigenvalue of the Hessian matrix of the validation loss with respect to architecture parameters. On the other hand, Wang et al.(2021a) turn to the magnitude-based architecture selection process, who empirically and theoretically show the magnitude of architecture parameters does not necessarily indicate how much the operation contributes to the supernet's performance. Chen & Hsieh (2020) observe a precipit","cbCaikCq9hzYkzPn","https://ap.wps.com/l/cbCaikCq9hzYkzPn","pdf",454249,1,18,"English","en",105,"# Introduction\n## Background: Neural Architecture Search and DARTS\n## Key Challenge: Hypergradient Computation\n## Proposed Method: iDARTS\n## Contributions and Organization","[{\"question\":\"What problem does iDARTS address in DARTS?\",\"answer\":\"iDARTS targets the difficulty of computing the architecture hypergradient, i.e., differentiating through the inner-loop optimization in DARTS to improve scalability and architecture quality.\"},{\"question\":\"How does iDARTS compute the hypergradient?\",\"answer\":\"iDARTS derives the hypergradient using the implicit function theorem so the computation depends on the inner-loop solution and is independent of the optimization path.\"},{\"question\":\"How does iDARTS reduce computational cost?\",\"answer\":\"It approximates the needed inverse via a Neumann series and further introduces a stochastic hypergradient approximation to reduce the computational burden for differentiable NAS.\"}]",1785675045,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"idarts-differentiable-architecture-search-with-stochastic-implicit-gradients","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/idarts-differentiable-architecture-search-with-stochastic-implicit-gradients/117296/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does iDARTS address in DARTS?","Question",{"text":74,"@type":75},"iDARTS targets the difficulty of computing the architecture hypergradient, i.e., differentiating through the inner-loop optimization in DARTS to improve scalability and architecture quality.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does iDARTS compute the hypergradient?",{"text":79,"@type":75},"iDARTS derives the hypergradient using the implicit function theorem so the computation depends on the inner-loop solution and is independent of the optimization path.",{"name":81,"@type":72,"acceptedAnswer":82},"How does iDARTS reduce computational cost?",{"text":83,"@type":75},"It approximates the needed inverse via a Neumann series and further introduces a stochastic hypergradient approximation to reduce the computational burden for differentiable NAS.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]