[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125889-en":3,"doc-seo-125889-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125889,1099523885336,"Violet","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Graph Machine Learning through the Lens of Bilevel Optimization","Bilevel optimization describes settings where the optimal solution of a lower-level energy problem provides input features to an upper-level objective. The work shows that many graph learning methods can be reformulated as bilevel optimization special cases or derived from simplifying assumptions. It derives a flexible family of energy functions that, with common descent steps, yield GNN message-passing layers while isolating approximation error sources. The framework also connects to knowledge graph embeddings, label propagation variants, and graph-regularized MLPs, supported by empirical results introducing BloomGML.","Graph Machine Learning through the Lens of Bilevel Optimization  \nAmber Yijia Zheng 1 Purdue University  \nTong He  \nAmazon Web Services  \nYixuan Qiu  \nShanghai University of Finance and Economics  \narXiv :2403 .04763v 1 [ cs .LG] 7 Mar 2024  \nMinjie Wang  \nAmazon Web Services  \nDavid Wipf  \nAmazon Web Services  \nAbstract  \nBilevel optimization refers to scenarios whereby the optimal solution of a lower-level energy function serves as input features to an upper-level objective of interest. These optimal features typically depend on tunable parameters of the lower-level energy in such a way that the entire bilevel pipeline can be trained end-to-end. Although not generally presented as such, this paper demonstrateshow a variety of graph learning techniques can be recast as special cases of bilevel optimization or simplifications thereof. In brief, building on prior work we first derive a more flexible class of energy functions that, when paired with various descent steps (e.g., gradient descent, proximal methods, momentum, etc.), form graph neural network (GNN) messagepassing layers; critically, we also carefully unpack where any residual approximation error lies with respect to the underlying constituent message-passing functions. We then probe several simplifications of this framework to derive close connections with non-GNN-based graph learning approaches, including knowledge graph embeddings, various forms of label propagation, and efficient graph-regularized MLP models. And finally, we present supporting empirical results that demonstrate the versatility of the proposed bilevel lens, which we refer to as BloomGML, referencing that BiLevel Optimization Offers More Graph Machine Learning. Our code is avail  \nable at [https://github.com/amberyzheng/](https://github.com/amberyzheng/)[ ](https://github.com/amberyzheng/) BloomGML. Let graph ML bloom.  \n1 Contribution during AWS Shanghai AI Lab internship.  \nProceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS) 2024, Valencia, Spain. PMLR: Volume 238 . Copyright 2024 by the author(s) .  \n1 Introduction  \nGraph machine learning covers a wide range of modeling tasks involving graph-structured data, where crossinstance/node dependencies are reflected by edges. As a classical example, label propagation (Zhou et al., 2003) and its many offshoots represent a semi-supervised learning approach whereby observed node labels are iteratively spread across the graph to unlabeled nodes. Ina related vein, various forms of graph-regularized MLP models (Ando and Zhang, 2006; Hu et al., 2021; Zhang et al., 2023) share node representations across edges to penalize misalignment with network structure during training. And as a third example more narrowly focused on certain heterogeneous graphs, non-parametric knowledge graph embedding (KGE) models (Bordes et al., 2013) produce node and relation-type embeddings that have been trained to differentiate factual knowledge triplets, composed of head and tail nodes connected by an edge relation, from spurious ones.  \nMore recently, graph neural networks (GNNs) have emerged as a promising class of predictive models for handling tasks such as node classification or link prediction (Kipf and Welling, 2016; Hamilton et al., 2017; Xu et al., 2019; Veličković et al., 2017; Zhou et al. , 2020) . Central to a wide variety of GNN architectures are layers composed of three components: a message function, which bundles information for sharing with neighbors, an aggregation function that fuses all the messages from neighbors, and an update function that computes the layer-wise output embedding for each node. Collectively, these functions enable the layer-bylayer propagation of information across the graph to facilitate downstream tasks (Kipf and Welling, 2016; Hamilton et al., 2017; Kearnes et al., 2016) .  \nIn the past, the graph ML models described above have primarily been motivated from diverse perspectives, without nece","cbCaifsSFqdxoIiR","https://ap.wps.com/l/cbCaifsSFqdxoIiR","pdf",1001510,5,1,23,"English","en",105,"# Abstract\n# Introduction\n## Background on graph machine learning\n## GNN layers and components\n## Unifying lens: bilevel optimization\n# Contributions Overview","[{\"question\":\"What does bilevel optimization mean in the context of graph machine learning?\",\"answer\":\"The lower-level objective is optimized first, and its solution becomes input features for an upper-level loss. This structure can be trained end-to-end and used to reinterpret graph learning architectures.\"},{\"question\":\"How does the paper connect graph neural networks to bilevel optimization?\",\"answer\":\"It derives energy functions that, combined with descent steps like gradient descent or proximal methods, produce GNN message-passing layers. It also analyzes where residual approximation error lies relative to the underlying message-passing functions.\"},{\"question\":\"What is BloomGML and what does it unify?\",\"answer\":\"BloomGML is a broadly applicable bilevel optimization framework introduced to unify multiple graph learning paradigms. The paper shows close connections with knowledge graph embedding methods, label propagation approaches, and efficient graph-regularized MLP models.\"}]","Graph Machine Learning through the Lens of Bilevel Optimization | PDF",1785901855,58,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"graph-machine-learning-through-the-lens-of-bilevel-optimization","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/graph-machine-learning-through-the-lens-of-bilevel-optimization/125889/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What does bilevel optimization mean in the context of graph machine learning?","Question",{"text":77,"@type":78},"The lower-level objective is optimized first, and its solution becomes input features for an upper-level loss. This structure can be trained end-to-end and used to reinterpret graph learning architectures.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the paper connect graph neural networks to bilevel optimization?",{"text":82,"@type":78},"It derives energy functions that, combined with descent steps like gradient descent or proximal methods, produce GNN message-passing layers. It also analyzes where residual approximation error lies relative to the underlying message-passing functions.",{"name":84,"@type":75,"acceptedAnswer":85},"What is BloomGML and what does it unify?",{"text":86,"@type":78},"BloomGML is a broadly applicable bilevel optimization framework introduced to unify multiple graph learning paradigms. The paper shows close connections with knowledge graph embedding methods, label propagation approaches, and efficient graph-regularized MLP models.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]