[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84073-en":3,"doc-seo-84073-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84073,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning","Generalization remains a central bottleneck in deep learning, where optimizers such as SGD frequently converge to sharp minima and thereby worsen performance on unseen data. Building on Sharpness-Aware Minimization (SAM), the work introduces Extragradient-Inspired Sharpness-Aware Minimization (EISAM), an optimizer that leverages an extragradient-style two-step update. A prediction step probes loss-landscape geometry, followed by a perturbation/correction step with a base optimizer. Results on benchmark datasets show EISAM outperforms SGD, Adam, and SAM in test accuracy and training efficiency, with theory tightening generalization bounds via flatter minima and reduced curvature.","arXiv :2607 .06 15 1v 1 [ cs .LG] 7 Jul 2026  \nLeveraging Extragradient for Effective Sharpness-Aware Minimization in Deep  \nLearning  \nYao Fu, Chunxia Zhang, Junmin Liu, Yihang Jin, Haishan Ye, and Yuanao Yang  \nAbstract—Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data. Building on Sharpness-Aware Minimization (SAM), for seeking flat minima associated with improved generalization, we propose the Extragradient-Inspired Sharpness-Aware Minimization (EISAM), a novel optimizer that enhances generalization via the extragradient technique. EISAM uses a two-step update process: a prediction step investigating the geometry of the loss landscape and a perturbation step that refines updates with a base optimizer. This approach achieves better generalization performance than SAM. Crucially, EISAM reduces sensitivity to the perturbation radius, enhancing robustness, and simplifying the tuning across diverse settings. Extensive experiments on benchmark datasets demonstrate that EISAM consistently outperforms SGD, Adaptive Moment Estimation (Adam), and SAM in test accuracy and training efficiency across various architectures. Theoretical analysis further confirms that EISAM tightens the  \ngeneralization bound by steering parameters toward flatter minima with reduced curvature. Accompanied by a thorough hyperparameter analysis, EISAM offers practical tuning guidance, establishing it as a robust, scalable, and broadly applicable optimization solution that advances both the theory and practice in deep learning.  \nIndex Terms—Deep neural networks, sharpness-aware minimization, extragradient method, excess risk analysis, generalization error.  \n~~ ~~ ✦ ~~ ~~  \n1 INTRODUCTION  \nOPTIMIZING deep neural networks to achieve better  \ngeneralization performance is a key challenge in deep learning. Traditional optimizers, such as SGD and Adam, have achieved significant success in training deep models, but they often converge to sharp minima of the loss function. Although the association between flat minima and generalization is controversial [1], studies have shown that sharp minima often lead to poor generalization on unseen data, while flat minima are associated with stronger generalization performance [2], [3] . Building on these insights, SAM [4] has emerged as an effective approach to seek flat minima by minimizing loss sharpness within a parameter neighborhood. Despite its benefits, SAM requires additional gradient computations at each step, resulting in approximately twice the computational cost compared to base optimizers like SGD and Adam, though this can be mitigated through faster convergence. Moreover, its performance is highly sensitive to the perturbation radius ρ, and the sharpness approximation may not accurately capture the loss landscape’s curvature, potentially yielding suboptimal generalization.    \n• Yao Fu, Junmin Liu, and Chunxia Zhang are with the School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an, 710049, China (Y. Fu and J. Liu also with SGIT AI Lab, State Grid Corporation of China, Xi’an, 710054, China). [E-mail: fyao56@stu.xjtu.edu.cn](E-mail: fyao56@stu.xjtu.edu.cn)., {junminliu, [cxzhang](cxzhang}@mail.xjtu.edu.cn)[}](cxzhang}@mail.xjtu.edu.cn)[@mail.xjtu.edu.cn](cxzhang}@mail.xjtu.edu.cn)  \n• Haishan Ye is with the School of Management, Xi’an Jiaotong University, Xi’an, 710049, China; SGIT AI Lab, State Grid Corporation of China, Xi’an, 710054, China. E-mail: hsye [cs@outlook.com](cs@outlook.com).  \n• Yihang Jin is with the School of Software, Xi’an Jiaotong University, Xi’an, 710049, [China. E-mail: yihangjin@stu.xjtu.edu.cn](China. E-mail: yihangjin@stu.xjtu.edu.cn).  \n• Yuanao Yang is with the State Key Laboratory of Multiphase Flow in Power Engineering, Xi’an Jiaotong University, Xi’an, 710049, China. E[mail: yya0407@stu.xjtu.e","cbCaivKOmZmE8QHl","https://ap.wps.com/l/cbCaivKOmZmE8QHl","pdf",27955678,4,1,24,"English","en",105,"# Introduction\n## Problem: sharp minima and generalization\n## Limitation of SAM\n# Proposed Method\n## Extragradient method overview\n## EISAM two-step update\n# Experimental and Theoretical Results","[{\"question\":\"What problem does EISAM aim to solve in deep learning optimization?\",\"answer\":\"EISAM targets poor generalization caused by convergence to sharp minima, which can lead to overfitting and reduced accuracy on unseen data.\"},{\"question\":\"How does EISAM modify Sharpness-Aware Minimization (SAM)?\",\"answer\":\"EISAM refines SAM by using a two-step process: a prediction step that probes local loss-landscape geometry, and a correction/perturbation step performed with a base optimizer to improve the update direction.\"},{\"question\":\"What advantages does EISAM provide over SGD, Adam, and SAM?\",\"answer\":\"EISAM achieves higher test accuracy and training efficiency in experiments, reduces sensitivity to the perturbation radius, and includes theoretical analysis showing tighter generalization bounds by steering toward flatter minima.\"}]",1784192505,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"leveraging-extragradient-for-effective-sharpness-aware-minimization-in-deep-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/leveraging-extragradient-for-effective-sharpness-aware-minimization-in-deep-learning/84073/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does EISAM aim to solve in deep learning optimization?","Question",{"text":75,"@type":76},"EISAM targets poor generalization caused by convergence to sharp minima, which can lead to overfitting and reduced accuracy on unseen data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EISAM modify Sharpness-Aware Minimization (SAM)?",{"text":80,"@type":76},"EISAM refines SAM by using a two-step process: a prediction step that probes local loss-landscape geometry, and a correction/perturbation step performed with a base optimizer to improve the update direction.",{"name":82,"@type":73,"acceptedAnswer":83},"What advantages does EISAM provide over SGD, Adam, and SAM?",{"text":84,"@type":76},"EISAM achieves higher test accuracy and training efficiency in experiments, reduces sensitivity to the perturbation radius, and includes theoretical analysis showing tighter generalization bounds by steering toward flatter minima.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]