[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85537-en":3,"doc-seo-85537-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85537,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints","Distributed stochastic minimax optimization under stochastic constraints is addressed with a first-order Softmax-Weighted Switching Gradient method designed for federated learning. With full client participation, the method attains oracle complexity O(ε−4) to meet a unified tolerance ε for both optimality gap and feasibility. The analysis extends to partial participation via a stochastic superiority assumption. Relaxing objective boundedness yields a tighter softmax hyperparameter lower bound, and a sharp O(log(1/δ)) high-probability convergence result. Experiments validate performance on Neyman–Pearson classification, fair classification, and federated safe reinforcement learning, using a stable primal-only switching alternative.","First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints  \nZhankun Luo, Antesh Upadhyay, Sang Bin Moon, Abolfazl Hashemi  \nSchool of Electrical and Computer Engineering, Purdue University {luo333, aantesh, moon182, [abolfazl}@purdue.edu](abolfazl}@purdue.edu)  \narXiv :2603 .05774v2 [ cs .LG] 13 Jul 2026  \nAbstract  \nThis paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints. We propose a novel first-order Softmax-Weighted Switching Gradient method tailored for federated learning. Under full client p˜articipation, our algorithm achieves the standard  \nO (ϵ−4) oracle complexity to satisfy a unified bound ϵ for both the optimality gap and feasibility tolerance. We extend our theoretical analysis to the practical partial participation regime by quantifying client sampling noise through a stochastic superiority assumption. Furthermore, by relaxing standard boundedness assumptions on the objective functions, we establish a strictly tighter lower bound for the softmax hyperparameter. We provide a unified error decomposition and establish a sharp O(log ~~1~~δ ) high-probability convergence guarantee. Ultimately, our framework demonstrates that a single-loop primal-only switching mechanism provides a stable alternative for optimizing worst-case client performance, effectively bypassing the hyperparameter sensitivity and convergence oscillations often encountered in traditional primaldual or penalty-based approaches. We verify the efficacy of our algorithm via experiment on the Neyman-Pearson (NP) classification, fair classification, and federated safe reinforcement learning tasks.  \n1 Introduction  \nFederated Learning (FL) aims to solve distributed optimization problems of the form  \nn  \nmwnΘ Xpifi (w), (1)  \ni=1  \nwhere Θ ⊆ Rd is a compact convex set, n is the number of clients, fi (w) := Eζ∼Di [fi (w,ζ)] denotes the local expected loss at client i and pi denote the probability weight of the client [McMahan et al., 2017, Kairouz and McMahan, 2021] . Under statistical heterogeneity, where local distributions {Di } are non-identical, this empirical risk minimization (ERM) objective inherently prioritizes average performance across clients [Li et al., 2020a, Mohri et al., 2019] . As a result, the learned model is biased toward dominant client distributions and may exhibit severely degraded performance on underrepresented or hard clients [Mohri et al., 2019, Hashimoto et al., 2018, Li et al., 2019] .  \nTo guarantee uniformly good performance across all devices, a powerful alternative is to frame the training process as a distributionally robust (or agnostic) optimization problem [Mohri et al., 2019, Deng et al., 2020, Duchi and Namkoong, 2021] . Instead of minimizing the average loss, the algorithm minimizes the maximum expected loss over a global set of adversarial weights  \nn  \nmwnΘ mλ X λifi (w), (2)  \ni=1  \nwhere Λ := 􀀈λ ∈ R : P λi = 1 􀀉 represents the probability weight assigned to each local client i. Intuitively, the inner maximization concentrates probability mass on the worst-performing clients. This recovers the equivalent minimax formulation minw∈Θ maxi∈I fi (w) with I :={1, 2 ,..., n} which directly enforces robustness to client heterogeneity.  \nExisting minimax formulations in federated settings typically optimize this worst-case loss in isolation. However, in many practical deployments, models must simultaneously satisfy strict client-wise operational requirements, such as fairness mandates, safety limits, resource budgets, or regulatory thresholds [Islamov et al., 2025b, Upadhyay et al., 2026] . Tracking n distinct stochastic constraints, i.e., gi (w) = Eζ∼Di [gi (w,ζ)] ≤ 0 , ∀i ∈ I, which encode client-specific operational constraints, is prohibitively expensive in federated environments, as it requires maintaining and synchro-  \nnizing n distinct dual variables across a network with intermittent cl","cbCaibYPogncfbyh","https://ap.wps.com/l/cbCaibYPogncfbyh","pdf",4693281,2,1,59,"English","en",105,"# Abstract\n# Introduction\n## Federated learning objective and heterogeneity\n## Distributionally robust minimax formulation\n## Stochastic minimax with stochastic constraints\n## Key challenges: non-smoothness, worst-case constraints, and coupling","[{\"question\":\"What optimization problem does the paper focus on?\",\"answer\":\"It studies distributed stochastic minimax optimization subject to stochastic constraints in federated learning, aiming to control both worst-case performance and constraint violations across heterogeneous clients.\"},{\"question\":\"What is the proposed algorithmic approach?\",\"answer\":\"The paper introduces a first-order Softmax-Weighted Switching Gradient method, tailored for federated learning, that uses a switching mechanism with a Softmax weighting to handle the worst-case structure.\"},{\"question\":\"What does the paper guarantee about complexity and convergence?\",\"answer\":\"Under full client participation, it achieves oracle complexity O(ε−4) to satisfy a unified bound ε for both optimality gap and feasibility. It also provides a sharp high-probability convergence guarantee with rate O(log(1/δ)).\"}]",1784204285,149,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"first-order-softmax-weighted-switching-gradient-method-for-distributed-stochastic-minimax-optimization-with-stochastic-constraints","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/first-order-softmax-weighted-switching-gradient-method-for-distributed-stochastic-minimax-optimization-with-stochastic-constraints/85537/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What optimization problem does the paper focus on?","Question",{"text":75,"@type":76},"It studies distributed stochastic minimax optimization subject to stochastic constraints in federated learning, aiming to control both worst-case performance and constraint violations across heterogeneous clients.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed algorithmic approach?",{"text":80,"@type":76},"The paper introduces a first-order Softmax-Weighted Switching Gradient method, tailored for federated learning, that uses a switching mechanism with a Softmax weighting to handle the worst-case structure.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper guarantee about complexity and convergence?",{"text":84,"@type":76},"Under full client participation, it achieves oracle complexity O(ε−4) to satisfy a unified bound ε for both optimality gap and feasibility. It also provides a sharp high-probability convergence guarantee with rate O(log(1/δ)).","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]