[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128623-en":3,"doc-seo-128623-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128623,962084925636,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Optimal Distributed Learning with Multi-pass Stochastic Gradient Methods","Optimal distributed learning with multi-pass stochastic gradient methods is studied for nonparametric regression in a reproducing kernel Hilbert space (RKHS). The work analyzes distributed stochastic gradient methods using mini-batches and multiple passes, showing that optimal generalization error bounds can be preserved when the partition level is sufficiently small. The derived theory improves state-of-the-art results, including settings where the target regression function may not belong to the hypothesis space. It also establishes lower theoretical computational complexity than distributed kernel ridge regression and classic stochastic gradient methods.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nOptimal Distributed Learning with Multi-pass Stochastic Gradient Methods  \nJunhong Lin 1 Volkan Cevher 1  \nAbstract  \nWe study generalization properties of distributed algorithms in the setting of nonparametric regression over a reproducing kernel Hilbert space (RKHS) . We investigate distributed stochastic gradient methods (SGM), with mini-batches and multi-passes over the data. We show that optimal generalization error bounds can be retained for distributed SGM provided that the partition level is not too large. Our results are superior to the state-of-the-art theory, covering the cases that the regression function may not be in the hypothesis spaces. Particularly, our results show that distributed SGM has a smaller theoretical computational complexity, compared with distributed kernel ridge regression (KRR) and classic SGM.  \n1. Introduction  \nIn statistical learning theory, a set of N input-output pairs from an unknown distribution is observed. The aim is to learn a function which can be used to predict future outputs given the corresponding inputs. The quality of a predictor is often measured in terms of the mean-squared error. In this case, the conditional mean, which is called as the regression function, is optimal among all the measurable functions (Cucker & Zhou, 2007 ; Steinwart & Christmann, 2008) . In nonparametric regression problems, the properties of the function to be estimated are not known a priori. Nonparametric approaches, which can adapt their complexity to the problem at hand, are key to good results. Kernel methods isone of the most common nonparametric approaches to learning (Schlkopf & Smola, 2002 ; Shawe-Taylor & Cristianini, 2004) . It is based on choosing a RKHS as the hypothesis space in the design of learning algorithms. With an appropri-  \n1 Laboratory for Information and Inference Systems, ´Ecole Polytechnique Fdrale de Lausanne, Lausanne, Switzerland. Correspondence to: Junhong Lin \u003Cjunhong.lin@epﬂ.ch>, Volkan Cevher \u003Cvolkan.cevher@epﬂ.ch> .  \nProceedings of the 35 th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018 . Copyright 2018 by the author(s) . This paper is a short report of“[https://arxiv.org/pdf/1801.07226.pdf](https://arxiv.org/pdf/1801.07226.pdf)”, submitted to Arxiv on January 23, 2018 .  \nate reproducing kernel, RKHS can be used to approximate any smooth function.  \nThe classical algorithms to perform learning task are regularized algorithms, such as KRR, kernel principal component regression (KPCR), and more generally, spectral regularization algorithms (SRA) . From the point of view of inverse problems, such approaches amount to solving an empirical, linear operator equation with the empirical covariance operator replaced by a regularized one (Engl et al., 1996 ; Bauer et al., 2007 ; Gerfo et al., 2008) . Here, the regularization term is used for controlling the complexity of the solution to against over-ﬁtting and for ensuring best generalization ability. Statistical results on generalization error had been developed in (Smale & Zhou, 2007 ; Caponnetto & De Vito, 2007) for KRR and in (Caponnetto, 2006 ; Bauer et al., 2007) for SRA.  \nAnother type of algorithms to perform learning tasks is based on iterative procedure (Engl et al., 1996) . In this kind of algorithms, an empirical objective function is optimized in an iterative way with no explicit constraint or penalization, and the regularization against overﬁtting is realized by early-stopping the empirical procedure. Statistical results on generalization error and the regularization roles of the number of iterations/passes have been investigated in (Zhang & Yu, 2005 ; Yao et al., 2007) for gradient methods (GM, also known as Landweber algorithm in inverse problems), in (Caponnetto, 2006 ; Bauer et al.,","cbCaiotaIdeFacOa","https://ap.wps.com/l/cbCaiotaIdeFacOa","pdf",694932,2,1,27,"English","en",105,"# Introduction\n## Nonparametric regression and RKHS\n## Regularized vs iterative learning algorithms\n## Computational motivation for distributed learning\n## Distributed stochastic gradient methods (multi-pass, mini-batches)","[{\"question\":\"What setting does the paper consider for the learning problem?\",\"answer\":\"It considers nonparametric regression over a reproducing kernel Hilbert space (RKHS) using distributed stochastic gradient methods.\"},{\"question\":\"Under what condition can optimal generalization error bounds be retained?\",\"answer\":\"The results show generalization bounds remain optimal as long as the partition level is not too large.\"},{\"question\":\"How do the proposed distributed stochastic gradient methods compare with kernel ridge regression?\",\"answer\":\"The paper shows distributed SGM can achieve smaller theoretical computational complexity than distributed kernel ridge regression (KRR) and classic SGM.\"}]","Optimal Distributed Learning with Multi-pass Stochastic Gradient Methods | PDF",1786002162,68,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"optimal-distributed-learning-with-multi-pass-stochastic-gradient-methods","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/optimal-distributed-learning-with-multi-pass-stochastic-gradient-methods/128623/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What setting does the paper consider for the learning problem?","Question",{"text":76,"@type":77},"It considers nonparametric regression over a reproducing kernel Hilbert space (RKHS) using distributed stochastic gradient methods.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Under what condition can optimal generalization error bounds be retained?",{"text":81,"@type":77},"The results show generalization bounds remain optimal as long as the partition level is not too large.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the proposed distributed stochastic gradient methods compare with kernel ridge regression?",{"text":85,"@type":77},"The paper shows distributed SGM can achieve smaller theoretical computational complexity than distributed kernel ridge regression (KRR) and classic SGM.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]