[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160252-en":3,"doc-seo-160252-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160252,962084931830,"Jacob","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Improved Distributed Principal Component Analysis","Improved Distributed Principal Component Analysis studies distributed computing where multiple servers each hold a subset of points and jointly compute functions over the union. The work focuses on principal component analysis (PCA), aiming to find a low-dimensional affine subspace capturing maximum variance, and extends approximate PCA to downstream tasks such as k-means clustering and low-rank approximation. The proposed distributed PCA algorithms optimize communication cost and computational efficiency for a target accuracy. Experiments on real-world data show order-of-magnitude speedups while maintaining communication and only negligible degradation in solution quality. Some methods also provide transformation guarantees for subspace embeddings that are independent of success probability.","Improved Distributed Principal Component Analysis  \nMaria-Florina Balcan  \nSchool of Computer Science Carnegie Mellon University [ninamf@cs.cmu.edu](ninamf@cs.cmu.edu)  \nVandana Kanchanapally  \nSchool of Computer Science Georgia Institute of Technology [vvandana@gatech.edu](vvandana@gatech.edu)  \nYingyu Liang  \nDepartment of Computer Science Princeton University [yingyul@cs.princeton.edu](yingyul@cs.princeton.edu)  \nDavid Woodruff  \nAlmaden Research Center IBM Research [dpwoodru@us.ibm.com](dpwoodru@us.ibm.com)  \nAbstract  \nWe study the distributed computing setting in which there are multiple servers, each holding a set of points, who wish to compute functions on the union of their point sets. A key task in this setting is Principal Component Analysis (PCA), in which the servers would like to compute a low dimensional subspace capturing as much of the variance of the union of their point sets as possible. Given a procedure for approximate PCA, one can use it to approximately solve problems such as k-means clustering and low rank approximation. The essential properties of an approximate distributed PCA algorithm are its communication cost and computational efﬁciency for a given desired accuracy in downstream applications. We give new algorithms and analyses for distributed PCA which lead to improved communication and computational costs for k-means clustering and related problems.  \nOur empirical study on real world data shows a speedup of orders of magnitude, preserving communication with only a negligible degradation in solution quality.  \nSome of these techniques we develop, such as a general transformation from a constant success probability subspace embedding to a high success probability subspace embedding with a dimension and sparsity independent of the success probability, may be of independent interest.  \n1 Introduction  \nSince data is often partitioned across multiple servers [20, 7, 18], there is an increased interest in computing on it in the distributed model. A basic tool for distributed data analysis is Principal Component Analysis (PCA) . The goal of PCA is to ﬁnd an r-dimensional (afﬁne) subspace that captures as much of the variance of the data as possible. Hence, it can reveal low-dimensional structure in very high dimensional data. Moreover, it can serve as a preprocessing step to reduce the data dimension in various machine learning tasks, such as Non-Negative Matrix Factorization (NNMF) [15] and Latent Dirichlet Allocation (LDA) [3] .  \nIn the distributed model, approximate PCA was used by Feldman et al. [9] for solving a number of shape ﬁtting problems such as k-means clustering, where the approximation is in the form of acoreset, and has the property that local coresets can be easily combined across servers into a global coreset, thereby providing an approximate PCA to the union of the data sets. Designing small coresets therefore leads to communication-efﬁcient protocols. Coresets have the nice property that their size typically does not depend on the number n of points being approximated. A beautiful property of the coresets developed in [9] is that for approximate PCA their size also only depends linearly on the dimension d, whereas previous coresets depended quadratically on d [8] . This gives the best known communication protocols for approximate PCA and k-means clustering.  \nDespite this recent exciting progress, several important questions remain. First, can we improve the communication further as a function of the number of servers, the approximation error, and other parameters of the downstream applications (such as the number k of clusters in k-means clustering)?  \nSecond, while preserving optimal or nearly-optimal communication, can we improve the computational costs of the protocols? We note that in the protocols of Feldman et al. each server has to run a singular value decomposition (SVD) on her local data set, while additional work needs to be performed to combine the outputs of each serve","cbCaiiKCnMjAQEul","https://ap.wps.com/l/cbCaiiKCnMjAQEul","pdf",451114,1,9,"English","en",105,"# Abstract\n# Introduction\n## Distributed data model\n## Approximate PCA and related objectives\n## Communication and computational challenges","[{\"question\":\"What is the distributed PCA setting studied in this paper?\",\"answer\":\"Multiple servers each store a local set of points and wish to compute over the union of all points. The goal is to produce a low-dimensional subspace that captures variance of the combined dataset.\"},{\"question\":\"How does approximate distributed PCA help downstream tasks like k-means?\",\"answer\":\"Given an approximate PCA procedure, the paper explains how it can be used to approximately solve problems such as k-means clustering and low-rank approximation. The effectiveness depends on communication and computational efficiency under target accuracy.\"},{\"question\":\"Which performance metrics does the paper emphasize for distributed PCA algorithms?\",\"answer\":\"The key properties are communication cost and computational efficiency for a desired accuracy in the downstream applications, including improvements that reduce communication as a function of servers and approximation parameters.\"}]","Improved Distributed Principal Component Analysis | PDF",1788052789,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improved-distributed-principal-component-analysis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improved-distributed-principal-component-analysis/160252/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the distributed PCA setting studied in this paper?","Question",{"text":75,"@type":76},"Multiple servers each store a local set of points and wish to compute over the union of all points. The goal is to produce a low-dimensional subspace that captures variance of the combined dataset.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does approximate distributed PCA help downstream tasks like k-means?",{"text":80,"@type":76},"Given an approximate PCA procedure, the paper explains how it can be used to approximately solve problems such as k-means clustering and low-rank approximation. The effectiveness depends on communication and computational efficiency under target accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"Which performance metrics does the paper emphasize for distributed PCA algorithms?",{"text":84,"@type":76},"The key properties are communication cost and computational efficiency for a desired accuracy in the downstream applications, including improvements that reduce communication as a function of servers and approximation parameters.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]