[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125538-en":3,"doc-seo-125538-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125538,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Fair Data Representation for Machine Learning at the Pareto Frontier","As machine learning increasingly drives everyday decision-making, ensuring fairness in upstream data processing becomes essential. The work introduces a pre-processing algorithm for fair data representation, showing how L2 (P)-objective supervised learning yields estimators of the Pareto frontier between prediction error and statistical disparity. Using optimal affine transport and Wasserstein barycenter post-processing, it links Wasserstein geodesics to the Pareto frontier, and numerical experiments highlight composability, privacy protection, computational efficiency in high dimensions, and relevance to fairness in L2-objective unsupervised learning.","arXiv :2201 .00292v2 [ stat .ML] 3 Nov 2022  \nFair Data Representation for Machine Learning at the Pareto Frontier  \nShizhou Xu [shzxu@ucdavis.edu](shzxu@ucdavis.edu)  \nDepartment of Mathematics University of California Davis Davis, CA 95616-5270, USA  \nThomas Strohmer [strohmer@math.ucdavis.edu](strohmer@math.ucdavis.edu)  \nDepartment of Mathematics  \nCenter of Data Science and Arti􀀌cial Intelligence Research University of California Davis  \nDavis, CA 95616-5270, USA  \nAbstract  \nAs machine learning powered decision-making becomes increasingly important in our daily lives, it is imperative to strive for fairness in the underlying data processing. We propose a pre-processing algorithm for fair data representation via which L2 (P)-objective supervised learning results in estimations of the Pareto frontier between prediction error and statistical disparity. Particularly, the present work applies the optimal a􀀎ne transport to approach the post-processing Wasserstein barycenter characterization of the optimal fair L2-objective supervised learning via a pre-processing data deformation. Furthermore, we show that the Wasserstein geodesics from learning outcome marginals to their barycenter characterizes the Pareto frontier between L2-loss and total Wasserstein distance among the marginals. Numerical simulations underscore the advantages: (1) the pre-processing step is compositive with arbitrary L2-objective supervised learning methods and unseen data; (2) the fair representation protects data privacy by preventing access to the sensitive information; (3) the optimal a􀀎ne maps are computationally e􀀎cient on high-dimensional data;  \n(4) experimental results shed light on the fairness of L2-objective unsupervised learning via the proposed fair data representation.  \nKeywords: statistical parity, Wasserstein barycenter, Wasserstein geodesics, optimal a􀀎ne transport, conditional expectation estimation  \n1. Introduction  \nOur society is increasingly in􀀍uenced by arti􀀌cial intelligence as decision-making processes become more reliant on statistical inference and machine learning. The potentially signi􀀌cant long-term impact from sequences of automated (facilitate of) decision-making has brought large concerns about bias and discrimination in machine learning [3, 34] . Machine learning based on unbiased algorithms can naturally inherit the historical biases that exist in data and hence reinforce the bias via automated decision-making process [9] .  \nOne straightforward partial remedy is to exclude the sensitive variables from the data set used in the learning and decision process. But such exclusion merely eliminates disparate treatment, which refers to direct discrimination, and leaves disparate impact, which refers to unintended or indirect discrimination, remaining in both data and learning outcome  \n© Shizhou Xu and Thomas Strohmer.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/) .  \nXu and Strohmer  \n[17] . Examples of the legal doctrine of disparate impact include Griggs v. Duke Powers Co.  \n[27] and Ricci v. DeStefano [28], where the decision is based on factors that are strongly correlated to race, such as intelligence quali􀀌cation in the former and the racially disproportionate test result in the latter, are ruled illegal by the US supreme court. As a result, along with the trending development of automated decision-making, the need for more sophisticated but practical techniques has made fairness in machine learning an important research area [30] .  \nTwo important but potentially con􀀍icting goals of fair machine learning are statistical parity (one of the most important de􀀌nitions of group fairness), which aims for similarity in predictions conditioned on sensitive information, and individual fairness, which aims for similar treatment of similar individuals regardless of the sensitive information. The present work targets statistical parity because it is closely","cbCaiiX2Ls5jZrHG","https://ap.wps.com/l/cbCaiiX2Ls5jZrHG","pdf",4562792,1,57,"English","en",105,"# Introduction\n# Fairness goals and related research\n# Statistical parity and disparate impact\n# Approaches to fair machine learning\n## Pre-processing\n## In-processing\n## Post-processing","[{\"question\":\"What problem does the paper address in machine learning fairness?\",\"answer\":\"It targets fairness in the data processing underlying machine learning decisions, focusing on how disparity and error trade off in the learned representation and outcomes.\"},{\"question\":\"How does the proposed method relate to the Pareto frontier?\",\"answer\":\"It constructs a pre-processing step so that L2 (P)-objective supervised learning estimates the Pareto frontier between prediction error and statistical disparity.\"},{\"question\":\"What role do Wasserstein barycenters and geodesics play?\",\"answer\":\"The paper uses optimal affine transport to connect a pre-processing deformation with the Wasserstein barycenter characterization of optimal fair learning; Wasserstein geodesics from outcome marginals to their barycenter are used to characterize the Pareto frontier.\"}]","Fair Data Representation for Machine Learning at the Pareto Frontier | PDF",1785899730,144,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"fair-data-representation-for-machine-learning-at-the-pareto-frontier","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fair-data-representation-for-machine-learning-at-the-pareto-frontier/125538/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in machine learning fairness?","Question",{"text":75,"@type":76},"It targets fairness in the data processing underlying machine learning decisions, focusing on how disparity and error trade off in the learned representation and outcomes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method relate to the Pareto frontier?",{"text":80,"@type":76},"It constructs a pre-processing step so that L2 (P)-objective supervised learning estimates the Pareto frontier between prediction error and statistical disparity.",{"name":82,"@type":73,"acceptedAnswer":83},"What role do Wasserstein barycenters and geodesics play?",{"text":84,"@type":76},"The paper uses optimal affine transport to connect a pre-processing deformation with the Wasserstein barycenter characterization of optimal fair learning; Wasserstein geodesics from outcome marginals to their barycenter are used to characterize the Pareto frontier.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]