[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124290-en":3,"doc-seo-124290-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124290,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Fair Soft Clustering","Scholars in machine learning study whether learning models behave fairly, including clustering algorithms. This work investigates fair clustering in a probabilistic (soft) regime where each observation can belong to multiple clusters via assignment probabilities. The paper defines new probabilistic fairness metrics that extend non-probabilistic fairness frameworks, and presents an algorithm to obtain a fair probabilistic cluster solution from a fairlet decomposition. Experiments use a fair Gaussian mixture model on three real-world datasets to verify the approach.","Fair Soft Clustering  \nRune D. Kjærsgaard∗,1 Pekka Parviainen2 Saket Saurabh2,3  \nMadhumita Kundu2 Line K. H. Clemmensen 1  \n1 DTU Compute, Technical University of Denmark, Denmark  \n2 Department of Informatics, University of Bergen, Norway  \n3 Theoretical Computer Science Group, The Institute of Mathematical Sciences, India  \n∗ Correspondence to: [rdokj@dtu.dk](rdokj@dtu.dk)  \nAbstract  \nScholars in the machine learning community have recently focused on analyzing the fairness of learning models, including clustering algorithms. In this work we study fair clustering in a probabilistic (soft) setting, where observations may belong to several clusters determined by probabilities. We introduce new probabilistic fairness metrics, which generalize and extend existing non-probabilistic fairness frameworks and propose an algorithm for obtaining a fair probabilistic cluster solution from a data representation known as a fairlet decomposition. Finally, we demonstrate our proposed fairness metrics and algorithm by constructing a fair Gaussian mixture model on three real-world datasets. We achieve this by identifying balanced micro-clusters which minimize the distances induced by the model, and on which traditional clustering can be performed while ensuring the fairness of the solution.  \n1 INTRODUCTION  \nDecision making systems based on machine learning (ML) applications have demonstrated unwanted consequences as a result of biased data (Phillips et al., 2011; Z. Obermeyer and Mullainan, 2019) . This has fostered efforts towards artificial intelligence (AI) alignment, wherein ML systems are aligned with their intended  \nProceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS) 2024, Valencia, Spain. PMLR: Volume 238 . Copyright 2024 by the author(s) .  \nobjectives. This includes ensuring decisions are fair and do not show bias against or for certain population sub-groups. Many of these fairness interventions are based on the Disparate Impact (DI) doctrine (Rutherglen, 1987), which prohibits discrimination between different groups of protected attributes such as race or sex. For clustering, this type of non-discrimination is denoted group-level fairness (Chhabra et al., 2021) .  \nClustering algorithms are an unsupervised ML approach used to partition a dataspace into clusters. These algorithms are widely used, particularly in settings where data labels are scarce. Here, clustering may be used as a feature engineering tool to supplement points with cluster assignments in an effort to increase expressive power of downstream models. If the underlying training data is unfair, this may propagate into the generated features and ultimately cause biased predictions. Fair clustering aims to prevent this.  \nThe topic of fairness for clustering was initiated in a seminal work by Chierichetti et al. (2017), which considered group-level fairness obtained by modifying the input data for traditional hard clustering algorithms like k-center and k-median. The literature on fair clustering is largely focused on such non-probabilistic algorithms, where point assignments are deterministic (Chhabra et al., 2021) . However, for a number of applications soft clustering is more appropriate. In our work, we consider group-level fair clustering in a probabilistic setting, where equal representation is ensured for protected groups in clusters found using soft clustering algorithms. As an example, a bank might use a dataset containing information about educational attainment and wages of individuals to train a model with the goal of identifying potential customers and offering them loans or credit opportunities. The bank then trains a soft clustering algorithm to group customers into low or high risk candidates, where the soft assignments  \ncould imply the probability (risk) that a given customer will default their loan. It should be pointed out, that a wage gap has been identified for women and people-of-color, who usual","cbCaitzGXf9e0kDp","https://ap.wps.com/l/cbCaitzGXf9e0kDp","pdf",7113951,1,12,"English","en",105,"# Introduction\n## Background on fairness in machine learning and clustering\n## Group-level fairness and the DI doctrine\n## Fair clustering in hard vs. soft (probabilistic) settings\n## Related work on probabilistic fairness formulations","[{\"question\":\"What does “fair soft clustering” mean in this paper?\",\"answer\":\"group-level fairness\"},{\"question\":\"What are the main contributions of the work?\",\"answer\":\"It also validates the metrics and algorithm with experiments on a fair Gaussian mixture model over three real-world datasets.\"},{\"question\":\"How do the experiments ensure both clustering quality and fairness?\",\"answer\":\"They identify balanced micro-clusters that minimize distances induced by the model, then perform traditional clustering on these micro-clusters while preserving group-level fairness in the solution.\"}]","Fair Soft Clustering | PDF",1785821411,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"fair-soft-clustering","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fair-soft-clustering/124290/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does “fair soft clustering” mean in this paper?","Question",{"text":75,"@type":76},"group-level fairness","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main contributions of the work?",{"text":80,"@type":76},"It also validates the metrics and algorithm with experiments on a fair Gaussian mixture model over three real-world datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the experiments ensure both clustering quality and fairness?",{"text":84,"@type":76},"They identify balanced micro-clusters that minimize distances induced by the model, then perform traditional clustering on these micro-clusters while preserving group-level fairness in the solution.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]