[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127295-en":3,"doc-seo-127295-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127295,2336475104957,"Seraphina","https://ap-avatar.wpscdn.com/avatar/22000c4c6bd8a5076e1?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786593998035447633",8,"Research & Report","A Novel Assurance Procedure for Fair Data Augmentation in Machine Learning","Machine learning performance depends on sufficient data, yet dataset augmentation can introduce or reinforce bias when synthetic samples do not respect fairness requirements. This work proposes a similarity network representation where each data point becomes a node and new synthetic points are generated near existing ones. A vector label propagation method, guided by an exponential kernel for adaptive link weights, labels synthetic samples accurately. The procedure reduces reliance on sensitive features without excluding them, supporting fairness and maintaining data variation. Implemented in a big data ecosystem, it enables continuous evaluation as domains evolve.","A Novel Assurance Procedure for Fair Data Augmentation in Machine Learning  \nSamira Maghool* , Paolo Ceravolo and Filippo Berto University of Milan, Department of Computer Science  \nAbstract  \nIn addressing the limited availability of data for predictive purposes with machine learning, we are concerned with potential biases arising from dataset augmentation. Despite advanced algorithms to generate synthetic data that can preserve the original data distribution, challenges remain, includingthe risk of perpetuating social biases. Our approach uses a similarity network representation that treats each data point as a node and strategically generates synthetic points near it. A vector label propagation algorithm, complemented by an exponential kernel for adjusting link weights, accurately labels these synthetic points. The primary goal is to reduce the system’s dependence on sensitive features without excluding them, thereby avoiding the risk of exacerbating biases or reducing data variation. Implemented in a big data ecosystem, our methodology enables continuous evaluation in an evolving domain, effectively addressing the challenges of data scarcity with a fairness-aware approach.  \nKeywords  \nMachine Learning, Fairness, Similarity Network, Data Augmentation  \n1. Introduction  \nThe widespread adoption of Machine Learning (ML) technologies across industries has ushered in a new era of data-driven decision-making [1] . While ML promises to increase efficiency and productivity, its application in decision-making processes presents a number of challenges, ranging from performance to regulatory compliance [2] . Regulatory frameworks, such as the European Union’s proposed Artificial Intelligence Act, emphasize the importance of fairness and accountability [3] . Overcoming these challenges requires industries to establish comprehensive testing frameworks that evaluate ML models’ performance, reliability, and generalization across multiple scenarios [4] . But developing frameworks for regulatory compliance is a complex task. As regulations evolve and new data becomes available, industries must establish mechanisms for continuously monitoring and updating ML models [5] . The different techniques used by designers to achieve specific properties in ML systems can conflict with each other. Fairness may come at the expense of accuracy, accuracy at the expense of transparency, and privacy compliance at the expense of explainability [6, 7] .  \nIn this paper, our contribution is to explore the delicate balance between data augmentation and fairness in tabular data, with the goal of developing a solution suitable for integration into  \nAIEB 2024: Workshop on Implementing AI Ethics through a Behavioural Lens | co-located with ECAI 2024, Santiago de Compostela, Spain  \n* Corresponding author.  \n$ [samira.maghool@unimi.it](samira.maghool@unimi.it) (S. Maghool); [paolo.ceravolo@unimi.it](paolo.ceravolo@unimi.it) (P. Ceravolo); [filippo.berto@unimi.it](filippo.berto@unimi.it) (F. Berto)  \n􀀚 0000-0001-8310-2050 (S. Maghool); 0000-0002-4519-0173 (P. Ceravolo); 0000-0002-2720-608X (F. Berto)  \n © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0) .  \n1  \nCEUR ~~  ~~[Workshop](Workshop ceur-ws.org)[ ceur-ws.org](Workshop ceur-ws.org)[ ](Workshop ceur-ws.org)[Proceedings](Proceedings ISSN 1613-0073)[ ISSN 1613-0073](Proceedings ISSN 1613-0073)   \na continuous assurance framework. ML model training is inherently data-hungry, requiring a significant amount of data for accurate results, often necessitating the expansion of the dataset through augmentation [8] . Traditional data augmentation techniques applied to tabular datasets focus on creating duplicates and ensuring their consistency with the data distribution. This is achieved by assigning values through random perturbations or by adhering to central tendency [9] . However, many established augmentation techniques do not exp","cbCaimnw2c8a8zpJ","https://ap.wps.com/l/cbCaimnw2c8a8zpJ","pdf",1453429,1,17,"English","en",105,"# Introduction\n## Data scarcity and fairness in augmentation\n## Regulatory context and continuous assurance needs\n## Bias, discrimination, and protected subgroups","[{\"question\":\"Why can data augmentation harm fairness in machine learning?\",\"answer\":\"Augmentation can perpetuate social biases already present in the original training data. If synthetic samples are generated without fairness considerations, biased patterns may be mirrored instead of mitigated.\"},{\"question\":\"What is the proposed mechanism for fair data augmentation?\",\"answer\":\"The method builds a similarity network where each data point is a node, then strategically generates synthetic points near existing ones. Vector label propagation with an exponential kernel assigns labels to synthetic points.\"},{\"question\":\"How does the approach avoid excluding sensitive features while improving fairness?\",\"answer\":\"The procedure aims to reduce dependence on sensitive features rather than remove them, helping avoid unintended consequences that can occur when features are ignored but bias remains in other correlated signals.\"}]","A Novel Assurance Procedure for Fair Data Augmentation in Machine Learning | PDF",1785938157,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-novel-assurance-procedure-for-fair-data-augmentation-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-novel-assurance-procedure-for-fair-data-augmentation-in-machine-learning/127295/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why can data augmentation harm fairness in machine learning?","Question",{"text":76,"@type":77},"Augmentation can perpetuate social biases already present in the original training data. If synthetic samples are generated without fairness considerations, biased patterns may be mirrored instead of mitigated.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the proposed mechanism for fair data augmentation?",{"text":81,"@type":77},"The method builds a similarity network where each data point is a node, then strategically generates synthetic points near existing ones. Vector label propagation with an exponential kernel assigns labels to synthetic points.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the approach avoid excluding sensitive features while improving fairness?",{"text":85,"@type":77},"The procedure aims to reduce dependence on sensitive features rather than remove them, helping avoid unintended consequences that can occur when features are ignored but bias remains in other correlated signals.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]