[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118692-en":3,"doc-seo-118692-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118692,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","A COMPARATIVE BENCHMARK OF FAIRNESS METRICS IN MACHINE LEARNING - Master’s Thesis","Machine learning increasingly supports high-stakes decisions such as recidivism prediction and loan approval, making errors and biases consequential for individuals and groups. This thesis investigates how the balance between accuracy and fairness changes across different standard models. Two sensitive datasets (COMPAS and Adult Income) are used to train and test nine binary classifiers. Models are evaluated with Demographic Parity Difference (DPD) and Equalized Odds Difference (EOD) alongside Accuracy and F1-score, showing architecture-dependent trade-offs.","Md Asif Shahariar  \nA COMPARATIVE BENCHMARK OF FAIRNESS METRICS IN MACHINE LEARNING  \nFaculty of Information Technology and Communication Sciences (ITC) Master’s thesis December 2025  \nABSTRACT  \nMd Asif Shahariar: A Comparative Benchmark of Fairness Metrics in Machine Learning  \nMaster’s thesis Tampere University  \nMaster’s Degree Program in Data Science December 2025  \nMachine learning is no longer limited to low-risk tasks, but it is also being used in critical decision-making systems. Consequently, machine learning has a direct effect on individuals’ lives in domains such as recidivism prediction (COMPAS) and loan approval. Thus, errors or biases generated by such algorithms may have severe consequences in the real world. Therefore, our expectations of ML are not limited to accuracy. In addition to making correct predictions, we expect the models to behave in an unbiased manner toward individuals and groups. Many mathematical definitions of fairness have been proposed by researchers, but there have been few large, systematic comparisons of fairness metrics applied to standard machine learning models. The main task of this thesis is to examine in detail how the balance or trade-off between accuracy and fairness occurs in different models.  \nThe study used two socially sensitive datasets, COMPAS recidivism and Adult Income. These two datasets were used to train and test nine machine learning models. The models used in this study include nine widely used binary classifiers, ranging from simple and interpretable algorithms (e.g., Logistic Regression and Support Vector Machine) to high-performance ensemble methods (e.g., Random Forest, XGBoost, and LightGBM) . Each model was carefully evaluated using Demographic Parity Difference (DPD) and Equalized Odds Difference (EOD) which are two important group fairness criteria, and traditional performance metrics including Accuracy and F1-score.  \nThe outcome of the experiment clearly shows a trade-off between accuracy and fairness despite the fact that the nature of this relationship varies greatly among the model architectures. Although ensemble algorithms including XGBoost and LightGBM achieved the highest predictive accuracy, they also had the largest fairness gaps. On the other hand, less complex models like Decision Trees and Multi-layer Perceptrons tended to be more fair and less accurate. One of the key findings in the hyperparameter sensitivity analysis is the fact that fairness is not fixed. It varied with hyperparameters for models like Logistic Regression and SVM, but was less responsive for others. Furthermore, statistical significance testing highlighted that while performance advantages were often robust, differences in fairness rankings  \nwere frequently statistically insignificant, cautioning against over-reliance on minor metric variations.  \nThis thesis develops a robust empirical foundation for fairness assessment in machine learning through broad benchmarking with both sensitivity analysis and statistical testing. This study serves as a methodological guide and a source of empirical evidence for practitioners and researchers, providing a framework for evaluating fairness based on evidence-based insights. These results indicate that fairness is not universal or transferable across models or datasets, it must be evaluated contextually, with awareness of model behavior and tuning sensitivity.  \nKeywords: Algorithmic Fairness, Fairness-Accuracy Trade-off, Demographic Parity, Equalized Odds, COMPAS, Adult Income, Hyperparameter Sensitivity, Statistical Significance Testing  \nThe originality of this thesis has been checked using the Turnitin Originality Check service.  \nUSE OF AI IN THESIS  \nI have utilised AI tools in my thesis:  \n□ No  \n⊠ Yes  \nThe AI tools utilised in my thesis, and their purposes, are described below: Names and versions of AI tools: ChatGPT 5.0, Gemini 3  \nPurpose of using AI tools: ChatGPT and Google Gemini were used as writingsupport tools during th","cbCailJHDKyWMdub","https://ap.wps.com/l/cbCailJHDKyWMdub","pdf",1505668,1,62,"English","en",105,"# Abstract\n## Datasets and models\n## Evaluation metrics\n## Accuracy-fairness trade-off results\n## Hyperparameter sensitivity and statistical testing\n## AI usage and responsibilities","[{\"question\":\"What problem does the thesis address in machine learning deployments?\",\"answer\":\"It addresses how fairness-related errors or biases in machine learning can harm individuals in real-world high-stakes decisions, beyond just achieving accuracy.\"},{\"question\":\"Which datasets and fairness criteria are used for benchmarking?\",\"answer\":\"The study uses COMPAS recidivism and Adult Income, evaluating models with Demographic Parity Difference (DPD) and Equalized Odds Difference (EOD) plus traditional performance metrics such as Accuracy and F1-score.\"},{\"question\":\"What main relationship is found between accuracy and fairness across models?\",\"answer\":\"A trade-off between accuracy and fairness is observed, but its strength depends on model architecture; ensemble models like XGBoost and LightGBM show higher accuracy with larger fairness gaps, while simpler models tend to be fairer and less accurate.\"}]","A COMPARATIVE BENCHMARK OF FAIRNESS METRICS IN MACHINE LEARNING - Master’s Thesis | PDF",1785684908,156,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-comparative-benchmark-of-fairness-metrics-in-machine-learning-masters-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-comparative-benchmark-of-fairness-metrics-in-machine-learning-masters-thesis/118692/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address in machine learning deployments?","Question",{"text":76,"@type":77},"It addresses how fairness-related errors or biases in machine learning can harm individuals in real-world high-stakes decisions, beyond just achieving accuracy.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which datasets and fairness criteria are used for benchmarking?",{"text":81,"@type":77},"The study uses COMPAS recidivism and Adult Income, evaluating models with Demographic Parity Difference (DPD) and Equalized Odds Difference (EOD) plus traditional performance metrics such as Accuracy and F1-score.",{"name":83,"@type":74,"acceptedAnswer":84},"What main relationship is found between accuracy and fairness across models?",{"text":85,"@type":77},"A trade-off between accuracy and fairness is observed, but its strength depends on model architecture; ensemble models like XGBoost and LightGBM show higher accuracy with larger fairness gaps, while simpler models tend to be fairer and less accurate.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]