[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126966-en":3,"doc-seo-126966-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126966,687207024478,"Liam","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Biases and Ethical considerations for Machine Learning pipelines in the Computational Social Sciences - Abstract","Computational social science studies using AI/ML to derive patterns and inferences from large datasets face bias risks across dataset construction, collection, and analysis. Interdisciplinary requirements—policy and rights context, rapidly evolving AI/ML paradigms, and dataset-specific pitfalls—affect where bias enters the pipeline. The chapter presents a bias taxonomy aligned with the AI/ML lifecycle, then discusses methods to detect and mitigate bias and promote responsible research and innovation.","Biases and Ethical considerations for Machine Learning pipelines in the Computational Social Sciences  \nSuparna De, Shalini Jangra, Vibhor Agarwal, Jon Johnson and Nishanth Sastry  \nAbstract Computational analyses driven by Artificial Intelligence (AI)/Machine Learning (ML) methods to generate patterns and inferences from big datasets in computational social science (CSS) studies can suffer from biases during the data construction, collection and analysis phases as well as encounter challenges of generalizability and ethics. Given the interdisciplinary nature of CSS, many factors such as the need for a comprehensive understanding of different facets such as the policy and rights landscape, the fast evolving AI/ML paradigms and dataset specific pitfalls influence the possibility of biases being introduced. This chapter identifies challenges faced by researchers in the CSS field and presents a taxonomy of biases that mayarise in AI/ML approaches. The taxonomy mirrors the various stages of common AI/ML pipelines: dataset construction and collection, data analysis and evaluation. With detecting and mitigating bias in AI an active area of research, this chapter seeks to highlight practices for incorporating responsible research and innovation into CSS practices.  \nSuparna De  \nUniversity of Surrey, UK. e-mail: [s.de@surrey.ac.uk](s.de@surrey.ac.uk)  \nShalini Jangra  \nUniversity of Surrey, UK. e-mail: [s.jangra@surrey.ac.uk](s.jangra@surrey.ac.uk)  \nVibhor Agarwal  \nUniversity of Surrey, UK. e-mail: [v.agarwal@surrey.ac.uk](v.agarwal@surrey.ac.uk)  \nJon Johnson  \nUniversity College, London (UCL), UK. e-mail: [jon.johnson@ucl.ac.uk](jon.johnson@ucl.ac.uk)  \nNishanth Sastry  \nUniversity of Surrey, UK. e-mail: [n.sastry@surrey.ac.uk](n.sastry@surrey.ac.uk)  \n2 Suparna De, Shalini Jangra, Vibhor Agarwal, Jon Johnson and Nishanth Sastry  \n1 Introduction  \nAdvances in communication networks and the growing use of social networking platforms means that there is an unprecedented amount of information that providesan important source for understanding a population [1, 2] . Computational tools have been successfully used to analyze the resulting structured and unstructured data, with the aim of understanding individuals, groups and their social practices. This well-studied field of computational social science (CSS) is characterized by: (1) the involvement of human subjects, with the resulting capabilities and tools also impacting individuals and communities,(2) the use of large and complex datasets, drawn from mixed methods data collection, incorporating both self-reporting through surveys and experiments, as well as through observation of ‘unconstrained’behaviour on social media platforms,(3) application of AI or ML-driven computational or algorithmic solutions to the resulting big data to generate insights, inferences and predictions about human behaviours, social networks and systems.  \nWe cannot use ML predictive models in a black box fashion for social science problems [24] . It is necessary to analyze the ethical implications and consequences of these models’ output as these may have real world consequences and impacts. Due to this human impact, computational research needs to be “ethical, trustworthy and responsible” [3] . However, this very human nature of the data means that it encounters issues of representativeness, uniformity and bias [1] . Thus, this chapter focuses on some of the key issues around ethics and generalizability confronting CSS researchers in the age of big data. These issues are analyzed through the lens of the data lifecycle in ML pipelines, as identified in existing literature [4], i.e. covering dataset creation/collection, data analysis and data (model) evaluation, as shown in Figure 1 . This is followed by a discussion of the strategies and existing initiatives to address the issue of bias in CSS ML pipelines.  \n2 Dataset Creation and Collection Bias  \nThe first stage of a typical ML pipeline starts with data ","cbCais52w9muOijj","https://ap.wps.com/l/cbCais52w9muOijj","pdf",445971,1,15,"English","en",105,"# Introduction\n## Dataset Creation and Collection Bias\n### Sampling Bias","[{\"question\":\"Why do ML models in computational social science require ethical and ethical-output analysis?\",\"answer\":\"ML predictions can have real-world consequences for humans and communities. Therefore, ethical implications and impacts of model outputs must be analyzed rather than using models as black boxes.\"},{\"question\":\"What stages of an AI/ML pipeline are connected to bias in computational social science?\",\"answer\":\"Bias can arise during dataset construction and collection, and again during data analysis and evaluation. The chapter’s taxonomy mirrors these common pipeline stages.\"},{\"question\":\"How does sampling bias affect generalization of AI models?\",\"answer\":\"Sampling bias occurs when some instances are selected more than others, leading to distorted representativeness. This distortion can compromise generalization and produce unintended poor performance for the broader population.\"}]","Biases and Ethical considerations for Machine Learning pipelines in the Computational Social Sciences - Abstract | PDF",1785935948,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"biases-and-ethical-considerations-for-machine-learning-pipelines-in-computational-social-sciences-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/biases-and-ethical-considerations-for-machine-learning-pipelines-in-computational-social-sciences-abstract/126966/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do ML models in computational social science require ethical and ethical-output analysis?","Question",{"text":75,"@type":76},"ML predictions can have real-world consequences for humans and communities. Therefore, ethical implications and impacts of model outputs must be analyzed rather than using models as black boxes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What stages of an AI/ML pipeline are connected to bias in computational social science?",{"text":80,"@type":76},"Bias can arise during dataset construction and collection, and again during data analysis and evaluation. The chapter’s taxonomy mirrors these common pipeline stages.",{"name":82,"@type":73,"acceptedAnswer":83},"How does sampling bias affect generalization of AI models?",{"text":84,"@type":76},"Sampling bias occurs when some instances are selected more than others, leading to distorted representativeness. This distortion can compromise generalization and produce unintended poor performance for the broader population.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]