[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125019-en":3,"doc-seo-125019-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125019,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","A Conceptual Framework for Ethical Evaluation of Machine Learning Systems","Responsible AI research has expanded principles and practices to guide the ethical use of machine learning systems, yet ethical concerns embedded in the evaluation design process remain underexamined. This work conceptualizes ethics-related issues in standard ML evaluation methods by proposing a utility framework that characterizes the central trade-off: maximizing information gain while limiting potential ethical harms. It provides a lens to separate competing considerations, draws parallels to domains like clinical trials and automotive crash testing, and argues for deliberate management of ethical complexities during evaluation, supported by institutional policies.","arXiv :2408 . 10239v1 [ cs .CY] 5 Aug 2024  \nA Conceptual Framework for Ethical Evaluation of Machine Learning Systems  \nNeha R. Gupta 1 , Jessica Hullman2 , Hari Subramonyam3  \n1 Carnegie Mellon University, Pittsburgh, USA  \n[nehagupt@andrew.cmu.edu](nehagupt@andrew.cmu.edu)  \n2 Northwestern University, Evanston, USA  \n[jhullman@northwestern.edu](jhullman@northwestern.edu)  \n3 Stanford University, Stanford, USA  \n[harihars@stanford.edu](harihars@stanford.edu)  \nAbstract  \nResearch in Responsible AI has developed a range of principles and practices to ensure that machine learning systems are used in a manner that is ethical and aligned with human values. However, a critical yet often neglected aspect of ethical ML is the ethical implications that appear when designing evaluations of ML systems. For instance, teams may have to balance a trade-off between highly informative tests to ensure downstream product safety, with potential fairness harms inherent to the implemented testing procedures. We conceptualize ethics-related concerns in standard ML evaluation techniques. Speci􀀂cally, we present a utility framework, characterizing the key trade-off in ethical evaluation as balancing information gain against potential ethical harms. The framework is then a tool for characterizing challenges teams face, and systematically disentangling competing considerations that teams seek to balance. Differentiating between different types of issues encountered in evaluation allows us to highlight best practices from analogous domains, such as clinical trials and automotive crash testing, which navigate these issues in ways that can offer inspiration to improve evaluation processes in ML. Our analysis underscores the critical need for development teams to deliberately assess and manage ethical complexities that arise during the evaluation of ML systems, and for the industry to move towards designing institutional policies to support ethical evaluations.  \nIntroduction  \nMachine learning (ML) model evaluation typically focuses on estimating errors of prediction or estimation via quanti􀀂 -able metrics. Given the increasing size and complexity of ML systems, comprehensive evaluations should ideally be multifaceted. For example, evaluations of large ML systems may include several methods, including A/B testing on live populations, adversarial testing to produce undesirable outputs, and comprehensive audits documenting outputs.  \nPotential ethical harms of ML systems have gained increasing attention in the broad Responsible AI community. However, even when evaluation metrics are expanded beyond performance to include factors like fairness, privacy loss, or other harms induced by the machine learning system, this is often focused on the ethical harms ofthe released  \nCopyright © 2024, Association for the Advancement of Arti􀀂cial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \nsystem, overlooking possible harms incurred during the machine learning development lifecycle itself. This is problematic because evaluation approaches do have the potential to cause ethical harm during evaluation. In a noteworthy example, Tesla’s autonomous vehicle live testing systems on real roadways in California, has been widely criticized for being involved in various crashes (Nayak, Laing, and Hull 2022) .  \nHow should practitioners evaluate large complex systems with potentially unknown ethical harms across the engineering lifecycle, including during the evaluation process? We provide a conceptual framework that casts the primary tradeoff in ethical evaluation decision-making as balancing the goal of optimizing for information gained in an evaluation, against the possible ethical harms that are induced.  \nBased on our sketch of this fundamental problem that practitioners face, we identify a series of challenges that can cause practitioners to stumble in selecting ethical evaluation practices. We illustrate these challenges using real-world examples of ","cbCaidasI4c2OdsC","https://ap.wps.com/l/cbCaidasI4c2OdsC","pdf",199138,1,13,"English","en",105,"# Introduction\n## Related Works: Ethical AI\n## Ethical Implications in ML Evaluation Design","[{\"question\":\"What is the central trade-off in ethical evaluation of machine learning systems proposed in the paper?\",\"answer\":\"The framework characterizes the key trade-off as balancing information gain against potential ethical harms induced by the evaluation procedures.\"},{\"question\":\"Why does the paper argue that ethical harms can occur during the ML development lifecycle?\",\"answer\":\"Even if evaluation metrics include fairness or privacy, common approaches may ignore harms created by the evaluation process itself, which can cause ethical harm during development.\"},{\"question\":\"How does the paper suggest improving ethical evaluation practices?\",\"answer\":\"It identifies challenges that lead teams to select ineffective practices, uses real-world evaluation examples, and draws lessons from analogous domains to inform mitigation and future policy or best practices.\"}]","A Conceptual Framework for Ethical Evaluation of Machine Learning Systems | PDF",1785896202,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-conceptual-framework-for-ethical-evaluation-of-machine-learning-systems","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-conceptual-framework-for-ethical-evaluation-of-machine-learning-systems/125019/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the central trade-off in ethical evaluation of machine learning systems proposed in the paper?","Question",{"text":75,"@type":76},"The framework characterizes the key trade-off as balancing information gain against potential ethical harms induced by the evaluation procedures.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the paper argue that ethical harms can occur during the ML development lifecycle?",{"text":80,"@type":76},"Even if evaluation metrics include fairness or privacy, common approaches may ignore harms created by the evaluation process itself, which can cause ethical harm during development.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper suggest improving ethical evaluation practices?",{"text":84,"@type":76},"It identifies challenges that lead teams to select ineffective practices, uses real-world evaluation examples, and draws lessons from analogous domains to inform mitigation and future policy or best practices.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]