[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121933-en":3,"doc-seo-121933-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121933,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","OPENHEXAI - An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning","Explainable AI (XAI) methods are increasingly needed to make machine learning behavior understandable in high-stakes settings, yet evaluating their effectiveness requires human subjects and human-centered benchmarks are difficult to run and reproduce. OpenHEXAI is an open-source framework that streamlines this process by providing diverse benchmark datasets, pre-trained models, post hoc explanation methods, an easy web application for user studies, comprehensive evaluation metrics, and guidance for experiment documentation. It also includes tools for power analysis and cost estimation, supporting reproducible comparisons and systematic benchmarking across human-AI decision-making tasks.","arXiv :2403 .05565v1 [ cs .HC] 20 Feb 2024  \nOPENHEXAI: AN OPEN-SOURCE FRAMEWORK FOR HUMAN-CENTERED EVALUATION OF EXPLAINABLE MACHINE  \nLEARNING  \nJiaqi Ma∗  \nUIUC  \nVivian Lai∗  \nVisa Research  \nYiming Zhang  \nCarnegie Mellon University  \nChacha Chen Paul Hamilton  \nUniversity of Chicago Harvard University  \nDavor Ljubenkov  \nHarvard University  \nHimabindu Lakkaraju† Harvard University  \nChenhao Tan† University of Chicago  \nABSTRACT  \nRecently, there has been a surge of explainable AI (XAI) methods driven by the need for understanding machine learning model behaviors in high-stakes scenarios. However, properly evaluating the effectiveness of the XAI methods inevitably requires the involvement of human subjects, and conducting human-centered benchmarks is challenging in a number of ways: designing and implementing user studies is complex; numerous design choices in the design space of user study lead to problems of reproducibility; and running user studies can be challenging and even daunting for machine learning researchers. To address these challenges, this paper presents OpenHEXAI, an open-source framework for human-centered evaluation of XAI methods. OpenHEXAI features (1)  \na collection of diverse benchmark datasets, pre-trained models, and post hoc explanation methods;  \n(2) an easy-to-use web application for user study; (3) comprehensive evaluation metrics for the effectiveness of post hoc explanation methods in the context of human-AI decision making tasks;(4) best practice recommendations of experiment documentation; and (5) convenient tools for power analysis and cost estimation. OpenHEAXI is the first large-scale infrastructural effort to facilitate human-centered benchmarks of XAI methods. It simplifies the design and implementation of user studies for XAI methods, thus allowing researchers and practitioners to focus on the scientific questions. Additionally, it enhances reproducibility through standardized designs. Based on OpenHEXAI, we further conduct a systematic benchmark of four state-of-the-art post hoc explanation methods and compare their impacts on human-AI decision making tasks in terms of accuracy, fairness, as well as users’ trust and understanding of the machine learning model.  \n1 Introduction  \nExplainable Machine Learning/Artificial Intelligence (XAI) methods aim to provide human-understandable insights on machine learning models. These insights are crucial in high-stakes applications such as hiring, loan approvals, or medical diagnosis, where the algorithmic decisions made by machine learning models can significantly impact people’s lives. Although there has been a surge in the development of XAI methods in recent years, it is highly non-trivial to evaluate and compare these methods properly. Unlike measuring the accuracy of AI predictions, quantifying explainability objectively is difficult since it is fundamentally dependent on human interpretation, and there is still little consensus on a specific set of objective metrics suitable for measuring explainability. As a result, human-centered evaluation (user study) is often adopted, which typically involves soliciting feedback from users who interact with an XAI method and then analyzing the feedback data to investigate how well the method is providing human-understandable information. However, designing and conducting user studies can be complicated and expensive, and lead to incomparable outcomes across different studies. To date, there has not been a systematic benchmark for human-centered evaluation of XAI  \n∗Equal contribution.†Equal supervision.  \nmethods. In this work, we propose OpenHEXAI, an Open-source framework for human-centered Evaluation of XAI methods, aiming to address the aforementioned challenges and establish a systematic and replicable benchmark.  \nDesigning and conducting user studies for evaluating XAI methods can be challenging for a few reasons. Firstly, one can measure the explainability of an XAI method from a diverse lens","cbCaibukyehrEYnD","https://ap.wps.com/l/cbCaibukyehrEYnD","pdf",949499,1,18,"English","en",105,"# Abstract\n# Introduction\n## Challenges of human-centered XAI evaluation\n## OpenHEXAI framework overview","[{\"question\":\"Why is evaluating explainable AI methods challenging in high-stakes applications?\",\"answer\":\"XAI effectiveness depends on human interpretation, so evaluation cannot rely solely on prediction accuracy. Human-centered benchmarks also require user studies that are complex, expensive, and hard to reproduce across different experimental designs.\"},{\"question\":\"What is OpenHEXAI and what problem does it address?\",\"answer\":\"OpenHEXAI is an open-source framework for human-centered evaluation of XAI methods. It targets the difficulty of designing, implementing, and documenting user studies for post hoc explanation methods.\"},{\"question\":\"How does OpenHEXAI support reproducible and scalable human-centered benchmarks?\",\"answer\":\"It provides standardized components, including benchmark datasets, pre-trained models, post hoc explanation methods, a web application with configurable user-study interfaces, comprehensive evaluation metrics, and best-practice recommendations for experiment documentation.\"}]","OPENHEXAI - An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning | PDF",1785807811,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"openhexai-an-open-source-framework-for-human-centered-evaluation-of-explainable-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/openhexai-an-open-source-framework-for-human-centered-evaluation-of-explainable-machine-learning/121933/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is evaluating explainable AI methods challenging in high-stakes applications?","Question",{"text":76,"@type":77},"XAI effectiveness depends on human interpretation, so evaluation cannot rely solely on prediction accuracy. Human-centered benchmarks also require user studies that are complex, expensive, and hard to reproduce across different experimental designs.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is OpenHEXAI and what problem does it address?",{"text":81,"@type":77},"OpenHEXAI is an open-source framework for human-centered evaluation of XAI methods. It targets the difficulty of designing, implementing, and documenting user studies for post hoc explanation methods.",{"name":83,"@type":74,"acceptedAnswer":84},"How does OpenHEXAI support reproducible and scalable human-centered benchmarks?",{"text":85,"@type":77},"It provides standardized components, including benchmark datasets, pre-trained models, post hoc explanation methods, a web application with configurable user-study interfaces, comprehensive evaluation metrics, and best-practice recommendations for experiment documentation.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]