[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127081-en":3,"doc-seo-127081-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127081,5909887256941,"Levi","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","EXACT - Towards a Platform for Empirically Benchmarking Machine Learning Model Explanation Methods","The evolving landscape of explainable artificial intelligence (XAI) seeks to improve interpretability of complex machine learning models, but faces difficulties in formalising and empirically validating explanation quality. This paper introduces an initial benchmarking platform, EXACT, which unifies benchmark datasets and quantitative performance metrics for XAI. It evaluates post-hoc XAI methods using datasets with ground-truth explanations for class-conditional features. Findings reveal limits of popular approaches, often failing to beat random baselines and producing explanations that depend on equally performing architectures.","XXIV IMEKO World Congress “Think Metrology ”  \nAugust 26-29, 2024, Hamburg, Germany  \nEXACT: TOWARDS A PLATFORM FOR EMPIRICALLY BENCHMARKING MACHINE LEARNING MODEL EXPLANATION METHODS  \nBenedict Clark a, *, Rick Wilming b, Artur Dox a,b,  \nPaul Eschenbach b, Sami Hachedb, Daniel Jin Wodkeb, Michias Taye Zewdie b, Uladzislau Bruilab, Marta Oliveira a, Hjalmar Schulz b,c, Luca Matteo Cornilsb, Danny Panknina, Ahcène Boubekkia, Stefan Haufea,b,c  \na Physikalisch-Technische Bundesanstalt, Abbestrasse 2-12, 10587 Berlin, Germany, email address: [stefan.haufe@ptb.de](stefan.haufe@ptb.de)b Technische Universität Berlin, Str. des 17. Juni 135, 10623 Berlin, Germany, email [address: haufe@tu-berlin.de](address: haufe@tu-berlin.de)c Charité – Universitätsmedizin Berlin, Charitéplatz 1, 10117 Berlin, Germany, email [address: haufe@tu-berlin.de](address: haufe@tu-berlin.de)  \n* Corresponding author  \nAbstract-The evolving landscape of explainable artificial intelligence (XAI) aims to improve the interpretability of intricate machine learning (ML) models, yet faces challenges in formalisation and empirical validation, being an inherently unsupervised process. In this paper, we bring together various benchmark datasets and novel performance metrics in an initial benchmarking platform, the Explainable AI Comparison Toolkit (EXACT), providing a standardised foundation for evaluating XAI methods. Our datasets incorporate ground truth explanations for class-conditional features, and leveraging novel quantitative metrics, this platform assesses the performance of post-hoc XAI methods in the quality of the explanations they produce. Our recent findings have highlighted the limitations of popular XAI methods, as they often struggle to surpass random baselines, attributing significance to irrelevant features. Moreover, we show the variability in explanations derived from different equally performing model architectures. This initial benchmarking platform therefore aims to allow XAI researchers to test and assure the high quality of their newly developed methods.  \nKeywords: Explainable AI, Benchmark, Explanation Performance, Deep Learning, Non-linear Problems, Suppressor Variables  \n1. INTRODUCTION  \nResearch in the field of Explainable AI (XAI) aims to‘explain’ the decisions of complicated Machine Learning (ML) models, with authors aiming to deploy their methods to high stakes domains such as medicine and law [1–3] . In recent years, a plethora of XAI methods have been developed to achieve this goal. In the past, methods such as SHAP [4] or LIME [5] have emerged as popular choices to assess the quality of ML models. In addition, the quality and robustness of such methods have been assessed by various supporting evaluation studies, already highlighting weaknesses of such methods. However, it remains unclear what to conclude from the output of XAI methods in general, since there is a lack of a formal problem definition of explainability. Being an inherently unsupervised task, formalization and empirical validation of the quality of explanations produced is difficult and limits their potential use for quality-control and  \ntransparency purposes. As such, current research often tends to rely on subjective evaluation of methods, for example through user studies on which of two given explanations appear better qualitatively [6] as well as evaluation of secondary properties of explanation methods [5] . Often these evaluation studies do not contain a formal definition of an explanation, but this is important to be able to interpret explanations correctly and understand their limits. When faced with so-called suppressor variables, in the context of explanations, high importance may be attributed to these types of variables although they lack any statistical relation to the prediction target [7] . The inclusion of suppressors may allow a model to remove unwanted noise, which can lead to improved prediction quality. While it is clear suppressors can be usefu","cbCaicJDR948Sra0","https://ap.wps.com/l/cbCaicJDR948Sra0","pdf",639627,1,9,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the EXACT platform address in explainable AI?\",\"answer\":\"EXACT targets the lack of formalisation and empirical validation for explanation quality, enabling objective benchmarking instead of primarily subjective assessments.\"},{\"question\":\"How does EXACT evaluate post-hoc XAI methods?\",\"answer\":\"It combines benchmark datasets with ground-truth explanations for class-conditional features and uses novel quantitative metrics to measure explanation quality.\"},{\"question\":\"What do the reported findings indicate about popular XAI methods?\",\"answer\":\"The results show that many popular methods may struggle to outperform random baselines and can attribute importance to irrelevant suppressor features.\"}]","EXACT - Towards a Platform for Empirically Benchmarking Machine Learning Model Explanation Methods | PDF",1785936747,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"exact-towards-a-platform-for-empirically-benchmarking-machine-learning-model-explanation-methods","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exact-towards-a-platform-for-empirically-benchmarking-machine-learning-model-explanation-methods/127081/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the EXACT platform address in explainable AI?","Question",{"text":75,"@type":76},"EXACT targets the lack of formalisation and empirical validation for explanation quality, enabling objective benchmarking instead of primarily subjective assessments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EXACT evaluate post-hoc XAI methods?",{"text":80,"@type":76},"It combines benchmark datasets with ground-truth explanations for class-conditional features and uses novel quantitative metrics to measure explanation quality.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the reported findings indicate about popular XAI methods?",{"text":84,"@type":76},"The results show that many popular methods may struggle to outperform random baselines and can attribute importance to irrelevant suppressor features.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]