[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119939-en":3,"doc-seo-119939-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119939,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","The Meta-Evaluation Problem in Explainable AI - Identifying Reliable Estimators with MetaQuantus","Explainable AI (XAI) focuses on improving transparency and trustworthiness of AI systems for human users. A key unresolved challenge is estimating how well XAI explanation methods perform for neural networks, where many competing metrics lack clear guidance on which evaluation method to prefer. This paper introduces MetaQuantus, a meta-evaluation framework that jointly assesses estimator resilience to noise and reaction to randomness, demonstrating reliability through experiments across open XAI questions.","The Meta-Evaluation Problem in Explainable AI: Identifying Reliable Estimators with MetaQuantus  \narXiv :2302 .07265v1 [ cs .LG] 14 Feb 2023  \nAnna Hedstr􀁿om 1 ;2 ;y Philine Bommer 1 ;2 Kristo􀀋er K. Wickstr􀀜m6 Wojciech Samek3 ;4 ;5 Sebastian Lapuschkin4 Marina M.-C. H􀁿ohne 1 ;2 ;5 ;6 ;7 ;y  \n1 Department of Machine Learning, Technische Universit􀁿at Berlin, 10587 Berlin, Germany  \n2 Understandable Machine Intelligence Lab, Department of Data Science, ATB, 14469 Potsdam, Germany  \n3 Department of Electrical Engineering and Computer Science, TU Berlin, 10587 Berlin, Germany  \n4 Department of Arti􀀌cial Intelligence, Fraunhofer Heinrich-Hertz-Institute, 10587 Berlin, Germany  \n5 BIFOLD { Berlin Institute for the Foundations of Learning and Data, 10587 Berlin, Germany  \n6 Machine Learning Group, UiT the Arctic University of Norway, 9037 Troms􀀜, Norway  \n7 Department of Computer Science, University of Potsdam, 14476 Potsdam, Germany  \ny corresponding authors  \nAbstract  \nExplainable AI (XAI) is a rapidly evolving 􀀌eld that aims to improve transparency and trustworthiness of AI systems to humans. One of the unsolved challenges in XAI is estimating the performance of these explanation methods for neural networks, which has resulted in numerous competing metrics with little to no indication of which one is to be preferred. In this paper, to identify the most reliable evaluation method in a given explainability context, we propose MetaQuantus|a simple yet powerful framework that meta-evaluates two complementary performance characteristics of an evaluation method: its resilience to noise and reactivity to randomness. We demonstrate the e􀀋ectiveness of our framework through a series of experiments, targeting various open questions in XAI, such as the selection of explanation methods and optimisation of hyperparameters of a given metric. We release our work under an open-source license1 to serve as a development tool for XAI researchers and Machine Learning (ML) practitioners to verify and benchmark newly constructed metrics (i.e., \\estimators\" of explanation quality) . With this work, we provide clear and theoretically-grounded guidance for building reliable evaluation methods, thus facilitating standardisation and reproducibility in the 􀀌eld of XAI.  \n1 Introduction  \nSince Explainable AI (XAI) is intended to increase trust and transparency in AI systems, it is necessary to evaluate the performance of proposed explanation methods to ensure their reliability. Apart from simpler or well-understood data domains where critical input features are known and models are interpretable (e.g. , linear functions and shallow decision trees), in the context of more complex Machine Learning (ML) models such as neural networks (NNs), there is generally an absence of ground truth labels for explanations [1] . This makes it di􀀎cult to evaluate the performance of explanation methods since the exact outcomes of explanations oftentimes remain unknown and thus unveri􀀌able [2] . Without consensus around how to de􀀌ne the quality or \\correctness\" of an explanation method, a variety of evaluation methods have been proposed. These e􀀋orts most commonly involve (i) measuring the extent to which desirable properties are ful􀀌lled, e.g., through faithfulness or robustness analysis [3, 4 , 5], (ii) generating well-de􀀌ned, synthetic settings where explanation labels are simulated [6, 7] or, (iii) evaluating explanations based on visual alignment with a human prior [8] . Most relevant to our work is the 􀀌rst category of evaluation techniques or \\metrics\"whose goal is to estimate the quality of attribution-based explanations. We henceforth refer to these XAI evaluation methods as \\quality estimators\", or simply \\estimators\".  \n1 Code released at the GitHub repository: [https://github.com/annahedstroem/MetaQuantus](https://github.com/annahedstroem/MetaQuantus).  \nFigure 1: An illustration of the Problem of Meta-Evaluation through three phases: (i) Modeling, (ii) Explaining a","cbCaioHb1ZiOuCPC","https://ap.wps.com/l/cbCaioHb1ZiOuCPC","pdf",8427132,1,30,"English","en",105,"# Abstract\n# Introduction\n## Evaluating explanation quality in the absence of ground truth\n## Existing evaluation approaches and quality estimators\n## Meta-evaluation motivation and gap in prior work","[{\"question\":\"What is the meta-evaluation problem in Explainable AI addressed by this work?\",\"answer\":\"It evaluates the evaluation methods themselves—i.e., how to judge the reliability of quality estimators used to compare, select, or reject explanation methods in XAI.\"},{\"question\":\"How does MetaQuantus assess the reliability of evaluation methods?\",\"answer\":\"It meta-evaluates two complementary performance characteristics of an estimator: resilience to noise and reactivity to randomness.\"},{\"question\":\"What kinds of experiments and XAI questions does the framework target?\",\"answer\":\"The paper reports experiments that support selection of explanation methods and optimization of hyperparameters for metrics, covering multiple open questions in XAI.\"}]","The Meta-Evaluation Problem in Explainable AI - Identifying Reliable Estimators with MetaQuantus | PDF",1785727102,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-meta-evaluation-problem-in-explainable-ai-identifying-reliable-estimators-with-metaquantus","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-meta-evaluation-problem-in-explainable-ai-identifying-reliable-estimators-with-metaquantus/119939/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the meta-evaluation problem in Explainable AI addressed by this work?","Question",{"text":75,"@type":76},"It evaluates the evaluation methods themselves—i.e., how to judge the reliability of quality estimators used to compare, select, or reject explanation methods in XAI.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MetaQuantus assess the reliability of evaluation methods?",{"text":80,"@type":76},"It meta-evaluates two complementary performance characteristics of an estimator: resilience to noise and reactivity to randomness.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of experiments and XAI questions does the framework target?",{"text":84,"@type":76},"The paper reports experiments that support selection of explanation methods and optimization of hyperparameters for metrics, covering multiple open questions in XAI.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]