[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124218-en":3,"doc-seo-124218-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124218,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","An Empirical Evaluation of the Rashomon Effect in Explainable Machine Learning","The Rashomon Effect captures how, for a given dataset, multiple models can achieve equally strong performance while relying on different solution strategies. This paper examines consequences for Explainable Machine Learning, focusing on the comparability of explanations. A unified framework is proposed across three comparison scenarios, and a quantitative study evaluates diverse datasets, model families, attribution methods, and similarity metrics. Findings show hyperparameter tuning influences explanation outcomes and that metric choice substantially affects evaluation reliability.","arXiv :2306 . 15786v2 [ cs .LG] 29 Jun 2023  \nAn Empirical Evaluation of the Rashomon Effect in Explainable Machine Learning  \nSebastian M¨uller (􀀀)1 ,4[0000−0002−0778−9695], Vanessa  \nToborek 1 ,4[0009−0009−8372−8251], Katharina Beckh3 ,4[0000−0002−7824−6647], Matthias Jakobs2 ,4[0000−0003−4607−8957], Christian Bauckhage 1 ,3 ,4[0000−0001−6615−2128], and Pascal Welke5[0000−0002−2123−3781]  \n1 University of Bonn, Bonn, Germany  \n2 TU Dortmund University, Dortmund, Germany  \n3 Fraunhofer IAIS, Sankt Augustin, Germany  \n4 Lamarr Institute, Germany  \n5 TU Wien, Vienna, Austria  \nAbstract. The Rashomon Effect describes the following phenomenon: for a given dataset there may exist many models with equally good performance but with different solution strategies. The Rashomon Effect has implications for Explainable Machine Learning, especially for the comparability of explanations. We provide a unified view on three different comparison scenarios and conduct a quantitative evaluation across different datasets, models, attribution methods, and metrics. We find that hyperparameter-tuning plays a role and that metric selection matters.  \nOur results provide empirical support for previously anecdotal evidence and exhibit challenges for both scientists and practitioners.  \nKeywords: Explainable ML · Interpretable ML · Attribution Methods  \n· Rashomon Effect · Disagreement Problem  \n1 Introduction  \nWe demonstrate the impact of the Rashomon Effect when analyzing ML models. The Rashomon Effect [8] describes the phenomenon that there may exist many models within a hypothesis class which solve a dataset equally well. The set of these models is referred to as the Rashomon Set [12, 37] . From a data-centric perspective this phenomenon is also called Predictive Multiplicity [23], meaning that there exist many strategies to solve a task on a dataset. Other works use Rashomon Sets to analyze and describe data [12,30] . Somewhat surprisingly, the Rashomon Effect has not yet found wider attention in the Explainable Machine Learning (XML) literature. Although a few works have observed the effect it was only anecdotally or without referring to its proper name [14,20,35] .  \nAccepted for presentation at European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD 2023)  \n2 M¨uller et al.  \nXML has recently become a very active area of research and numerous explanation methods exist [1, 9, 24] . Many approaches explain black-box models in a post-hoc manner by providing attribution scores [22,27] which assign each input dimension a numerical value that represents this feature’s importance with respect to the model decision. Attribution scores are used to answer questions such as “What feature was the most important in this input sample?” and have been used to uncover spurious correlations in the data [29] and biased behavior of models [21] . However, attribution scores are sometimes ambiguous and their interpretation depends on the application context. It is hard to decide at what magnitude a feature is still important, particularly, if magnitudes of attribution scores can be sorted into an evenly descending order. It follows that the task of comparing different attribution methods is a difficult problem. Several works touch upon the problem of explanation comparison [5,7,19,26,35] from different perspectives.  \nOur main contribution is an empirical analysis of one novel and two existing perspectives, 1) demonstrating model-specific sensitivity regarding the hyperparameter choice for explanation methods, 2) comparison of different explanations from the same attribution method on differently initialized but otherwise identical model architectures [5,35] and 3) the disagreement between different explanations applied to the same architecture and parameterization [19,26] . We place these three perspectives into a unified framework to investigate how the Rashomon Effect manifests itself in each situation. ","cbCaihalUtocj8nH","https://ap.wps.com/l/cbCaihalUtocj8nH","pdf",398484,1,17,"English","en",105,"# Introduction\n## Rashomon Effect and XML context\n# Comparing Attribution Scores\n## Attribution score dependencies\n# Experimental setup and scenarios\n## Same model, same sample, and attribution method\n## Hyperparameter tuning and metric selection\n# Results and discussion\n## Perspective-specific findings\n# Conclusion","[{\"question\":\"What is the Rashomon Effect in Explainable Machine Learning?\",\"answer\":\"For a fixed dataset, many models can reach comparable performance but use different solution strategies. This creates a Rashomon Set of equally good yet behaviorally diverse models, which complicates how explanations should be compared.\"},{\"question\":\"Why is comparing attribution methods difficult according to the paper?\",\"answer\":\"Attribution scores can be ambiguous and their interpretation depends on application context and on how magnitudes relate to importance. This makes it hard to determine whether two explanations agree or are meaningfully different.\"},{\"question\":\"What factors most affect explanation comparisons in the study?\",\"answer\":\"The evaluation finds that hyperparameter tuning for explanation methods plays an important role. It also shows that the selection of similarity/assessment metrics can change conclusions about explanation quality or disagreement.\"}]","An Empirical Evaluation of the Rashomon Effect in Explainable Machine Learning | PDF",1785821067,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"an-empirical-evaluation-of-the-rashomon-effect-in-explainable-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/an-empirical-evaluation-of-the-rashomon-effect-in-explainable-machine-learning/124218/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the Rashomon Effect in Explainable Machine Learning?","Question",{"text":75,"@type":76},"For a fixed dataset, many models can reach comparable performance but use different solution strategies. This creates a Rashomon Set of equally good yet behaviorally diverse models, which complicates how explanations should be compared.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is comparing attribution methods difficult according to the paper?",{"text":80,"@type":76},"Attribution scores can be ambiguous and their interpretation depends on application context and on how magnitudes relate to importance. This makes it hard to determine whether two explanations agree or are meaningfully different.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors most affect explanation comparisons in the study?",{"text":84,"@type":76},"The evaluation finds that hyperparameter tuning for explanation methods plays an important role. It also shows that the selection of similarity/assessment metrics can change conclusions about explanation quality or disagreement.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]