[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119625-en":3,"doc-seo-119625-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119625,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Explainable Machine Learning - Approximating Shapley Values for Dependent Predictors","Modern machine learning improves predictive accuracy at the cost of interpretability, which limits trust in high-stakes decisions across enterprises and institutions. The thesis focuses on Shapley values from game theory as a universal explanation method, but addresses their high computational cost through approximations. It reviews Shapley theory, evaluates original Kernel SHAP against recent approaches that handle dependent predictors using three real-world datasets, and reports reduced approximation error alongside higher computational demand. Future improvements are discussed.","Explainable Machine Learning:  \nApproximating Shapley Values for Dependent Predictors  \nThesis submitted in partial fulfillment of the requirements for the degree  \nMaster of Science in Statistics  \nby  \nJan Kasperek  \nJanuary 2024  \nSupervisor:  \nProf. tit. Dr. Alina Matei  \nUniversité de Neuchâtel  \nFaculty of Science  \nInstitute of Statistics  \nAbstract  \nModern Machine Learning algorithms often outperform classical statistical methods in predictive accuracy. This comes at the expense of model interpretability. As businesses and institutions increasingly rely on Machine Learning to support and automate decision making processes to reap the benefits of more accurate predictions, explaining these model outputs becomes more important. A universally applicable approach to explaining such complex models is based on the Shapley value, a concept originating from game theory. However, its calculation is very computer-intensive, so approximations have to be used. The state-of-the-art approach, Kernel SHAP, assumes independence of the predictors, which is unrealistic in practice. Recent research has developed improvements to incorporate dependencies between predictors. After a review of the theoretical underpinnings, the original KernelSHAP method is compared with improved versions in realistic settings, using three real-world datasets. While the improved versions are found to have smaller approximation error to exact Shapley values, they are also more computationally demanding. Further improvements are discussed and possible research directions are suggested. The thesis is structured as follows: After introducing explainable machine learning in chapter 1, the Shapley value and its applications to model explainability are explored in chapter 2 . Chapter 3 presents methods to approximate Shapley values as well as recent improvements to these methods, which are tested on real datasets in chapter 4 . Some possible directions for future research are pointed out in chapter 5, before giving a final conclusion in chapter 6 . Code for the experiments of chapter 4 is found in the appendix.  \nContents  \n1. Introduction: Explainable Machine Learning 1  \n1.1 Model-Specific Approaches .............................. 2  \n1.1.1 Example: Generalized Linear Models .................... 3  \n1.2 Model-Agnostic Approaches .............................. 4  \n2. The Shapley Value 6  \n2.1 The Shapley Value in Game Theory .......................... 6  \n2.2 The Shapley Value in Machine Learning ....................... 8  \n3. Approximating Shapley Values 12  \n3. 1 Kernel SHAP . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13  \n3.1.1 Accounting for Dependent Predictors .................... 15  \n4. Applications to Real Datasets 18  \n4.1 Data Descriptions .................................... 19  \n4.2 Model Training ..................................... 25  \n4.3 Comparison of Shapley Value Approximations ................... 27  \n5. Discussion of Further Improvements 35  \n6. Conclusions 37  \nReferences 38  \nAppendix: R Code 41  \n1. Introduction: Explainable Machine Learning  \nWith ever-increasing computational resources at decreasing costs, the use of complex predictive models becomes more widespread. Such machine learning models often offer an advantage when it comes to predictive accuracy, making it an attractive choice for prediction problems in enterprises and other institutions over classical statistical models. However, higher model complexity comes at the cost of lower interpretability: It becomes less clear how the model inputs (predictors) actually relate to the output (i.e., the predictions) . For instance, in a classical linear regression model, it is immediately visible how any given predictor affects the predictions: Both are linked by a simple linear function. Such a simple link does not exist for many other commonly used predictive models such as neural networks or tree ensembles such as random forests.  \nSince such “blac","cbCaiuUb6EQdKheS","https://ap.wps.com/l/cbCaiuUb6EQdKheS","pdf",5348550,1,64,"English","en",105,"# 1. Introduction: Explainable Machine Learning\n## 1.1 Model-Specific Approaches\n## 1.2 Model-Agnostic Approaches\n# 2. The Shapley Value\n## 2.1 The Shapley Value in Game Theory\n## 2.2 The Shapley Value in Machine Learning\n# 3. Approximating Shapley Values\n## 3.1 Kernel SHAP\n## 3.1.1 Accounting for Dependent Predictors\n# 4. Applications to Real Datasets\n## 4.1 Data Descriptions\n## 4.2 Model Training\n## 4.3 Comparison of Shapley Value Approximations\n# 5. Discussion of Further Improvements\n# 6. Conclusions\n# References\n# Appendix: R Code","[{\"question\":\"Why are approximations needed for Shapley values in machine learning explanations?\",\"answer\":\"Exact Shapley value calculation is very computer-intensive, so practical methods rely on approximations to make explanations feasible for real models.\"},{\"question\":\"What limitation of Kernel SHAP motivates the use of improved methods?\",\"answer\":\"Kernel SHAP assumes predictors are independent, which is often unrealistic in practice, motivating methods that incorporate dependencies between predictors.\"},{\"question\":\"How do the improved dependent-predictor Shapley approximations compare to the original KernelSHAP method?\",\"answer\":\"The improved versions achieve smaller approximation error relative to exact Shapley values, but they are also more computationally demanding.\"}]","Explainable Machine Learning - Approximating Shapley Values for Dependent Predictors | PDF",1785725363,161,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"explainable-machine-learning-approximating-shapley-values-for-dependent-predictors","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/explainable-machine-learning-approximating-shapley-values-for-dependent-predictors/119625/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are approximations needed for Shapley values in machine learning explanations?","Question",{"text":75,"@type":76},"Exact Shapley value calculation is very computer-intensive, so practical methods rely on approximations to make explanations feasible for real models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation of Kernel SHAP motivates the use of improved methods?",{"text":80,"@type":76},"Kernel SHAP assumes predictors are independent, which is often unrealistic in practice, motivating methods that incorporate dependencies between predictors.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the improved dependent-predictor Shapley approximations compare to the original KernelSHAP method?",{"text":84,"@type":76},"The improved versions achieve smaller approximation error relative to exact Shapley values, but they are also more computationally demanding.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]