[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127911-en":3,"doc-seo-127911-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127911,137451207643,"Noah","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Interpreting Black-box Machine Learning Models for High Dimensional Datasets - Research and Analysis","High-dimensional datasets often contain many irrelevant features, adding noise and increasing computational complexity. Deep neural networks achieve strong performance but act as black-box systems due to non-linearity and higher-order feature interactions. This paper proposes a method that trains a black-box on full features, probes and perturbs it to obtain global top-k feature importance, and then trains an interpretable surrogate on the top-k space. Local decision rules and counterfactuals are derived to approximate and explain outcomes.","Interpreting Black-box Machine Learning Models for High Dimensional Datasets  \nMd. Rezaul Karim∗†, Md Shajalal†‡, Alexander Graß†∗ , Till Dhmen†, Sisay Adugna Chala†∗ ,  \nAlexander Boden†¶ , Christian Beecks§†, and Stefan Decker∗†  \n∗ Computer Science 5 -Information Systems and Databases, RWTH Aachen University, Germany † Fraunhofer -Institute for Applied Information Technology FIT, Germany ‡ University of Siegen, Germany  \n§ University of Hagen, Germany  \n¶ Bonn-Rhein-Sieg University of Applied Sciences, Germany  \narXiv :2208 . 13405v4 [ cs .LG] 21 Nov 2023  \nAbstract—Many datasets are of increasingly high dimensionality, where a large number of features could be irrelevant to the learning task. The inclusion of such features would not only introduce unwanted noise but also increase computational complexity. Deep neural networks (DNNs) outperform machine learning (ML) algorithms in a variety of applications due to their effectiveness in modelling complex problems and handling highdimensional datasets. However, due to non-linearity and higherorder feature interactions, DNN models are unavoidably opaque, making them black-box methods. In contrast, an interpretable model can identify statistically significant features and explain the way they affect the model’s outcome. In this paper1 , we propose a novel method to improve the interpretability of blackbox models in the case of high-dimensional datasets. First, a black-box model is trained on full feature space that learns useful embeddings on which the classification is performed. To decompose the inner principles of the black-box and to identify top-k important features (global explainability), probing and perturbing techniques are applied. An interpretable surrogate model is then trained on top-k feature space to approximate the black-box. Finally, decision rules and counterfactuals are derived from the surrogate to provide local decisions. Our approach outperforms tabular learners, e.g., TabNet and XGboost, and SHAPbased interpretability techniques, when tested on a number of datasets having dimensionality between 54 and 20,5312.  \nIndex Terms—Curse of dimensionality, Black-box models, Interpretability, Attention mechanism, Model surrogation.  \nI. INTRODUCTION  \nHigh availability and easy access to large datasets, AI accelerators, and state-of-the-art machine learning (ML) and deep learning (DNNs) algorithms paved the way for performing predictive modelling at scale. However, in the case of high-dimensional datasets (e.g., omics), the feature space exponentially increases. Principal component analysis (PCA) and isometric feature mapping (Isomap) are widely used to tackle the curse of dimensionality [1] . Although they preserve inter-point distances, they are fundamentally limited to linear embedding and tend to lose useful information, which makes them less effective in dimensionality reduction [2] . The inclusion of a large number of irrelevant features not  \n1 This paper is accepted and included in proceedings of 2023 IEEE 10th International Conference on Data Science and Advanced Analytics (DSAA’2023) 2 GitHub: [https://github.com/rezacsedu/DeepExplainHidim](https://github.com/rezacsedu/DeepExplainHidim)  \nonly introduces unwanted noise but also increases computational complexity as the data becomes sparser. With increased modelling complexity involving hundreds of features and their interactions, making a general conclusion or interpreting the black-box model’s outcome becomes increasingly difficult, whereas many approaches do not take into account understanding the inner structure of opaque models.  \nIn contrast, DNNs benefit from higher pattern recognition capabilities during learning useful representation from such datasets. With multiple hidden layers and non-linear activation functions within layers, autoencoder (AEs) can model complex and higher-order feature interactions. Learning non-linear mappings allow embedding input feature space into a lowerdimensional laten","cbCaisEVWWdckK8i","https://ap.wps.com/l/cbCaisEVWWdckK8i","pdf",2169373,2,1,10,"English","en",105,"# Abstract\n# Introduction\n## Curse of dimensionality and opaque models\n## Need for explainable AI\n## Global vs local interpretability","[{\"question\":\"Why are high-dimensional datasets challenging for machine learning models?\",\"answer\":\"High dimensionality can include many irrelevant features, introducing unwanted noise and increasing computational complexity, which makes interpretation and deriving conclusions more difficult.\"},{\"question\":\"What does the proposed method do to interpret a black-box model?\",\"answer\":\"It trains a black-box on the full feature space, uses probing and perturbing to identify top-k important features for global explainability, then trains an interpretable surrogate on the top-k feature space to approximate the black-box.\"},{\"question\":\"How are local explanations produced in the approach?\",\"answer\":\"The method derives decision rules and counterfactuals from the surrogate model to provide local decisions explaining individual outcomes.\"}]","Interpreting Black-box Machine Learning Models for High Dimensional Datasets - Research and Analysis | PDF",1785942902,25,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"interpreting-black-box-machine-learning-models-for-high-dimensional-datasets-research-and-analysis","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/interpreting-black-box-machine-learning-models-for-high-dimensional-datasets-research-and-analysis/127911/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are high-dimensional datasets challenging for machine learning models?","Question",{"text":76,"@type":77},"High dimensionality can include many irrelevant features, introducing unwanted noise and increasing computational complexity, which makes interpretation and deriving conclusions more difficult.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What does the proposed method do to interpret a black-box model?",{"text":81,"@type":77},"It trains a black-box on the full feature space, uses probing and perturbing to identify top-k important features for global explainability, then trains an interpretable surrogate on the top-k feature space to approximate the black-box.",{"name":83,"@type":74,"acceptedAnswer":84},"How are local explanations produced in the approach?",{"text":85,"@type":77},"The method derives decision rules and counterfactuals from the surrogate model to provide local decisions explaining individual outcomes.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":22,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]