[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128305-en":3,"doc-seo-128305-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128305,962085564381,"Clementine","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Explainability in Machine Learning Models - A Shap-Driven Framework for Misclassification Analysis and Feature Selection","Increasing complexity of high-performing machine learning models can produce opaque “black-box” behavior, especially when misclassifying instances. This dissertation proposes a unified SHAP-driven methodology to improve explainability by analyzing misclassifications first and then optimizing feature sets using those insights, even for pre-optimized industrial models. A Misclassification Explanation Framework uses SHAP values and instance clustering to generate hierarchical, error-type-aware explanations and quantify feature contributions under distinct data conditions. An Impact-Based Recursive Feature Selection strategy, using a metric-specific Net Impact score, iteratively identifies effective feature subsets while preserving or improving predictive performance and improving parsimony via a LightGBM tree-dropping variant. Experiments on three real-world datasets validate improved error diagnostics and feature refinement for gradient boosting models. The work also connects model transparency improvements to relevant UN SDGs.","FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO  \nExplainability in Machine Learning Models: A Shap-Driven Framework for Misclassification Analysis and Feature  \nSelection  \nHenrique Oliveira Silva  \nMestrado em Engenharia Informática e Computação Supervisor: Prof. Ana Paula Rocha Second Supervisor: Maria Manuel Castro  \nJuly 23, 2025  \nExplainability in Machine Learning Models: AShap-Driven Framework for Misclassification Analysis  \nand Feature Selection  \nHenrique Oliveira Silva  \nMestrado em Engenharia Informática e Computação  \nApproved in oral examination by the committee:  \nPresident: Prof. Rui Camacho  \nReferee: Prof. Ana Paula Rocha  \nReferee: Dr. André Carreiro  \nJuly 23, 2025  \nAbstract  \nThe increasing complexity of high-performing machine learning (ML) models often results in\"black-box\" systems whose decision-making processes are opaque, particularly when they misclassify instances. This dissertation introduces a unified, SHAP-driven approach to enhance ML model explainability by first providing deep insights into misclassifications and then leveraging these insights for feature set optimization, even on pre-optimized industrial models.  \nThe core methodology begins with a Misclassification Explanation Framework (MEF) that employs SHAP values and instance clustering to dissect model errors. This framework generates hierarchical explanations, from specific error clusters to global patterns, quantifying feature contributions to different types of misclassifications and identifying associated data conditions. Building upon the understanding of feature impacts derived from this misclassification analysis, an Impact-Based Recursive Feature Selection (RFS) strategy was developed. This RFS component utilizes a metric-specific \"Net Impact\" score, also informed by SHAP values related to correct and incorrect predictions, to iteratively identify optimal feature subsets. The goal is to maintain or enhance predictive performance (according to metrics like accuracy or F1-score) while improving model parsimony. An efficient tree-dropping variant for LightGBM models is also presented as part of the RFS strategy.  \nEvaluations on three diverse real-world datasets demonstrate the utility of this integrated approach in providing error diagnostics and effectively refining feature sets for complex gradient boosting models. The findings contribute practical tools for deeper, more cohesive model understanding and targeted improvement in the field of explainable AI.  \nUN Sustainable Development Goals  \nThe United Nations Sustainable Development Goals (SDGs) provide a global framework to achieve a better and more sustainable future for all. It includes 17 goals to address the world’s most pressing challenges, including poverty, inequality, climate change, environmental degradation, peace, and justice.  \nThis dissertation contributes to this global agenda by developing methodologies that enhance machine learning models’ explainability, reliability, and efficiency. Such advancements are crucial for fostering trust in Artificial Intelligence, enabling more responsible innovation, and ensuring that complex AI systems can be developed and deployed in a manner that is transparent and accountable. This work supports the creation of more robust and understandable AI by providing tools to dissect model errors and optimize their design, a foundational element for sustainable technological progress.  \nThe Sustainable Development Goals (SDGs) to which this work most directly relates are:  \nSDG 9 Industry, Innovation, and Infrastructure: Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation.  \nSDG 16 Peace, Justice, and Strong Institutions: Promote peaceful and inclusive societies for sustainable development, provide access to justice for all and build effective, accountable and inclusive institutions at all levels.  \niii  \n\n| SDG | Target | Contribution of this Dissertation | Potential Performance ","cbCaigzua2ltmPhR","https://ap.wps.com/l/cbCaigzua2ltmPhR","pdf",2231321,3,1,116,"English","en",105,"# Abstract\n## Misclassification Explanation Framework (MEF)\n## Impact-Based Recursive Feature Selection (RFS)\n## Tree-dropping variant for LightGBM\n## Evaluations on real-world datasets\n## Connection to UN Sustainable Development Goals (SDGs)","[{\"question\":\"What is the main contribution of the dissertation?\",\"answer\":\"It introduces a unified SHAP-driven framework that explains misclassifications and then uses those insights to optimize feature selection, including for pre-optimized industrial models.\"},{\"question\":\"How does the Misclassification Explanation Framework (MEF) work?\",\"answer\":\"MEF combines SHAP values with instance clustering to dissect model errors and produce hierarchical explanations from error clusters to global patterns, quantifying feature contributions for different misclassification types.\"},{\"question\":\"What is Impact-Based Recursive Feature Selection (RFS)?\",\"answer\":\"RFS uses a metric-specific Net Impact score informed by SHAP values for both correct and incorrect predictions to iteratively select feature subsets that maintain or improve metrics like accuracy or F1-score.\"}]","Explainability in Machine Learning Models - A Shap-Driven Framework for Misclassification Analysis and Feature Selection | PDF",1785946723,292,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"explainability-in-machine-learning-models-a-shap-driven-framework-for-misclassification-analysis-and-feature-selection","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/explainability-in-machine-learning-models-a-shap-driven-framework-for-misclassification-analysis-and-feature-selection/128305/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main contribution of the dissertation?","Question",{"text":76,"@type":77},"It introduces a unified SHAP-driven framework that explains misclassifications and then uses those insights to optimize feature selection, including for pre-optimized industrial models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the Misclassification Explanation Framework (MEF) work?",{"text":81,"@type":77},"MEF combines SHAP values with instance clustering to dissect model errors and produce hierarchical explanations from error clusters to global patterns, quantifying feature contributions for different misclassification types.",{"name":83,"@type":74,"acceptedAnswer":84},"What is Impact-Based Recursive Feature Selection (RFS)?",{"text":85,"@type":77},"RFS uses a metric-specific Net Impact score informed by SHAP values for both correct and incorrect predictions to iteratively select feature subsets that maintain or improve metrics like accuracy or F1-score.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]