[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121340-en":3,"doc-seo-121340-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121340,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Framework for Benchmarking Machine Learning Models - Integrating Performance Metrics, Explainability Techniques, and Robustness Assessments - Completed Research Full Paper","Machine learning (ML) models are widely applied to high-stakes domains because they can process large-scale data and deliver strong predictive accuracy. Traditional benchmarking often emphasizes performance metrics such as accuracy and precision, yet these measures do not fully capture transparency and robustness as models become more complex and operate under shifting conditions. This research presents a comprehensive benchmarking framework for supervised ML models that integrates performance metrics, explainability techniques, and robustness assessments to evaluate efficiency, interpretability, and stability under noise and data shifts, supporting trustworthy, resilient model selection for organizational decision-making.","Association for Information Systems  \nAIS Electronic Library (AISeL)  \n\n| AMCIS 2025 Proceedings | Americas Conference on Information Systems\u003Cbr>(AMCIS) |\n| --- | --- |\n| August 2025\u003Cbr>Framework for Benchmarking Machine Learning Models: Integrating Performance Metrics, Explainability Techniques, and Robustness Assessments\u003Cbr>Deependra Malla\u003Cbr>Dakota State University, [deependra.malla@trojans.dsu.edu](deependra.malla@trojans.dsu.edu)\u003Cbr>Omar El-Gayar\u003Cbr>Dakota State University, [omar.el-gayar@dsu.edu](omar.el-gayar@dsu.edu)\u003Cbr>Follow this and additional works at: [https://aisel.aisnet.org/amcis2025](https://aisel.aisnet.org/amcis2025) |  |\n\nRecommended Citation  \nMalla, Deependra and El-Gayar, Omar, \"Framework for Benchmarking Machine Learning Models: Integrating Performance Metrics, Explainability Techniques, and Robustness Assessments\" (2025) . AMCIS 2025 Proceedings. 7.  \n[https://aisel.aisnet.org/amcis2025/data_science/sig_dsa/7](https://aisel.aisnet.org/amcis2025/data_science/sig_dsa/7)  \nThis material is brought to you by the Americas Conference on Information Systems (AMCIS) at AIS Electronic Library (AISeL) . It has been accepted for inclusion in AMCIS 2025 Proceedings by an authorized administrator of AIS Electronic Library (AISeL) . For more information, please contact [elibrary@aisnet.org](elibrary@aisnet.org).  \nFramework for Benchmarking Machine Learning Models: Integrating Performance Metrics, Explainability Techniques, and Robustness Assessments  \nCompleted Research Full Paper  \nDeependra Malla  \nDakota State University [deependra.malla@trojans.dsu.edu](deependra.malla@trojans.dsu.edu)  \nOmar El-Gayar  \nDakota State University [omar.el-gayar@dsu.edu](omar.el-gayar@dsu.edu)  \nAbstract  \nMachine learning (ML) models are widely used across various critical domains for their ability to process large-scale data and deliver high predictive accuracy. While traditional benchmarking of ML models focuses on performance metrics like accuracy and precision, these metrics often fall short in accounting for the model’s transparency and robustness when the models get complex with varying conditions. This research proposes a comprehensive framework for benchmarking supervised machine learning models that incorporate the model’s performance metrics, explainability techniques, and robustness assessments. The proposed framework combines traditional accuracy- and precision-based performance evaluation with explainability and robustness to ensure the model's efficiency, transparency, and stability in the presence of noise and data shifts. By holistically addressing performance, explainability, and robustness in tandem, the framework supports data-driven decision-making in selecting the model appropriate to organizational contexts, addressing stakeholder concerns related to model interpretability, trustworthiness, and resilience—particularly in critical domains such as healthcare.  \nKeywords  \nSupervised machine learning, performance, explainability, robustness  \nIntroduction  \nThe rapid advancements in machine learning technologies have given rise to their widespread adoption across various domains, including healthcare, cybersecurity, finance, and autonomous systems. ML models have demonstrated their significant capabilities in processing large-scale data, identifying complex patterns, and making accurate predictions, thereby transforming decision-making processes. Among the several types of machine learning, supervised learning models have gained prominence due to their ability to deliver high predictive accuracy on well-labeled datasets. These models are widely used for classification and regression tasks, with performance typically measured by metrics such as accuracy and F1-score. While these models have achieved high predictive performance, their growing complexity has raised concerns about transparency and interpretability, especially in critical domains.  \nTraditionally, the research in machine learning models is mostly foc","cbCaihCgTKlXLrcT","https://ap.wps.com/l/cbCaihCgTKlXLrcT","pdf",508531,1,11,"English","en",105,"# Introduction\n## Motivation for integrated benchmarking\n## Limitations of accuracy-only evaluation\n## Importance of explainability\n## Importance of robustness\n# Proposed Research Framework\n## Integrating performance, explainability, and robustness\n## Evaluating stability under noise and data shifts\n# Expected Impact\n## Supporting transparent and resilient decision-making\n## Relevance to critical domains","[{\"question\":\"Why do accuracy-focused benchmarking approaches fall short for complex ML models?\",\"answer\":\"Accuracy and precision metrics often do not account for transparency and robustness, which are increasingly needed as models become more complex and behave like black boxes under varied conditions.\"},{\"question\":\"What does the proposed framework integrate when benchmarking supervised ML models?\",\"answer\":\"It combines traditional performance evaluation (e.g., accuracy and precision) with explainability techniques and robustness assessments to evaluate efficiency, transparency, and stability.\"},{\"question\":\"How does robustness assessment support real-world deployment of ML models?\",\"answer\":\"Robustness helps models maintain performance during feature-level perturbations and unpredictable data changes, which is crucial in high-stake domains such as healthcare and finance.\"}]","Framework for Benchmarking Machine Learning Models - Integrating Performance Metrics, Explainability Techniques, and Robustness Assessments - Completed Research Full Paper | PDF",1785735141,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"framework-for-benchmarking-machine-learning-models-integrating-performance-metrics-explainability-techniques-and-robustness-assessments-completed-research-full-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/framework-for-benchmarking-machine-learning-models-integrating-performance-metrics-explainability-techniques-and-robustness-assessments-completed-research-full-paper/121340/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do accuracy-focused benchmarking approaches fall short for complex ML models?","Question",{"text":75,"@type":76},"Accuracy and precision metrics often do not account for transparency and robustness, which are increasingly needed as models become more complex and behave like black boxes under varied conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed framework integrate when benchmarking supervised ML models?",{"text":80,"@type":76},"It combines traditional performance evaluation (e.g., accuracy and precision) with explainability techniques and robustness assessments to evaluate efficiency, transparency, and stability.",{"name":82,"@type":73,"acceptedAnswer":83},"How does robustness assessment support real-world deployment of ML models?",{"text":84,"@type":76},"Robustness helps models maintain performance during feature-level perturbations and unpredictable data changes, which is crucial in high-stake domains such as healthcare and finance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]