[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125617-en":3,"doc-seo-125617-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125617,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","A Holistic Assessment of the Reliability of Machine Learning Systems","As machine learning systems increasingly operate in high-stakes domains such as healthcare, transportation, military, and national security, reliability concerns intensify. Performance may degrade due to adversarial attacks or environmental changes, producing overconfident predictions, missed input faults, and weak generalization. The paper introduces a holistic reliability assessment framework covering five properties: in-distribution accuracy, distribution-shift robustness, adversarial robustness, calibration, and out-of-distribution detection. It also defines a reliability score and reports analysis over 500 models, showing trade-offs and cross-metric improvements.","arXiv :2307 . 10586v 1 [ cs .LG] 20 Jul 2023  \nA Holistic Assessment of the Reliability of Machine Learning Systems  \nA Holistic Assessment of the Reliability of Machine Learning  \nSystems  \nAnthony Corso [acorso@stanford.edu](acorso@stanford.edu)  \nAeronautics and Astronautics, Stanford University, Stanford, CA 94305, USA  \nDavid Karamadian [dk11@alumni.stanford.edu](dk11@alumni.stanford.edu)  \nStatistics, Stanford University, Stanford, CA 94305, USA  \nRomeo Valentin [romeov@stanford.edu](romeov@stanford.edu)  \nAeronautics and Astronautics, Stanford University, Stanford, CA 94305, USA  \nMary Cooper [marycoop@alumni.stanford.edu](marycoop@alumni.stanford.edu)  \nAeronautics and Astronautics, Stanford University, Stanford, CA 94305, USA  \nMykel J. Kochenderfer [mykel@stanford.edu](mykel@stanford.edu)  \nAeronautics and Astronautics, Stanford University, Stanford, CA 94305, USA  \nAbstract  \nAs machine learning (ML) systems increasingly permeate high-stakes settings such as healthcare, transportation, military, and national security, concerns regarding their reliability have emerged. Despite notable progress, the performance of these systems can signi􀀌cantly diminish due to adversarial attacks or environmental changes, leading to overcon􀀌dent predictions, failures to detect input faults, and an inability to generalize in unexpected scenarios. This paper proposes a holistic assessment methodology for the reliability of ML systems. Our framework evaluates 􀀌ve key properties: in-distribution accuracy, distribution-shift robustness, adversarial robustness, calibration, and out-of-distribution detection. A reliability score is also introduced and used to assess the overall system reliability. To provide insights into the performance of di􀀋erent algorithmic approaches, we identify and categorize state-of-the-art techniques, then evaluate a selection on real-world tasks using our proposed reliability metrics and reliability score. Our analysis of over  \n500 models reveals that designing for one metric does not necessarily constrain others but certain algorithmic techniques can improve reliability across multiple metrics simultaneously. This study contributes to a more comprehensive understanding of ML reliability and provides a roadmap for future research and development.  \n1. Introduction  \nRapid progress in the capabilities of machine learning (ML) technology has prompted its nascent or proposed use in high-stakes settings such as transportation (Ma et al., 2020; Sridhar, 2020), healthcare (Jiang et al., 2017; K.-H. Yu et al., 2018), military (Morgan et al., 2020), and national security (Sayler & Hoadley, 2020) . Despite this progress, ML components still su􀀋er from various reliability issues, which compromise the ability for the  \nCorso, Karamadian, Valentin, Cooper & Kochenderfer  \nsystem to perform its intended function during deployment. Adversarial or natural environmental changes can result in signi􀀌cant drops in performance, while poor calibration of uncertainty estimates can lead to overcon􀀌dent predictions even in scenarios where accuracy is compromised. In the extreme case, ML components fail to detect input faults that should cause them to cease operation. So how should we evaluate ML models?  \nRecent guidance and regulation on arti􀀌cial intelligence (AI) systems (3000.09, 2023; European Commission, 2021; Executive Order, 2020; NIST, 2022; OECD, 2019), which are typically enabled by ML technology, have outlined several key requirements for ensuring their safety, fairness, and security. While guidance and regulations vary, we summarize many of the key points addressed as follows. AI systems must have a well-de􀀌ned objective and produce accurate, reliable outputs that generalize well and be robust to new scenarios. In the event of unexpected scenarios, AI systems must fail gracefully and minimize harm. To ensure safety, AI systems should undergo rigorous validation to identify failure modesand corresponding harms and be regular","cbCaiv6muyhH87ta","https://ap.wps.com/l/cbCaiv6muyhH87ta","pdf",1720259,1,49,"English","en",105,"# Introduction\n## Motivation and high-stakes deployments\n## Regulatory and safety requirements for AI systems\n## Reliability properties of ML components\n# Holistic assessment methodology\n## Five key reliability properties\n## Reliability score and evaluation across algorithms\n## Experimental analysis over real-world tasks","[{\"question\":\"Why is ML system reliability important in high-stakes settings?\",\"answer\":\"As ML is deployed in areas like healthcare, transportation, military, and national security, performance drops from adversarial attacks or environmental changes can cause overconfident errors and missed input faults.\"},{\"question\":\"What five properties does the proposed framework evaluate?\",\"answer\":\"The framework evaluates in-distribution accuracy, distribution-shift robustness, adversarial robustness, calibration, and out-of-distribution detection.\"},{\"question\":\"How does the paper assess overall reliability?\",\"answer\":\"It introduces a reliability score to quantify overall system reliability and uses it to evaluate algorithmic approaches on real-world tasks.\"}]","A Holistic Assessment of the Reliability of Machine Learning Systems | PDF",1785900237,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-holistic-assessment-of-the-reliability-of-machine-learning-systems","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-holistic-assessment-of-the-reliability-of-machine-learning-systems/125617/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is ML system reliability important in high-stakes settings?","Question",{"text":75,"@type":76},"As ML is deployed in areas like healthcare, transportation, military, and national security, performance drops from adversarial attacks or environmental changes can cause overconfident errors and missed input faults.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What five properties does the proposed framework evaluate?",{"text":80,"@type":76},"The framework evaluates in-distribution accuracy, distribution-shift robustness, adversarial robustness, calibration, and out-of-distribution detection.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper assess overall reliability?",{"text":84,"@type":76},"It introduces a reliability score to quantify overall system reliability and uses it to evaluate algorithmic approaches on real-world tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]