[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117257-en":3,"doc-seo-117257-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},117257,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Unveiling the robustness of machine learning families - Paper","Machine learning systems are often assessed only on clean, curated datasets, so reported performance may not match real-world robustness when data distributions shift between training and deployment. A central contributor is instance difficulty: some inputs trigger failures that are more unexpected than others. This work introduces an item response theory-based framework to estimate instance difficulty for supervised tasks and to evaluate robustness by measuring performance deviations under perturbations that emulate noise and deployment variability. The results enable a taxonomy linking model-family robustness to instance difficulty, highlighting strengths and vulnerabilities.","PAPER • OPEN ACCESS  \nUnveiling the robustness of machine learning families  \nTo cite this article: R Fabra-Boluda et al 2024 Mach. Learn. : Sci. Technol. 5 035040  \nView the article online for updates and enhancements.  \nYou may also like  \n-On the robustness of deep learning-based lung-nodule classification for CT images with respect to image noise  \nChenyang Shen, Min-Yu Tsai, Liyuan Chen et al.  \n-Improving robustness of a deep learningbased lung-nodule classification model of CT images with respect to image noise  \nYin Gao, Jennifer Xiong, Chenyang Shenet al.  \n-Improving the attack tolerance of scalefree networks by adding and hiding edges  \nYue Zhuo, Yunfeng Peng, Chang Liu et al.  \nThis content was downloaded from IP address [158.42.234.60](158.42.234.60) on 13/09/2024 at 08:28  \n Mach. Learn.: Sci. Technol. 5 (2024) 035040 [https://doi.org/10.1088/2632-2153/ad62ab](https://doi.org/10.1088/2632-2153/ad62ab)  \nOPEN ACCESS  \nRECEIVED  \n22 December 2023  \nREVISED  \n1 May 2024  \nACCEPTED FOR PUBLICATION 12 July 2024  \nPUBLISHED  \n8 August 2024  \nOriginal Content from this work may be used under the terms of the  \nCreative Commons Attribution 4 .0 licence.  \nAny further distribution of this work must maintain attribution to the author(s) and the title of the work, journal citation and DOI.  \nPAPER  \nUnveiling the robustness of machine learning families  \nR Fabra-Boluda∗􀁂, C Ferri, M J Ramírez-Quintana and F Martínez-Plumed􀁂 Valencian Research Institute for Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain ∗ Author to whom any correspondence should be addressed.  \n[E-mail: rafabbo@dsic.upv.es](E-mail: rafabbo@dsic.upv.es), [cferri@dsic.upv.es](cferri@dsic.upv.es), mramirez@dsic.upv.es and fmartinez@dsic.upv.es  \nKeywords: robustness, noise, instance difficulty, supervised learning, item response theory  \nAbstract  \nThe evaluation of machine learning systems has typically been limited to performance measures on clean and curated datasets, which may not accurately reflect their robustness in real-world situations where data distribution can vary from learning to deployment, and where truthfully predict some instances could be more difficult than others. Therefore, a key aspect in understanding robustness is instance difficulty, which refers to the level of unexpectedness of system failure on a specific instance. We present a framework that evaluates the robustness of different ML models using item response theory-based estimates of instance difficulty for supervised tasks. This framework evaluates performance deviations by applying perturbation methods that simulate noise and variability in deployment conditions. Our findings result in the development of a comprehensive taxonomy of ML techniques, based on both the robustness of the models and the difficulty of the instances, providing a deeper understanding of the strengths and limitations of specific families of ML models. This study is a significant step towards exposing vulnerabilities of particular families of ML models.  \n1. Introduction  \nThe proliferation of machine learning (ML) systems has transformed various fields, including medicine, finance, social media, and autonomous transport, integrating into our daily lives and shaping  \ndecision-making processes. With the growing influence of these systems, it is imperative to have reliable and robust ML systems that can function correctly under different conditions and inputs [1] . Robustness, in this context, refers to the ability of a ML system consistently maintain its predictions despite variations or perturbations [2] .  \nTraditional evaluations of ML robustness have predominantly focused on resistance to adversarial examples—deliberately manipulated inputs designed to trick models into making incorrect predictions. These studies often involve the introduction of noise during training and testing phases to test the model’s defences against such attacks [1] . These examples are commonly know","cbCaijyLv5rKGPp8","https://ap.wps.com/l/cbCaijyLv5rKGPp8","pdf",7906590,1,33,"English","en",105,"# Introduction\n## Robustness in real-world variation\n## Adversarial examples vs prediction consistency\n## Adding instance difficulty to robustness evaluation","[{\"question\":\"What does “instance difficulty” mean in this study?\",\"answer\":\"Instance difficulty describes how unexpected a system failure is for a specific input instance, reflecting intrinsic or extrinsic challenge beyond mere performance on curated data.\"},{\"question\":\"How does the proposed framework evaluate robustness?\",\"answer\":\"It estimates instance difficulty using item response theory and then measures performance deviations when applying perturbation methods that simulate noise and variability in deployment conditions.\"},{\"question\":\"Why do the authors argue existing robustness evaluations are incomplete?\",\"answer\":\"They note that many studies emphasize resistance to adversarial examples, which can overshadow prediction consistency under input variations, and they further emphasize that instance difficulty has not been used as an evaluation criterion for robustness.\"}]",1785674738,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"unveiling-the-robustness-of-machine-learning-families-paper","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/unveiling-the-robustness-of-machine-learning-families-paper/117257/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What does “instance difficulty” mean in this study?","Question",{"text":74,"@type":75},"Instance difficulty describes how unexpected a system failure is for a specific input instance, reflecting intrinsic or extrinsic challenge beyond mere performance on curated data.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed framework evaluate robustness?",{"text":79,"@type":75},"It estimates instance difficulty using item response theory and then measures performance deviations when applying perturbation methods that simulate noise and variability in deployment conditions.",{"name":81,"@type":72,"acceptedAnswer":82},"Why do the authors argue existing robustness evaluations are incomplete?",{"text":83,"@type":75},"They note that many studies emphasize resistance to adversarial examples, which can overshadow prediction consistency under input variations, and they further emphasize that instance difficulty has not been used as an evaluation criterion for robustness.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]