[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123448-en":3,"doc-seo-123448-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123448,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Understanding Prediction Discrepancies in Machine Learning Classifiers","A multitude of classifiers can be trained on the same dataset and still reach similar test-time performance while learning substantially different decision patterns. This mismatch, termed prediction discrepancies, leaves practitioners unable to observe where models agree, where they differ, and what limits apply. Choosing a single model then drives concrete outcomes for instances in disagreement regions, potentially harming fairness or opportunity. The work analyzes discrepancies within a pool of best-performing models and proposes a model-agnostic method (DIG) to capture and locally explain discrepancies via intervals, supporting informed model selection and mitigation actions.","arXiv :2104 .05467v1 [ cs .LG] 12 Apr 2021  \nUnderstanding Prediction Discrepancies in Machine Learning Classi􀀌ers  \nXavier Renard 1∗, Thibault Laugel 1􀀃 , and Marcin Detyniecki 1 ;2 ;3  \n1 AXA, Paris, France  \nfxavier.renard,[thibault.laugel](thibault.laugelg@axa.com)[g](thibault.laugelg@axa.com)[@axa.com](thibault.laugelg@axa.com)  \n2 Sorbonne Universit􀀓e, CNRS, LIP6, F-75005, Paris, France  \n3 Polish Academy of Science, IBS PAN, Warsaw, Poland  \nAbstract. A multitude of classi􀀌ers can be trained on the same data to achieve similar performances during test time, while having learned signi􀀌cantly di􀀋erent classi􀀌cation patterns. This phenomenon, which we call prediction discrepancies, is often associated with the blind selection of one model instead of another with similar performances. When making a choice, the machine learning practitioner has no understanding on the di􀀋erences between models, their limits, where they agree and where they don't. But his/her choice will result in concrete consequences for instances to be classi􀀌ed in the discrepancy zone, since the 􀀌nal decision will be based on the selected classi􀀌cation pattern. Besides the arbitrary nature of the result, a bad choice could have further negative consequences such as loss of opportunity or lack of fairness. This paper proposes to address this question by analyzing the prediction discrepancies in a pool of best-performing models trained on the same data. A model-agnostic algorithm, DIG, is proposed to capture and explain discrepancies locally, to enable the practitioner to make the best educated decision when selecting a model by anticipating its potential undesired consequences. All the code to reproduce the experiments is available4  \nKeywords: machine learning interpretability · model discrepancy  \n1 Introduction  \nThe machine learning practice leverages a large variety of models and techniques to tackle, among others, classi􀀌cation tasks. The optimization of all the parameters and hyper-parameters involved leads to an arbitrary large number of models that turn to achieve similar classi􀀌cation performances, while having learned signi􀀌cantly di􀀋erent classi􀀌cation patterns. One of the intuitions behind why this phenomenon arises is the fact that any dataset is an imperfect sampling of aclassi􀀌cation task and the evaluation of the performances is done over a non exhaustive validation set.  \n∗ equal contribution  \n4[https://github.com/axa-rev-research/discrepancies-in-machine-learning](https://github.com/axa-rev-research/discrepancies-in-machine-learning)  \n2 X. Renard et al.  \nFar from being an insigni􀀌cant issue, a simple experiment shows that bestperforming classi􀀌ers, on the same task and with less than a 5% di􀀋erence inclassi􀀌cation performance, disagree over 10:64% to 28:81% of validation instanceson the tested datasets5 .  \nThis phenomenon, that we call prediction discrepancies, questions the selection of one model over others with similar performances. Indeed, this choice is made blindly, since the di􀀋erences between models and their limits are generally not observable. This is all the more problematic at a time of widespread development of machine learning-based applications, when society calls for an increased level of responsibility around the deployment of AI systems. Prediction discrepancies may result in concrete negative consequences (e.g. loss of opportunity, fairness), and should therefore be addressed. Beyond our present issue of prediction discrepancies, the paradigm of selecting the best performing model solely in terms of predictive performances is being increasingly questioned by the machine learning community [16,5] .  \nDiscrepancies in models have not always been seen as an issue: aggregating the diverse predictions of weak but diverse classi􀀌ers is for instance the principle of ensemble learning. Nonetheless, more recently, undesired consequences have been identi􀀌ed, such as the threat of fairwashing and explanation manipulation [1,2","cbCaioNMhifnXQRi","https://ap.wps.com/l/cbCaioNMhifnXQRi","pdf",1214317,1,20,"English","en",105,"# Introduction\n## Prediction discrepancies and their consequences\n## Limitations of existing assessment tools\n# Proposed approach\n## DIG: Discrepancy Intervals Generation\n## Model-agnostic explanations and computation\n# Experimental protocol","[{\"question\":\"What are prediction discrepancies in machine learning classifiers?\",\"answer\":\"Prediction discrepancies refer to cases where multiple classifiers trained on the same data achieve similar test performance but learn significantly different classification patterns, leading to disagreement on a portion of validation instances.\"},{\"question\":\"Why is blind model selection a problem when predictions disagree?\",\"answer\":\"Blindly choosing one model prevents practitioners from knowing where models agree or diverge and what limitations apply, which can produce concrete negative impacts for instances in the discrepancy zone, such as unfairness or loss of opportunity.\"},{\"question\":\"How does the proposed DIG method help practitioners?\",\"answer\":\"DIG captures and explains discrepancies locally using discrepancy intervals, allowing practitioners to select among candidate classifiers more responsibly and to take actions on instances affected by disagreements.\"}]","Understanding Prediction Discrepancies in Machine Learning Classifiers | PDF",1785816580,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"understanding-prediction-discrepancies-in-machine-learning-classifiers","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/understanding-prediction-discrepancies-in-machine-learning-classifiers/123448/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are prediction discrepancies in machine learning classifiers?","Question",{"text":75,"@type":76},"Prediction discrepancies refer to cases where multiple classifiers trained on the same data achieve similar test performance but learn significantly different classification patterns, leading to disagreement on a portion of validation instances.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is blind model selection a problem when predictions disagree?",{"text":80,"@type":76},"Blindly choosing one model prevents practitioners from knowing where models agree or diverge and what limitations apply, which can produce concrete negative impacts for instances in the discrepancy zone, such as unfairness or loss of opportunity.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed DIG method help practitioners?",{"text":84,"@type":76},"DIG captures and explains discrepancies locally using discrepancy intervals, allowing practitioners to select among candidate classifiers more responsibly and to take actions on instances affected by disagreements.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]