[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125850-en":3,"doc-seo-125850-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125850,1099523882367,"Hazel","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Data Banzhaf - A Robust Data Valuation Framework for Machine Learning - paper","Data valuation supports machine learning workflows such as improving data quality and enabling economic incentives for data sharing. This work analyzes robustness when model performance scores are noisy, which arises naturally from stochastic training like stochastic gradient descent. The study shows that randomness can cause existing valuation notions, including Shapley value and leave-one-out error, to yield inconsistent data value rankings across runs. It introduces safety margin to quantify robustness and proves that the Banzhaf value attains the largest margin among semivalues while offering efficient estimation and strong performance on multiple ML tasks.","arXiv :2205 . 15466v7 [ cs .LG] 18 Dec 2023  \nData Banzhaf: A Robust Data Valuation Framework for Machine  \nLearning  \nJiachen T. Wang 1 and Ruoxi Jia2  \n1 Princeton University  \n2Virginia Tech  \n[tianhaowang@princeton.edu](tianhaowang@princeton.edu) , [ruoxijia@vt.edu](ruoxijia@vt.edu)  \nAbstract  \nData valuation has wide use cases in machine learning, including improving data quality and creating economic incentives for data sharing. This paper studies the robustness of data valuation to noisy model performance scores. Particularly, we find that the inherent randomness of the widely used stochastic gradient descent can cause existing data value notions (e.g., the Shapley value and the Leave-one-out error) to produce inconsistent data value rankings across different runs. To address this challenge, we introduce the concept of safety margin, which measures the robustness of a data value notion. We show that the Banzhaf value, a famous value notion that originated from cooperative game theory literature, achieves the largest safety margin among all semivalues (a class of value notions that satisfy crucial properties entailed by ML applications and include the famous Shapley value and Leave-one-out error) . We propose an algorithm to efficiently estimate the Banzhaf value based on the Maximum Sample Reuse (MSR) principle. Our evaluation demonstrates that the Banzhaf value outperforms the existing semivalue-based data value notions on several ML tasks such as learning with weighted samples and noisy label detection. Overall, our study suggests that when the underlying ML algorithm is stochastic, the Banzhaf value is a promising alternative to the other semivalue-based data value schemes given its computational advantage and ability to robustly differentiate data quality.1  \n1 Introduction  \nData valuation, i.e., quantifying the usefulness of a data source, is an essential component in developing machine learning (ML) applications. For instance, evaluating the worth of data plays a vital role in cleaning bad data [Tang et al. , 2021 , Karlaš et al. , 2022] and understanding the model’s test-time behavior [Koh and Liang, 2017] . Furthermore, determining the value of data is crucial in creating incentives for data sharing and in implementing policies regarding the monetization of personal data [Ghorbani and Zou, 2019 , Zhu et al. , 2019] .  \nDue to the great potential in real applications, there has been a surge of research efforts on developing data value notions for supervised ML [Jia et al. , 2019b, Ghorbani and Zou, 2019 , Yan and Procaccia, 2020 , Ghorbani et al. , 2021 , Kwon and Zou, 2021 , Yoon et al. , 2020] . In the ML context, a data point’s value depends on other data points used in model training. For instance, a data point’s value will decrease if we add extra data points that are similar to the existing one into the training set. To accommodate this interplay, current data valuation techniques typically start by defining the“utility” of a set of data points, and then measure the value of an individual data point based on  \n1 Code is available at [https://github.com/Jiachen-T-Wang/data-banzhaf](https://github.com/Jiachen-T-Wang/data-banzhaf) .  \nthe change of utility when the point is added to an existing dataset. For ML tasks, the utility of adataset is naturally chosen to be the performance score (e.g., test accuracy) of a model trained on the dataset.  \nHowever, the utility scores can be noisy and unreliable. Stochastic training methods such as stochastic gradient descent (SGD) are widely adopted in ML, especially for deep learning. The models trained with stochastic methods are inherently random, and so are their performance scores. This, in turn, makes the data values calculated from the performance scores noisy. Despite being ignored in past research, we find that the noise in a typical learning process is actually substantial enough to make different runs of the same data valuation algorithm produce inconsistent","cbCainigxAwEFzlx","https://ap.wps.com/l/cbCainigxAwEFzlx","pdf",2754636,5,1,46,"English","en",105,"# Abstract\n# Introduction\n## Motivation for data valuation\n## Robustness under noisy performance scores\n# Main contributions\n## Safety margin as robustness measure\n## Banzhaf value as most robust semivalue\n## Efficient estimation via Maximum Sample Reuse","[{\"question\":\"Why can data valuation produce inconsistent value rankings in stochastic machine learning?\",\"answer\":\"Because stochastic training methods like stochastic gradient descent introduce inherent randomness, the model performance scores become noisy. This noise can change the computed data values across different runs, leading to inconsistent rankings.\"},{\"question\":\"What is safety margin in this paper?\",\"answer\":\"Safety margin measures the largest perturbation of model performance scores that can be tolerated without changing the value order of any pair of data points. It provides a mathematical way to quantify robustness.\"},{\"question\":\"How does the Banzhaf value compare to Shapley value and leave-one-out error?\",\"answer\":\"The paper shows that the Banzhaf value achieves the largest safety margin among all semivalues considered, and its safety margin is exponentially larger than those of both the Shapley value and the leave-one-out error.\"}]","Data Banzhaf - A Robust Data Valuation Framework for Machine Learning - paper | PDF",1785901577,116,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"data-banzhaf-a-robust-data-valuation-framework-for-machine-learning-paper","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/data-banzhaf-a-robust-data-valuation-framework-for-machine-learning-paper/125850/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why can data valuation produce inconsistent value rankings in stochastic machine learning?","Question",{"text":77,"@type":78},"Because stochastic training methods like stochastic gradient descent introduce inherent randomness, the model performance scores become noisy. This noise can change the computed data values across different runs, leading to inconsistent rankings.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What is safety margin in this paper?",{"text":82,"@type":78},"Safety margin measures the largest perturbation of model performance scores that can be tolerated without changing the value order of any pair of data points. It provides a mathematical way to quantify robustness.",{"name":84,"@type":75,"acceptedAnswer":85},"How does the Banzhaf value compare to Shapley value and leave-one-out error?",{"text":86,"@type":78},"The paper shows that the Banzhaf value achieves the largest safety margin among all semivalues considered, and its safety margin is exponentially larger than those of both the Shapley value and the leave-one-out error.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]