[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121606-en":3,"doc-seo-121606-105":31,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},121606,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","CHG Shapley - Efficient Data Valuation and Selection towards Trustworthy Machine Learning","Understanding the decision-making process of machine learning models is essential for trustworthy machine learning. Data Shapley quantifies each datum’s contribution to model performance, but repeated model retraining makes it computationally impractical for large datasets. This work introduces the CHG (compound of Hardness and Gradient) utility to approximate subset utility at every training epoch. A closed-form CHG Shapley value is derived for each data point, reducing computation to a single retraining and improving efficiency over marginal methods. CHG Shapley is further used for real-time data selection and evaluated on standard, label-noise, and class-imbalance datasets to identify both high-value and noisy data, supported by available code.","arXiv :2406 . 11730v3 [ cs .GT] 22 Jan 2025  \nCHG SHAPLEY: EFFICIENT DATA VALUATION AND SELECTION TOWARDS TRUSTWORTHY MACHINE LEARNING  \nHuaiguang Cai  \nState Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences [caihuaiguang@gmail.com](caihuaiguang@gmail.com)  \nABSTRACT  \nUnderstanding the decision-making process of machine learning models is crucial for ensuring trustworthy machine learning. Data Shapley, a landmark study on data valuation, advances this understanding by assessing the contribution of each datum to model performance. However, the resource-intensive and timeconsuming nature of multiple model retraining poses challenges for applying Data Shapley to large datasets. To address this, we propose the CHG (compound of Hardness and Gradient) utility function, which approximates the utility of each data subset on model performance in every training epoch. By deriving the closedform Shapley value for each data point using the CHG utility function, we reduce the computational complexity to that of a single model retraining, achieving a quadratic improvement over existing marginal contribution-based methods.  \nWe further leverage CHG Shapley for real-time data selection, conducting experiments across three settings: standard datasets, label noise datasets, and class imbalance datasets. These experiments demonstrate its effectiveness in identifying high-value and noisy data. By enabling efficient data valuation, CHG Shapley promotes trustworthy model training through a novel data-centric perspective. Our codes are available at [https://github.com/caihuaiguang/](https://github.com/caihuaiguang/)  \nCHG-Shapley-for-Data-Valuation and [https://github.com/](https://github.com/)[ ](https://github.com/)caihuaiguang/CHG-Shapley-for-Data-Selection.  \n1 INTRODUCTION  \nThe central problem of trustworthy machine learning is explaining the decision-making process of models to enhance the transparency of data-driven algorithms. However, the high complexity of machine learning model training and inference processes obscures an intuitive understanding of their internal mechanisms. Approaching trustworthy machine learning from a data-centric perspective (Liu et al., 2023) offers a new perspective for research. For trustworthy model inference, a representative algorithm is the SHAP (Lundberg & Lee, 2017), which quantitatively attributes model outputs to input features, clarifying which features influence specific results the most. SHAP and its variants(Kwon & Zou, 2022b) are widely applied in data analysis and healthcare. For trustworthy model training, the Data Shapley algorithm (Ghorbani & Zou, 2019) stands out. It quantitatively attributes a model’s performance to each training data point, identifying valuable data that improves performance and noisy data that degrades it. Data valuation reveals how much each training sample affects model performance, serving as a foundation for tasks like data selection, acquisition, and cleaning, while also facilitating the creation of data markets (Mazumder et al., 2023) .  \nThe impressive effectiveness of both SHAP and Data Shapley algorithms is rooted in the Shapley value (Shapley, 1953) . The unique feature of the Shapley value lies in its ability to accurately and fairly allocate contributions to each factor in decision-making processes where multiple factors interact with each other. However, the exact computation of the Shapley value is O(2n ), where nis the number of factors. Even with estimation techniques such as linear least squares regression (Lundberg & Lee, 2017) or Monte Carlo methods (Ghorbani & Zou, 2019), the efficiency of SHAP and Data Shapley algorithms struggles to scale with high-dimensional inputs or large datasets.  \nIn this paper, we focus on the efficiency problem of data valuation on large-scale datasets. Instead of investigating more efficient and robust algorithms to approximate the Shapley value as in (Lundberg ","cbCaiicuuPcIXDqQ","https://ap.wps.com/l/cbCaiicuuPcIXDqQ","pdf",860647,3,1,20,"English","en",105,"# Introduction\n## Trustworthy machine learning from a data-centric perspective\n## Shapley value and scalability challenges\n# Proposed Method: CHG Shapley\n## CHG utility function for efficient valuation\n## Closed-form Shapley value derivation\n# Real-Time Data Selection Experiments\n## Standard datasets\n## Label noise datasets\n## Class imbalance datasets\n# Contributions\n## Efficient large-scale data valuation\n## Parameter-free data-centric improvement","[{\"question\":\"What problem does CHG Shapley address in Data Shapley?\",\"answer\":\"Data Shapley requires multiple model retraining to evaluate subsets, which becomes resource-intensive on large datasets. CHG Shapley reduces the computation by using an epoch-wise CHG utility and deriving a closed-form Shapley value per data point.\"},{\"question\":\"How does the CHG utility function speed up data valuation?\",\"answer\":\"The CHG utility approximates the utility of each data subset on model performance in every training epoch. This enables computing Shapley values for individual data points with the complexity of a single model retraining.\"},{\"question\":\"How is CHG Shapley used for data selection, and what do experiments show?\",\"answer\":\"CHG Shapley supports real-time data selection by identifying high-value and noisy samples efficiently. Experiments across standard, label-noise, and class-imbalance settings demonstrate its effectiveness in selecting useful data.\"}]","CHG Shapley - Efficient Data Valuation and Selection towards Trustworthy Machine Learning | PDF",1785736449,50,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":29},"chg-shapley-efficient-data-valuation-and-selection-towards-trustworthy-machine-learning","",{"@graph":37,"@context":85},[38,54,68],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/chg-shapley-efficient-data-valuation-and-selection-towards-trustworthy-machine-learning/121606/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CHG Shapley address in Data Shapley?","Question",{"text":75,"@type":76},"Data Shapley requires multiple model retraining to evaluate subsets, which becomes resource-intensive on large datasets. CHG Shapley reduces the computation by using an epoch-wise CHG utility and deriving a closed-form Shapley value per data point.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the CHG utility function speed up data valuation?",{"text":80,"@type":76},"The CHG utility approximates the utility of each data subset on model performance in every training epoch. This enables computing Shapley values for individual data points with the complexity of a single model retraining.",{"name":82,"@type":73,"acceptedAnswer":83},"How is CHG Shapley used for data selection, and what do experiments show?",{"text":84,"@type":76},"CHG Shapley supports real-time data selection by identifying high-value and noisy samples efficiently. Experiments across standard, label-noise, and class-imbalance settings demonstrate its effectiveness in selecting useful data.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":30,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":47,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":47,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":47,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]