[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117260-en":3,"doc-seo-117260-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},117260,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","ECOVAL - An Efficient Data Valuation Framework for Machine Learning","Quantifying data value in a machine learning workflow supports better decisions in model development, data pricing, and governance. Existing Shapley value-based data valuation methods are computationally expensive because they require many repeated training runs to approximate Shapley contributions. EcoVal introduces an efficient framework that assigns value at a cluster level for groups of similar samples, propagates value to cluster members, and derives overall worth via intrinsic and extrinsic value using a production-function formulation. The paper includes formal proof, explains acceleration mechanisms, and demonstrates effectiveness for both in-distribution and out-of-sample data. Code is provided.","ECOVAL: AN EFFICIENT DATA VALUATION FRAMEWORK FOR  \nMACHINE LEARNING  \narXiv :2402 .09288v 5 [ cs .LG] 9 Jul 2024  \nAyush K Tarun∗  \nRespAI Lab, India [ayushtarun210@gmail.com](ayushtarun210@gmail.com)  \nHong Ming Tan  \nNUS Business School National University of Singapore [thm@nus.edu.sg](thm@nus.edu.sg)  \nVikram S Chundawat∗  \nRespAI Lab, India [vikram2000b@gmail.com](vikram2000b@gmail.com)  \nMurari Mandal †  \nRespAI Lab, KIIT Bhubaneswar, India [murari.mandalfcs@kiit.ac.in](murari.mandalfcs@kiit.ac.in)  \nBowei Chen  \nAdam Smith Business School University of Glasgow [bowei.chen@glasgow.ac.uk](bowei.chen@glasgow.ac.uk)  \nMohan Kankanhalli  \nSchool of Computing National University of Singapore [mohan@comp.nus.edu.sg](mohan@comp.nus.edu.sg)  \nABSTRACT  \nQuantifying the value of data within a machine learning workflow can play a pivotal role in making more strategic decisions in machine learning initiatives. The existing Shapley value based frameworks for data valuation in machine learning are computationally expensive as they require considerable amount of repeated training of the model to obtain the Shapley value. In this paper, we introduce an efficient data valuation framework EcoVal, to estimate the value of data for machine learning models in a fast and practical manner. Instead of directly working with individual data sample, we determine the value of a cluster of similar data points. This value is further propagated amongst all the member cluster points. We show that the overall value of the data can be determined by estimating the intrinsic and extrinsic value of each data. This is enabled by formulating the performance of a model as a production function, a concept which is popularly used to estimate the amount of output based on factors like labor and capital in a traditional free economic market. We provide a formal proof of our valuation technique and elucidate the principles and mechanisms that enable its accelerated performance. We demonstrate the real-world applicability of our method by showcasing its effectiveness for both in-distribution and out-of-sample data. This work addresses one of the core challenges of efficient data valuation at scale in machine learning models. The code is available at [https://github.com/respai-lab/ecoval](https://github.com/respai-lab/ecoval).  \n1 Introduction  \nData valuation is a pivotal concern in modern machine learning (ML) and data analytics, where the quality and worth of data have profound implications for decision-making, model performance, and data marketplace. Quantifying the worth of data plays an important role in data pricing and regulation compliance [1, 2], removing low-value/noisy data from the training set [3, 4], and incentivizing data sharing by personal data monetization [5, 6, 7, 8] . In a ML framework, the quality of data determines the effectiveness of the final model. Therefore, identifying high and low value data through data valuation would yield significant benefits for a wide range of machine learning applications.  \nBackground: In recent studies, a cooperative game theory concept, Shapley value [9] has been frequently used for data valuation in supervised ML [6, 5, 7] . It offers a desirable property of equitable reward allocation. The data Shapley and its extensions [6, 7, 10, 11] have empirically shown the effectiveness of Shapley value based valuation ina fixed dataset as well as in a particular distribution of data, allowing for out-of-time data valuation. The value of a data point in ML relies on its individual contribution to the model’s performance and its relationship with other data points utilized during training. The presence of similar data in the training set can dilute the significance of individual points. To account for these interactions, data Shapley methods evaluate the contribution of each point by determining how  \n*  \nThese authors contributed equally to this work  \n†Corresponding author  \nits absence affects the overall performanc","cbCaijO03ferPF1x","https://ap.wps.com/l/cbCaijO03ferPF1x","pdf",4177340,1,16,"English","en",105,"# Introduction\n## Data valuation in machine learning\n## Background: Shapley value for data valuation\n## Motivation: computational cost and scalability\n## Our contribution: cluster-level valuation with production functions","[{\"question\":\"Why are existing Shapley value-based data valuation methods expensive?\",\"answer\":\"They require many repeated training runs to estimate Shapley contributions, typically involving repeated model evaluations over many excluded subsets of data, leading to high computational cost.\"},{\"question\":\"How does EcoVal reduce the computation in data valuation?\",\"answer\":\"EcoVal performs valuation at the cluster level by grouping similar data points, drastically reducing the number of samples involved during repeated training. The cluster value is then propagated to member points.\"},{\"question\":\"How is overall data value computed in EcoVal?\",\"answer\":\"EcoVal determines overall value by estimating intrinsic and extrinsic value for each data cluster, enabled by modeling model performance with a production-function formulation.\"}]",1785674772,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"ecoval-an-efficient-data-valuation-framework-for-machine-learning","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/ecoval-an-efficient-data-valuation-framework-for-machine-learning/117260/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why are existing Shapley value-based data valuation methods expensive?","Question",{"text":74,"@type":75},"They require many repeated training runs to estimate Shapley contributions, typically involving repeated model evaluations over many excluded subsets of data, leading to high computational cost.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does EcoVal reduce the computation in data valuation?",{"text":79,"@type":75},"EcoVal performs valuation at the cluster level by grouping similar data points, drastically reducing the number of samples involved during repeated training. The cluster value is then propagated to member points.",{"name":81,"@type":72,"acceptedAnswer":82},"How is overall data value computed in EcoVal?",{"text":83,"@type":75},"EcoVal determines overall value by estimating intrinsic and extrinsic value for each data cluster, enabled by modeling model performance with a production-function formulation.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]