[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85848-en":3,"doc-seo-85848-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85848,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","How Much Data Do We Need Sequential Data Collection for Stochastic Programming","Data-driven optimization depends on estimating uncertain model parameters from collected data, yet acquiring additional data can be costly and constrained in practice. This work studies an optimal stopping problem for sequential data collection in stochastic optimization under parameter uncertainty. A Bayesian learning model updates beliefs as new observations arrive. At each step, the decision maker compares expected marginal information benefit to unit sampling cost, then chooses to continue sampling or stop and solve. Stopping policies are tested on a newsvendor case with experiments against fixed-budget and hindsight baselines.","arXiv :2607 . 10207v1 [math .OC] 11 Jul 2026  \nHow much Data do We Need? Sequential Data Collection for Stochastic Programming  \nXin Li, Juergen Branke, Xuan Vinh Doan University of Warwick  \n{[xin. li.2](xin. li.2) , juergen. branke, [xuan. doan](xuan. doan}@wbs. ac. uk)[}](xuan. doan}@wbs. ac. uk)[@wbs. ac. uk](xuan. doan}@wbs. ac. uk)  \nAbstract  \nData-driven optimization often requires collecting data to estimate uncertain model parameters before solving the underlying decision problem. In practice, however, data acquisition may incur non-negligible costs, making it critical to determine when to stop additional data collection. In this paper, we study an optimal stopping problem for sequential data collection in stochastic optimization under parameter uncertainty. We propose a benefit-driven stopping framework that balances information gain and sampling cost. We model the unknown distribution parameter within a Bayesian learning framework and update beliefs sequentially as new observations are collected. At each iteration, the decision maker evaluates the expected marginal benefit of additional data relative to the unit sampling cost and determines whether to continue sampling or stop and implement the optimization decision. Based on this framework, we develop several stopping policies. The proposed policies are evaluated through a newsvendor problem with exponentially distributed demand. Numerical experiments compare the policies with fixed-budget and hindsight benchmark strategies. The results show that benefit-driven stopping rules can substantially reduce unnecessary data collection while achieving near-optimal decision performance, demonstrating the effectiveness of adaptive stopping in data-driven optimization.  \n1 Introduction  \nMathematical programming models typically rely on input parameters that are either specified by domain experts or estimated from historical data. In practice, these parameters are rarely known with certainty. When the estimated parameters deviate from their true values, the resulting optimization decisions may be sub-optimal (Lam, 2016) . This challenge is commonly referred to as optimization under input uncertainty (Birge and Louveaux, 2011; Shapiro et al., 2021) and has received increasing attention in stochastic optimization and data-driven decision making.  \nA natural approach to mitigating input uncertainty is to collect additional data. More data improve parameter estimates and, consequently, decision quality. However, real-world data collection is often costly and subject to practical constraints (Ungredda et al., 2022) . In applications such as healthcare, supply chain management, manufacturing, and service operations, collecting additional data can be costly due to testing expenses, operational disruptions, limited access to data, or time constraints (Xu et al., 2023; Fu and Zhu, 2010) . This creates a fundamental trade-off between the value of information obtained from new data and the cost of collecting it. In this paper, we are interested in finding the trade-off between the benefit and the collection cost of new data.  \nMost existing studies in data-driven optimization assume that data are given in a single batch and remain fixed throughout the decision-making process (Song and Shanbhag, 2019; He and Song, 2024) . In contrast, many real-world applications allow data to be collected sequentially over time. For example, demand information may be obtained incrementally through repeated market studies or ongoing observations (Besbes and Muharremoglu, 2013) . As more data are collected, parameter estimates become more accurate, and the resulting decisions improve. With infinite data, the estimates converge to the true parameters, and thus the corresponding solution converges to the true optimal solution (Kleywegt et al., 2002; Kim et al., 2014; He et al., 2024) . Nevertheless, data collection incurs cost, the decision maker (DM) must balance the benefit of reduced input uncertainty a","cbCaicWbrfoWm49y","https://ap.wps.com/l/cbCaicWbrfoWm49y","pdf",1715790,4,1,40,"English","en",105,"# Introduction\n## Optimization under input uncertainty\n## Data collection trade-off\n## Sequential data vs batch data\n## Optimal stopping formulation\n## Regret, information benefit, and sampling cost","[{\"question\":\"Why is stopping additional data collection an important decision in data-driven optimization?\",\"answer\":\"Because more data can improve parameter estimates and decision quality, but acquiring data may incur non-negligible costs and practical constraints. The key is balancing information gain against sampling cost.\"},{\"question\":\"How does the proposed framework decide whether to continue sampling or stop?\",\"answer\":\"It uses Bayesian learning to update beliefs sequentially, then evaluates the expected marginal benefit of additional data relative to a unit sampling cost at each iteration. The decision maker continues sampling or stops accordingly.\"},{\"question\":\"How are the stopping policies evaluated in the paper?\",\"answer\":\"The policies are assessed through a newsvendor problem with exponentially distributed demand. Numerical experiments compare benefit-driven stopping rules with fixed-budget and hindsight benchmark strategies, showing reduced unnecessary data collection while maintaining near-optimal performance.\"}]",1784206684,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"how-much-data-do-we-need-sequential-data-collection-for-stochastic-programming","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/how-much-data-do-we-need-sequential-data-collection-for-stochastic-programming/85848/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is stopping additional data collection an important decision in data-driven optimization?","Question",{"text":75,"@type":76},"Because more data can improve parameter estimates and decision quality, but acquiring data may incur non-negligible costs and practical constraints. The key is balancing information gain against sampling cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework decide whether to continue sampling or stop?",{"text":80,"@type":76},"It uses Bayesian learning to update beliefs sequentially, then evaluates the expected marginal benefit of additional data relative to a unit sampling cost at each iteration. The decision maker continues sampling or stops accordingly.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the stopping policies evaluated in the paper?",{"text":84,"@type":76},"The policies are assessed through a newsvendor problem with exponentially distributed demand. Numerical experiments compare benefit-driven stopping rules with fixed-budget and hindsight benchmark strategies, showing reduced unnecessary data collection while maintaining near-optimal performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]