[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125198-en":3,"doc-seo-125198-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125198,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Active sampling - A machine-learning-assisted framework for finite population inference with optimal subsamples","Data subsampling helps overcome computational and economic bottlenecks when analyzing massive datasets. This work develops adaptive design methods for estimating finite population characteristics by combining active learning with adaptive importance sampling. An active sampling strategy is proposed that alternates between estimation and new data collection, selecting optimally sized subsamples. Machine learning predictions guide which unseen data are most informative. The approach is demonstrated on simulation-based traffic safety assessment for advanced driver assistance systems, showing substantial gains over traditional subsampling.","arXiv :2212 . 10024v3 [ stat .ME] 4 Jul 2024  \nActive sampling: A  \nmachine-learning-assisted framework for finite population inference with optimal  \nsubsamples  \nHenrik Imberg 1 , Xiaomi Yang2 , Carol Flannagan2 ,3 , Jonas B¨argman2  \n1 Department of Mathematical Sciences,  \nChalmers University of Technology and University of Gothenburg  \n2 Division of Vehicle Safety, Chalmers University of Technology  \n3 University of Michigan Transportation Research Institute  \nAbstract  \nData subsampling has become widely recognized as a tool to overcome computational and economic bottlenecks in analyzing massive datasets. We contribute to the development of adaptive design for estimation of finite population characteristics, using active learning and adaptive importance sampling. We propose an active sampling strategy that iterates between estimation and data collection with optimal subsamples, guided by machine learning predictions on yet unseen data. The method is illustrated on virtual simulation-based safety assessment of advanced driver assistance systems. Substantial performance improvements are demonstrated compared to traditional sampling methods.  \nKeywords: active learning, adaptive importance sampling, computer simulation experiments, inverse probability weighting, optimal design, traffic safety assessment  \n1 Introduction  \nWe consider a deterministic computer simulation experiment which for a given input z returns a fixed output y. The input space is assumed to be discrete and the simulation experiment hence characterized by the set of complete input-output pairs { (zi , yi)}, where N is the size of the experiment. The aim our experiment it to calculate a characteristic  \nN  \nθ = h(ty) , ty =X yi , (1)  \ni=1  \nfor some differentiable function h : Rd → R and d-dimensional vector of totals ty . Examples of such a characteristic include, e.g. , a total, mean, ratio, or correlation coefficient. This is also known as a finite population inference problem (Beaumont and Haziza, 2022) . We further assume that N is large, as is the computational cost of each single experiment, rendering complete enumeration unfeasible. In such circumstances, researches often resort to subsampling.  \nSubsampling methods have seen a huge increase in popularity over the past few years across many different areas of statistics. For instance, Ma et al. (2015 , 2022) introduced leverage sampling for big data regression, which subsequently inspired similar developments for logistic regression (Wang et al. , 2018; Yao and Wang, 2019) generalized linear models (Ai et al. , 2021b; Yu et al. , 2022), and quantile regression (Ai et al. , 2021a; Wang et al. , 2021) . Similarly, Dai et al. (2022) developed an optimal subsampling method for regression using a minimum energy criterion. Sometimes subsampling is induced by economical rather than computational constraints. In this setting, Imberg et al. (2022) developed an optimal subsampling method for two-phase sampling experiments. A similar measurement-  \nconstrained experiment problem was addressed by Zhang et al. (2021) using a sequential subsampling procedure and by Meng et al. (2021) using a space-filling Latin hypercube sampling method.  \nFor computer simulation experiments, subsampling methods using adaptive design for Gaussian process response surface modeling are commonly employed. Together with active learning and Bayesian optimization, this provides a powerful framework for computer experiment emulation (Gramacy and Apley, 2015; Sun et al. , 2017; Lei et al. , 2021; Limet al. , 2021) . Another popular approach is model-free space-filling methods using, e.g. , Latin hypercube sampling designs (see, e.g. , Cioppa and Lucas, 2007; Zhang et al. , 2024; Zhou et al. , 2024) . Others have utilized methods based on optimal transport, e.g., for kernel density estimation (Zhang et al. , 2023) . For estimating a simple statistic, such as a mean or ratio, however, importance sampling and adaptive importance sampling ","cbCaitBl3qvXFuBj","https://ap.wps.com/l/cbCaitBl3qvXFuBj","pdf",10665664,1,68,"English","en",105,"# Abstract\n# Introduction\n## Finite population inference and computational bottlenecks\n## Subsampling and optimal subsampling methods\n## Adaptive design, active learning, and importance sampling","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It targets finite population inference when estimating characteristics of a large deterministic simulation experiment is computationally expensive, making full enumeration infeasible.\"},{\"question\":\"How does the proposed method choose data to sample?\",\"answer\":\"It uses an iterative active sampling strategy that alternates between estimation and data collection, selecting optimal subsamples guided by machine learning predictions.\"},{\"question\":\"Where is the method demonstrated and how does it perform?\",\"answer\":\"It is illustrated in simulation-based safety assessment of advanced driver assistance systems, where results show substantial performance improvements compared with traditional sampling methods.\"}]","Active sampling - A machine-learning-assisted framework for finite population inference with optimal subsamples | PDF",1785897336,171,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"active-sampling-a-machine-learning-assisted-framework-for-finite-population-inference-with-optimal-subsamples","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/active-sampling-a-machine-learning-assisted-framework-for-finite-population-inference-with-optimal-subsamples/125198/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"It targets finite population inference when estimating characteristics of a large deterministic simulation experiment is computationally expensive, making full enumeration infeasible.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method choose data to sample?",{"text":80,"@type":76},"It uses an iterative active sampling strategy that alternates between estimation and data collection, selecting optimal subsamples guided by machine learning predictions.",{"name":82,"@type":73,"acceptedAnswer":83},"Where is the method demonstrated and how does it perform?",{"text":84,"@type":76},"It is illustrated in simulation-based safety assessment of advanced driver assistance systems, where results show substantial performance improvements compared with traditional sampling methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]