[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122013-en":3,"doc-seo-122013-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122013,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Effect of the Sampling of a Dataset in the Hyperparameter Optimization Phase over the Efficiency of a Machine Learning Algorithm","The research examines how partitioning a dataset influences model efficiency during the hyperparameter optimization (HPO) stage. Hyperparameter selection directly affects predictive performance but is costly due to many configurations and substantial computational resources. By applying nonparametric inference, the study quantifies accuracy, time cost, and spatial complexity across different partitions and the full dataset. It assigns a gain level to each partition to identify sampling strategies that provide more profitable subsets. Analyses use five cybersecurity datasets, emphasizing efficiency for actionable AI.","Hindawi  \nComplexity  \nVolume 2019, Article ID 6278908, 16 pages [https://doi.org/10.1155/2019/6278908](https://doi.org/10.1155/2019/6278908)  \nResearch Article  \nEffect of the Sampling of a Dataset in  \nthe Hyperparameter Optimization Phase over the Efficiency of a Machine Learning Algorithm  \nNoem-DeCastro-Garc-a,1 Ángel Luis Muñoz Castañeda,2  \nDavid Escudero Garc-a,2 and Miguel V. Carriegos1  \n1 Departamento de Matem ticas, Universidad de Le n, Campus de Vegazana s/n, 24071 Le n, Spain  \n2 Research Institute on Applied Sciences in Cybersecurity, Universidad de Le  n, Campus de Vegazana s/n, 24071 Le n, Spain Correspondence should be addressed to Noem´ı DeCastro-Garca; [ncasg@unileon.es](ncasg@unileon.es)  \nReceived 7 December 2018; Accepted 17 January 2019; Published 4 February 2019  \nGuest Editor: Fernando Snchez Lasheras  \nCopyright © 2019 NoemDeCastro-Garc´ıaetal. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.  \nSelecting the best configuration of hyperparameter values for a Machine Learning model yields directly in the performance of the model on the dataset. It is a laborious task that usually requires deep knowledge of the hyperparameter optimizations methods and the Machine Learning algorithms. Although there exist several automatic optimization techniques, these usually take significant resources, increasing the dynamic complexity in order to obtain a great accuracy. Since one of the most critical aspects in this computational consume is the available dataset, among others, in this paper we perform a study of the effect of using different partitions of a dataset in the hyperparameter optimization phase over the efficiency of a Machine Learning algorithm. Nonparametric inference has been used to measure the rate of different behaviorsofthe accuracy, time, and spatial complexity that are obtained among the partitions and the whole dataset. Also, a level of gain is assigned to each partition allowing us to study patterns and allocate whose samples are more profitable. Since Cybersecurity is a discipline in which the efficiency of Artificial Intelligence techniques is a key aspect in order to extract actionable knowledge, the statistical analyses have been carried out over five Cybersecurity datasets.  \n1. Introduction  \nA Machine Learning (ML) solution for a classification problem is effective if it works efficiently in terms of accuracy and the required computational cost. The improvement ofthe first factor is faced on by several points of view that could affect to the second one in different forms.  \nThe simplest way to get a ML model with a good accuracy is by testing and comparing different ML algorithms for the same problem and choosing, finally, the one that performs better. However, it is clear that, for instance, a decision tree model does not require, in general, as much computational time and memory tobe trained asa Multilayer Perceptron. So, we will need to adjust the achieved accuracy with the available resources.  \nAnother usual effective approach to reacha high accuracy is working with large training datasets. Nevertheless, this  \nsolution is limited because of the associated computational cost (obtaining and storing the data, cleaning and transformation processes, and learning from the data) . A possible alternative to the mentioned problem is to reduce the training without losing too much information [1, 2] . However, these kinds of solutions used to need an expensive data preprocessing phase.  \nThe research related to this aspect, in addition to usual filtering the data, is focused on how to optimize the training set, and not only reduce it. The progressive sampling method shows that the performance with random samples with determined sizes is equal or more effective than working with the entire dataset [3] . Also, this solution used to","cbCaidxWNJZR0fCW","https://ap.wps.com/l/cbCaidxWNJZR0fCW","pdf",1653538,1,16,"English","en",105,"# Introduction\n# Complexity","[{\"question\":\"Why does dataset sampling matter in the hyperparameter optimization phase?\",\"answer\":\"Sampling determines how training data are partitioned during HPO, which impacts accuracy and computational demands. Different partitions can lead to measurable differences in efficiency and complexity.\"},{\"question\":\"How are the effects across dataset partitions measured?\",\"answer\":\"Nonparametric inference is used to compare accuracy behavior, time cost, and spatial complexity between partitions and the whole dataset.\"},{\"question\":\"What is the practical purpose of assigning a gain level to each partition?\",\"answer\":\"The gain level helps identify which subsets are more profitable, revealing patterns that guide choosing sampling strategies during HPO.\"}]","Effect of the Sampling of a Dataset in the Hyperparameter Optimization Phase over the Efficiency of a Machine Learning Algorithm | PDF",1785808289,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"effect-of-the-sampling-of-a-dataset-in-the-hyperparameter-optimization-phase-over-the-efficiency-of-a-machine-learning-algorithm","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/effect-of-the-sampling-of-a-dataset-in-the-hyperparameter-optimization-phase-over-the-efficiency-of-a-machine-learning-algorithm/122013/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does dataset sampling matter in the hyperparameter optimization phase?","Question",{"text":75,"@type":76},"Sampling determines how training data are partitioned during HPO, which impacts accuracy and computational demands. Different partitions can lead to measurable differences in efficiency and complexity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the effects across dataset partitions measured?",{"text":80,"@type":76},"Nonparametric inference is used to compare accuracy behavior, time cost, and spatial complexity between partitions and the whole dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the practical purpose of assigning a gain level to each partition?",{"text":84,"@type":76},"The gain level helps identify which subsets are more profitable, revealing patterns that guide choosing sampling strategies during HPO.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]