[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120687-en":3,"doc-seo-120687-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120687,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","A Novel Algorithm Can Generate Data to Train Machine Learning Models in Conditions of Extreme Scarcity of Real World Data","Training effective machine learning models depends on access to large datasets, yet real-world data collection and management are constrained by cost, ethical and legal risks, and limited availability. The presented approach introduces a genetic-algorithm-based method that evolves candidate synthetic datasets by repeatedly training a neural network and using its performance on real data as a fitness surrogate. Selection pressure discards unfit datasets and improves fitness across generations. Experiments on Iris and Breast Cancer datasets show comparable accuracy under abundance and significant gains under simulated extreme scarcity.","A novel algorithm can generate data to train machine learning models  \nin conditions of extreme scarcity of real world data  \nOlivier Niel, MD, PhD (1)(2) *  \n(1) Kannerklinik, 4 rue Barblé, L1210 Luxembourg, Luxembourg  \n(2) Faculty of Science, Technology and Medicine, University of Luxembourg, 2 Avenue de l'Université, L4365, Esch-sur-Alzette, Luxembourg  \n* Correspondence: o.r.p.niel@free.fr (preferred), [olivier.niel@ext.uni.lu](olivier.niel@ext.uni.lu)  \nKeywords: data generation, artificial intelligence, machine learning, neural network, genetic algorithm, data augmentation, synthetic data generation, Iris dataset, Breast cancer dataset  \nAbstract  \nBackground: Training machine learning models requires large datasets. However, collecting, curating, and operating large and complex sets of real world data poses several problems, in terms of costs, ethical and legal issues, and data availability. To overcome these difficulties, we propose a novel algorithm which can generate large artificial datasets to train machine learning models, even in conditions of extreme scarcity of real world data.  \nMethods: The data generation algorithm is based on a genetic algorithm, which mutates randomly generated datasets subsequently used for training a neural network. After training, the performance of the neural network on a batch of real world data is considered a surrogate for the fitness of the generated dataset used for its training. As selection pressure is applied to the population of generated datasets, unfit generated datasets are discarded, and the fitness of the fittest generated datasets increases throughout generations.  \nResults: The performance of the data generation algorithm was measured on the Iris dataset and on the Breast Cancer Wisconsin diagnostic dataset. In conditions of real world data abundance, mean accuracy of machine learning models trained on generated data was comparable to mean accuracy of machine learning models trained on real world data (0.956 in both cases on the Iris dataset, p = 0.6996, and 0.9377 versus 0.9472 on the Breast Cancer dataset, p = 0.1189) . In conditions of simulated extreme scarcity of real world data, mean accuracy of machine learning models trained on generated data was significantly higher than mean accuracy of comparable machine learning models trained on scarce real world data (0.9533 versus 0.9067 on the Iris dataset, p \u003C 0.0001, and 0.8692 versus 0.7701 on the Breast Cancer dataset, p = 0.0091) .  \nConclusion: Here we propose a novel algorithm, which can generate large artificial datasets to train machine learning models, in conditions of extreme scarcity of real world data, and also when cost or data sensitivity becomes an obstacle to the collection of large real world datasets.  \n1. Introduction  \nArtificial intelligence algorithms are now being used extensively in most scientific disciplines. In particular, machine learning models have proven remarkably effective at solving complex classification or regression problems [1]. It is noteworthy that machine learning models need intensive training before they can be used; effective training requires significant computational resources, and, for intrinsic reasons, extremely large datasets [1] .  \nHowever, collecting, curating, and operating large and complex sets of real world data poses several problems. First, these operations have substantial material and human costs. As an example, in 2021, global spending on big data and analytics reached 241 billion US dollars, and should peak above 655 billion US dollars by 2029 [2] . Second, data collection often raises ethical and legal issues, in terms of privacy and data protection. Indeed, when the nature of data becomes sensitive, which is the case with most personal data like medical health records, complex and costly arrangements are required to comply with local regulations and legislations, such as the General Data Protection Regulation in European countries, and to ensure that data","cbCaimB6flLBIw6n","https://ap.wps.com/l/cbCaimB6flLBIw6n","pdf",1198191,1,24,"English","en",105,"# Abstract\n## Background\n## Methods\n## Results\n## Conclusion\n# Introduction\n# Material and methods\n## Implementation","[{\"question\":\"Why is data scarcity a problem for training machine learning models?\",\"answer\":\"Machine learning models require intensive training and large datasets. Collecting, curating, and operating real-world data can be expensive, raises privacy and legal issues, and may be impossible to obtain in scarce settings such as rare disease cohorts.\"},{\"question\":\"How does the proposed algorithm generate synthetic data?\",\"answer\":\"A genetic algorithm applies random mutations to candidate generated datasets. A neural network is trained using each dataset, and its performance on a batch of real-world data is used as a surrogate fitness measure for selecting fitter datasets.\"},{\"question\":\"What results does the method achieve on the Iris and Breast Cancer datasets?\",\"answer\":\"With real-world data abundance, models trained on generated data reach mean accuracies comparable to those trained on real data. Under simulated extreme scarcity, models trained on generated data show significantly higher mean accuracy than those trained on scarce real data.\"}]","A Novel Algorithm Can Generate Data to Train Machine Learning Models in Conditions of Extreme Scarcity of Real World Data | PDF",1785731500,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-novel-algorithm-can-generate-data-to-train-machine-learning-models-in-conditions-of-extreme-scarcity-of-real-world-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-novel-algorithm-can-generate-data-to-train-machine-learning-models-in-conditions-of-extreme-scarcity-of-real-world-data/120687/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is data scarcity a problem for training machine learning models?","Question",{"text":75,"@type":76},"Machine learning models require intensive training and large datasets. Collecting, curating, and operating real-world data can be expensive, raises privacy and legal issues, and may be impossible to obtain in scarce settings such as rare disease cohorts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed algorithm generate synthetic data?",{"text":80,"@type":76},"A genetic algorithm applies random mutations to candidate generated datasets. A neural network is trained using each dataset, and its performance on a batch of real-world data is used as a surrogate fitness measure for selecting fitter datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does the method achieve on the Iris and Breast Cancer datasets?",{"text":84,"@type":76},"With real-world data abundance, models trained on generated data reach mean accuracies comparable to those trained on real data. Under simulated extreme scarcity, models trained on generated data show significantly higher mean accuracy than those trained on scarce real data.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]