[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124634-en":3,"doc-seo-124634-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124634,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","The Effect of Dataset Size and the Process of Big Data Mining for Investigating Solar-Thermal Desalination by Machine Learning","An effective interdisciplinary study between machine learning and solar-thermal desalination depends on sufficiently large, well-analyzed experimental datasets. A modified dataset collection and analysis process is proposed, using an optimized water condensation and collection method to accelerate acquisition and reduce collection time by 83.3%, yielding over one thousand datasets. Feature effects are then evaluated with ANN, multiple linear regressions, and random forests, focusing on dataset size and range impacts on prediction accuracy, factor-importance ranking, and generalization. Larger datasets improve accuracy, strongly influence importance rankings, and affect extrapolation accuracy for ANN.","Optimized data collection and analysis process for studying solar-thermal desalination by machine learning  \nGuilong Peng\\#1,2, Senshan Sun\\#2, Yangjun Qin2, Zhenwei Xu2, Juxin Du2, Swellam W.  \nsharshir2,3, A.W. Kandel2,3, A.E. Kabeel2,4,5, Nuo Yang*2  \n1School of Mechanical and Energy Engineering, Shaoyang University, Shaoyang 422000, China  \n2State Key Laboratory of Coal Combustion, Huazhong University of Science and Technology, Wuhan 430074, China  \n3Mechanical Engineering Department, Faculty of Engineering, Kafrelsheikh University, Kafrelsheikh 33516, Egypt  \n4Mechanical Power Engineering Department, Faculty of Engineering, Tanta University, Tanta, Egypt  \n5Faculty of Engineering, Delta University for Science and Technology, Gamasa, Egypt  \n\\#Guilong Peng and Senshan Sun contribute equally to this work  \n*Corresponding email: Nuo Yang ([nuo@hust.edu.cn](nuo@hust.edu.cn))  \nAbstract  \nAn effective interdisciplinary study between machine learning and solar-thermal desalination requires a sufficiently large and well-analyzed experimental datasets. This study develops a modified dataset collection and analysis process for studying solar-thermal desalination by machine learning. Based on the optimized water condensation and collection process, the proposed experimental method collects over one thousand datasets, which is ten times more than the average number of datasets in previous works, by accelerating data collection and reducing the time by 83.3% . On the other hand, the effects of dataset features are investigated by using three different algorithms, including artificial neural networks, multiple linear regressions, and random forests. The investigation focuses on the effects of dataset size and range on prediction accuracy, factor importance ranking, and the model's generalization ability. The results demonstrate that a larger dataset can significantly improve prediction accuracy when using artificial neural networks and random forests. Additionally, the study highlights the significant impact of dataset size and range on ranking the importance of influence factors. Furthermore, the study reveals that the extrapolation data range significantly affects the extrapolation accuracy of artificial neural networks. Based on the results, massive dataset collection and analysis of dataset feature effects are important steps in an effective and consistent machine learning process flow for solar-thermal desalination, which can promote machine learning as a more general tool in the field of solar-thermal desalination.  \nKeywords: Solar desalination; Machine learning; Dataset collection; Production prediction; Artificial neural network.  \nNomenclature  \n෡  \nB  \nc1  \nc2  \nD  \n푒  \n퐸  \n퐸푚푖푛  \nf푖  \n퐻퐹  \n퐺퐼  \nℎ  \ni  \nj  \nk  \nl  \nm  \n푀푚̇   \n푚푡  \nn  \nN  \n푁푀  \n표푀  \n푒  \n푃퐹  \n푞  \n푅 푗푅푀 (푗)  \n푅푙  \n푅2  \n∆푡  \n푇푎푚푏  \nList of regression coefficients Average values of productivity in 푅1 푗 in RF  \nAverage values of productivity in 푅2 푗 in RF  \nNormalized dataset  \nLabel of node  \nOutput error signal in BP-ANN Threshold of root mean square error in BP-ANN  \nPredicted value in ML Fan height above the basin  \nGini impurity  \nHyperparameters of BP-ANN Label of neurons in current layer in BP-ANN  \nLabel of sample in RF  \nNumber of independent variables Label of neurons in previous layer in BP-ANN  \nNumber of neurons in current layer in BP-ANN  \nlabel of regions in RF Productivity  \nTotal mass of the collected freshwater increases with time Number of DTs in RF  \nSample size  \nNumber of elements in region 푀 in RF  \nAverage output value in DT Estimated probability that sample belongs to any class at node 푒 in RF  \nPower of the fan  \nDimension label  \nRegions sliced by 푗th sample Region of label 푀 in RF Random number  \nCoefficient of determination Given period  \nAmbient temperature  \n푇푔  \nTss  \n푇푤  \nVIM퐷푖푇  \n푤  \n푥  \nX  \n푦퐿푅  \n푦  \n푛푒푢푦  \n푁푁푦  \n푅퐹푦  \n푦푖  \nY  \nGlass cover temperature Solar still types  \nWater temperature  \nImportance of one variable at node 푒 i","cbCaiaKefnwqMbdb","https://ap.wps.com/l/cbCaiaKefnwqMbdb","pdf",4100210,1,46,"English","en",105,"# Introduction\n## Safe drinking water and desalination needs\n## Solar-thermal desalination and its advantages\n## Machine learning methods for scientific data analysis\n## Applications of machine learning in solar energy fields","[{\"question\":\"Why is a large and well-analyzed dataset important for applying machine learning to solar-thermal desalination?\",\"answer\":\"Sufficient dataset size and quality are necessary to enable effective interdisciplinary modeling and reliable prediction from experimental conditions.\"},{\"question\":\"How does the proposed method improve dataset collection efficiency?\",\"answer\":\"It uses an optimized water condensation and collection process that accelerates data acquisition and reduces the collection time by 83.3%, producing over one thousand datasets.\"},{\"question\":\"What does the study find about the effects of dataset size and range on model performance?\",\"answer\":\"A larger dataset significantly improves prediction accuracy for ANN and random forests, strongly impacts factor-importance ranking, and the extrapolation data range notably affects ANN extrapolation accuracy.\"}]","The Effect of Dataset Size and the Process of Big Data Mining for Investigating Solar-Thermal Desalination by Machine Learning | PDF",1785893428,116,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-effect-of-dataset-size-and-the-process-of-big-data-mining-for-investigating-solar-thermal-desalination-by-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-effect-of-dataset-size-and-the-process-of-big-data-mining-for-investigating-solar-thermal-desalination-by-machine-learning/124634/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is a large and well-analyzed dataset important for applying machine learning to solar-thermal desalination?","Question",{"text":75,"@type":76},"Sufficient dataset size and quality are necessary to enable effective interdisciplinary modeling and reliable prediction from experimental conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method improve dataset collection efficiency?",{"text":80,"@type":76},"It uses an optimized water condensation and collection process that accelerates data acquisition and reduces the collection time by 83.3%, producing over one thousand datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study find about the effects of dataset size and range on model performance?",{"text":84,"@type":76},"A larger dataset significantly improves prediction accuracy for ANN and random forests, strongly impacts factor-importance ranking, and the extrapolation data range notably affects ANN extrapolation accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]