[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124984-en":3,"doc-seo-124984-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124984,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Farthest Point Sampling in Property Designated Chemical Feature Space - as a General Strategy for Enhancing Machine Learning Model Performance for Small Scale Chemical Dataset","Machine learning model development in chemistry and materials science often faces small, unbalanced labeled datasets that reduce generalization and increase overfitting risk. This study evaluates farthest point sampling (FPS) in property-designated chemical feature spaces to construct well-distributed training sets. Across artificial neural networks, support vector machines, and random forests using physicochemical properties (e.g., boiling points and vaporization enthalpy), FPS-based models deliver higher predictive accuracy and robustness than random sampling, with stronger gains for smaller datasets due to increased feature-space diversity.","Farthest Point Sampling in Property Designated Chemical Feature Space as a General Strategy for Enhancing the Machine Learning Model Performance for Small Scale Chemical Dataset  \nYuze Liua, b, Xi Yua, b*  \na Key Laboratory of Organic Integrated Circuit, Ministry of Education & Tianjin Key Laboratory of Molecular Optoelectronic Sciences, Department of Chemistry, School of Science, Tianjin University, Tianjin 300072, China  \nb Collaborative Innovation Center of Chemical Science and Engineering (Tianjin), Tianjin 300072, China  \n* [Email: xi.yu@tju.edu.cn](Email: xi.yu@tju.edu.cn)  \nMachine Learning, Cheminformatics, Farthest Point Sampling, Small Data Set, Chemical Database  \nABSTRACT: Machine learning model development in chemistry and materials science often grapples with the challenge of smallscale, unbalanced labelled datasets, a common limitation in scientific experiments. These dataset imbalances can precipitate overfitting and diminish model generalization. Our study explores the efficacy of the farthest point sampling (FPS) strategy within targeted chemical feature spaces, demonstrating its capacity to generate well-distributed training datasets and consequently enhance model performance. We rigorously evaluated this strategy across various machine learning models, including artificial neural networks (ANN), support vector machines (SVM), and random forests (RF), using datasets encapsulating physicochemical properties like standard boiling points and enthalpy of vaporization. Our findings reveal that FPS-based models consistently surpass those trained via random sampling, exhibiting superior predictive accuracy and robustness, alongside a marked reduction in overfitting. This improvement is particularly pronounced in smaller training datasets, attributable to increased diversity within the training data's chemical feature space. Consequently, FPS emerges as a universally effective and adaptable approach in approaching high performance machine learning models by small and biased experimental datasets prevalent in chemistry and materials science.   \n1. INTRODUCTION  \nMachine learning (ML) has significantly advanced the fields of chemistry and material science 1-4, propelling the study of cheminformatics and enabling rapid structureproperty prediction and design5-11. However, the inherent requirement on the extensive dataset for ML study raised significant challenge for practical application of ML in experimental science12. Often, labelled experimental chemical and material datasets are limited in size and coverage, and most significantly imbalanced 13, 14, due to constraints in data acquisition, like time, cost, and technical barriers. Consequently, ML models trained on these datasets, which are frequently subsampled randomly for training and testing, are prone to overfitting and exhibit diminished generalization capabilities due to the imbalanced nature of the data, where certain types of observations are disproportionately represented compared to others. The complexity of these challenges is further magnified by the high dimensionality of chemical data and the intricate nature of chemical scenarios.  \nTo mitigate these issues, various sampling methods have been employed to achieve data balance and curtail the risk of overfitting 15, 16. Conventional methods like oversampling and under-sampling17, 18 directly manipulate the dataset's size to address class imbalances but can lead to loss of information or overfitting, while stratified sampling 19 maintains the proportion of classes but doesn't necessarily enhance data diversity. Advanced methods bring nuanced solutions with their own  \ntrade-offs. The Genetic Algorithm (GA) method20, while optimizing data point selection for diversity, can be computationally intensive and may require careful tuning to avoid converging on suboptimal solutions. Active learning21-24 effectively refines models iteratively by selecting informative data, yet this comes with the cost of increase","cbCait0On5U5IPRA","https://ap.wps.com/l/cbCait0On5U5IPRA","pdf",1343210,1,9,"English","en",105,"# Introduction\n## Dataset challenges in cheminformatics\n## Sampling methods and limitations\n## Proposed strategy: FPS in property-designated chemical feature space\n# Evaluation approach and models\n## Compared learning models","[{\"question\":\"Why do small, imbalanced chemical datasets hurt machine learning in chemistry and materials science?\",\"answer\":\"They commonly lead models toward overfitting and weaker generalization because random subsampling cannot correct the imbalance and data coverage limitations.\"},{\"question\":\"What is the core idea of farthest point sampling (FPS) in this work?\",\"answer\":\"FPS selects training samples that are furthest apart within a targeted chemical feature space, aiming to represent the dataset with fewer points while preserving diversity.\"},{\"question\":\"How does FPS-based training compare with random sampling in the reported experiments?\",\"answer\":\"Models trained using FPS consistently outperform random-sampling baselines in predictive accuracy and robustness, and they show a marked reduction in overfitting, especially when the training set is small.\"}]","Farthest Point Sampling in Property Designated Chemical Feature Space - as a General Strategy for Enhancing Machine Learning Model Performance for Small Scale Chemical Dataset | PDF",1785895842,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"farthest-point-sampling-in-property-designated-chemical-feature-space-as-a-general-strategy-for-enhancing-machine-learning-model-performance-for-small-scale-chemical-dataset","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/farthest-point-sampling-in-property-designated-chemical-feature-space-as-a-general-strategy-for-enhancing-machine-learning-model-performance-for-small-scale-chemical-dataset/124984/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do small, imbalanced chemical datasets hurt machine learning in chemistry and materials science?","Question",{"text":75,"@type":76},"They commonly lead models toward overfitting and weaker generalization because random subsampling cannot correct the imbalance and data coverage limitations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea of farthest point sampling (FPS) in this work?",{"text":80,"@type":76},"FPS selects training samples that are furthest apart within a targeted chemical feature space, aiming to represent the dataset with fewer points while preserving diversity.",{"name":82,"@type":73,"acceptedAnswer":83},"How does FPS-based training compare with random sampling in the reported experiments?",{"text":84,"@type":76},"Models trained using FPS consistently outperform random-sampling baselines in predictive accuracy and robustness, and they show a marked reduction in overfitting, especially when the training set is small.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]