[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123863-en":3,"doc-seo-123863-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123863,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Predicting the Activity of Chemical Compounds Based on Machine Learning Approaches","Machine learning methods are applied to cheminformatics to improve prediction of chemical compound activity. Experiments evaluate 100 combinations of existing ML techniques, and the best solutions are selected using G-mean, F1-score, and AUC metrics. Model validation is performed on a PubChem-derived dataset containing about 10,000 compounds labeled by activity status, enabling assessment under real-world class imbalance conditions and support for reliable ranking of candidate molecules.","PREDICTING THE ACTIVITY OF CHEMICAL COMPOUNDS BASED ON  \nMACHINE LEARNING APPROACHES  \nDo Hoang Tu1, Tran Van Lang2*, Pham Cong Xuyen1, Le Mau Long1,3,  \n1 Lac Hong University, Vietnam  \n2 HCMC University of Foreign Languages-Information Technology, Vietnam  \n3 Nguyen Tat Thanh University, Vietnam  \n[dhtus.vn@gmail.com](dhtus.vn@gmail.com), [langtv@huflit.edu.vn](langtv@huflit.edu.vn), [pcxuyen@lhu.edu.vn](pcxuyen@lhu.edu.vn), [lmaulong@gmail.com](lmaulong@gmail.com)  \nABSTRACT—Exploring methods and techniques of machine learning (ML) to address specific challenges in various fields is essential. In this work, we tackle a problem in the domain of Cheminformatics; that is, providing a suitable solution to aid in predicting the activity of a chemical compound to the best extent possible. To address the problem at hand, this study conducts experiments on 100 different combinations of existing techniques. These solutions are then selected based on a set of criteria that includes the Gmeans, F1-score, andAUC metrics. The results have been tested on a dataset of about 10,000 chemical compounds from PubChem that have been classified according to their activity.  \nKEYWORDS— Cheminformatics, Data imbalance, Lossfunction, GAN model, Ensemble learning  \nI. INTRODUCTION  \nIn datasets used in biological experiments for measuring the activity of various compounds against different biological targets, often used in screening, there is usually a significant imbalance between active and inactive compounds, with the number of inactive data points being much larger. Therefore, training requires the use of suitable machine learning models. Additionally, preprocessing before using machine learning methods for training is also a crucial issue. The following issues are approached to address the problem of predicting the activity of chemical compounds using chemistry-related datasets:  \n• Investigating the dependency of attributes or features in the dataset to potentially reduce the number of features. This can be done using methods such as ANOVA F-test to assess the dependency of each feature on the target variable or by using correlation coefficients.  \n• Data normalization to mitigate the variance of observations (dataset) within the same feature.  \n• Handling data imbalance using resampling techniques such as oversampling or undersampling.  \n• Combining grid search, cross-validation, and Bayesian optimization techniques to search for hyperparameters, including sensitive parameters like learning rate, epochs, batch size, and insensitive parameters to find suitable hyperparameters for the model.  \nII. RELATED WORKS  \nPredicting the activity of chemical compounds is a crucial area in pharmacology and chemoinformatics research. Several noteworthy research works in this field are as follows. The paper \"A Deep Learning Approach to Antibiotic Discovery\"  \n[1] explores the use of deep learning to predict the antibiotic activity of compounds, aiding in identifying compounds with antibacterial properties. The study experimented with over 107 million molecules from the ZINC15 database. The results identified eight antibacterial compounds with distinct structures from known antibiotics. This research highlights the role of deep learning methods in expanding the arsenal of antibiotics by discovering structurally different antibacterial molecules. It represents a significant breakthrough in antibiotic discovery, potentially aiding in finding new compounds to combat antibiotic-resistant bacteria. The paper \"Predicting Antitumor Activity of Peptides by Consensus of Regression Models Trained on a Small Data Sample\" [2] focuses on solving regression problems with a small dataset of only 429 compounds. The study employs various methods such as linear regression, polynomial regression, Gaussian kernel regression, neural networks, k-nearest neighbors (kNN), and support vector machines (SVM) to address overfitting. It develops a method to predict the anti-tumor activi","cbCaibFOvAJ1EOEO","https://ap.wps.com/l/cbCaibFOvAJ1EOEO","pdf",486692,1,10,"English","en",105,"# Introduction\n## Related Works\n## Machine Learning Problem Setup\n## Preprocessing and Modeling Strategies","[{\"question\":\"What is the main goal of this study?\",\"answer\":\"To use machine learning approaches to predict the activity of chemical compounds in cheminformatics as accurately as possible.\"},{\"question\":\"How are candidate ML solutions selected in the experiments?\",\"answer\":\"Solutions are evaluated using G-mean, F1-score, and AUC, then selected according to these criteria.\"},{\"question\":\"What dataset is used to test the models?\",\"answer\":\"The models are tested on a dataset of about 10,000 chemical compounds from PubChem classified according to their activity.\"}]","Predicting the Activity of Chemical Compounds Based on Machine Learning Approaches | PDF",1785818948,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predicting-the-activity-of-chemical-compounds-based-on-machine-learning-approaches","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/predicting-the-activity-of-chemical-compounds-based-on-machine-learning-approaches/123863/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of this study?","Question",{"text":75,"@type":76},"To use machine learning approaches to predict the activity of chemical compounds in cheminformatics as accurately as possible.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are candidate ML solutions selected in the experiments?",{"text":80,"@type":76},"Solutions are evaluated using G-mean, F1-score, and AUC, then selected according to these criteria.",{"name":82,"@type":73,"acceptedAnswer":83},"What dataset is used to test the models?",{"text":84,"@type":76},"The models are tested on a dataset of about 10,000 chemical compounds from PubChem classified according to their activity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]