[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121483-en":3,"doc-seo-121483-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121483,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Non-small Cell Lung Cancer Active Compounds Discovery - Protein Expressions Integration Using Machine Learning Models","Computational methods reshape drug discovery by enabling rapid identification of promising compounds through machine learning. A comparative study trains multiple algorithms to predict active compounds targeting non-small cell lung cancer while integrating protein expression information. Bioactivity data are extracted from the ChEMBL database, molecular descriptors are computed to capture structure–activity relationships, and models are evaluated by performance metrics. The multilayer perceptron achieves the highest F1 score (0.861). A list of 10 literature-supported candidate drugs is provided.","Non-small cell lung cancer active compounds discovery holding on protein expression using machine learning models  \nHamza Hanafi1, M’hamed Aït Kbir1, Badr Dine Rossi Hassani2  \n1Intelligent Automation and BioMedGenomics Laboratory, STSM Doctoral Center, Abdelmalek Essaadi University, Tangier, Morocco 2LABIPHABE Laboratory, STI Doctoral Center, Abdelmalek Essaadi University, Tangier, Morocco  \nArticle Info ABSTRACT  \nArticle history:  \nReceived May 23, 2024 Revised Feb 25, 2025 Accepted Mar 15, 2025  \nKeywords:  \nDrug discovery  \nLung cancer  \nMachine learning models Precision medicine Protein expressions  \nCorresponding Author:  \nComputational methods have transformed the field of drug discovery, which significantly helped in the development of new treatments. Nowadays, researchers are exploring a wide ranger of opportunities to identify new compounds using machine learning. We conducted a comparative study between multiple models capable of predicting compounds to target nonsmall cell lung cancer, we focused on integrating protein expressions to identify potential compounds that exhibit a high efficacy in targeting lung cancer cells. A dataset was constructed based on the trials available in the ChEMBL database. Then, molecular descriptors were calculated to extract structure-activity relationships from the selected compounds and feed into several machine learning models to learn from. We compared the performance of various algorithms. The multilayer perceptron model exhibited the highest F1 score, achieving an outstanding value of 0,861 . Moreover, we present a list of 10 drugs predicted as active in lung cancer, all of which are supported by relevant scientific evidence in the medical literature. Our study showcases the potential of combining protein expression analysis and machine learning techniques to identify novel drugs. Our analytical approach contributes to the drug discovery pipeline, and opens new opportunities to explore and identify new targeted therapies.  \nThis is an open access article under the CC BY-SA license.  \nHamza Hanafi  \nIntelligent Automation and BioMedGenomics Laboratory, STSM Doctoral Center Abdelmalek Essaadi University  \nTangier, Morocco  \n[Email: hamza.hanafi@etu.uae.ac.ma](Email: hamza.hanafi@etu.uae.ac.ma)  \n1. INTRODUCTION  \nDrug discovery plays a fundamental role in the healthcare sector, as developing new compounds demands a multidisciplinary approach to provide novel therapeutic interventions. Despite this, the process is often complex, time-consuming, and requires an enormous effort to validate new treatments. Moreover, traditional methods of drug discovery are not only resource-intensive but also limited in their scope [1] .  \nRecent advancements in computational biology have completely transformed drug discovery pipelines. The combination of biology with computational methods offers new insights to accelerate the identification and evaluation of novel compounds. Therefore, computational techniques have emerged as powerful tools in the field of pharmacological medicine [2], and revealed great success compared to traditional methods. Besides, these techniques have found widespread application in various healthcare domains, including disease classification [3] and surgical enhancements [4] .  \nNowadays, a large amount of biological data is stored in public databases and enables researchers to explore a wide range of methodologies. Furthermore, the integration and analysis of this biological data ease  \nthe study of new hypotheses [5], for example, predictive modeling using machine learning (ML) techniques is one of the most explored methodologies and has gained prominence. ML models can effectively classify drugs into relevant therapeutic categories, accurately detect and classify tumor stages [6], and design new drugs based on chemical properties [7] . Consequently, ML-based methods are capable in detecting patterns and identifying correlations within large and complex datasets with numer","cbCaik17shFimL5F","https://ap.wps.com/l/cbCaik17shFimL5F","pdf",903195,1,11,"English","en",105,"# Introduction\n## Computational biology and drug discovery pipelines\n## Public biological data and machine learning applications\n## Bioinformatics integration for precision medicine\n## Challenges in building ML models for drug discovery\n# Methodology\n## Dataset curation from ChEMBL using NSCLC protein expression\n## Molecular descriptor computation and feature preparation\n## Model training and comparative evaluation","[{\"question\":\"What is the main objective of the study?\",\"answer\":\"To develop and compare machine learning models that predict active compounds targeting non-small cell lung cancer by integrating protein expression information.\"},{\"question\":\"How was the dataset constructed?\",\"answer\":\"Bioactivity trials were extracted from the ChEMBL database based on proteins expressed in non-small cell lung cancer, then used to derive molecular descriptors.\"},{\"question\":\"Which machine learning model performed best?\",\"answer\":\"The multilayer perceptron model achieved the highest F1 score, reaching 0.861.\"}]","Non-small Cell Lung Cancer Active Compounds Discovery - Protein Expressions Integration Using Machine Learning Models | PDF",1785735856,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"non-small-cell-lung-cancer-active-compounds-discovery-protein-expressions-integration-using-machine-learning-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/non-small-cell-lung-cancer-active-compounds-discovery-protein-expressions-integration-using-machine-learning-models/121483/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main objective of the study?","Question",{"text":75,"@type":76},"To develop and compare machine learning models that predict active compounds targeting non-small cell lung cancer by integrating protein expression information.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was the dataset constructed?",{"text":80,"@type":76},"Bioactivity trials were extracted from the ChEMBL database based on proteins expressed in non-small cell lung cancer, then used to derive molecular descriptors.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine learning model performed best?",{"text":84,"@type":76},"The multilayer perceptron model achieved the highest F1 score, reaching 0.861.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]