[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127922-en":3,"doc-seo-127922-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127922,687207024478,"Liam","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Comparative Analysis of Data Preprocessing Methods, Feature Selection Techniques and Machine Learning Models for Improved Classification and Regression Performance on Imbalanced Genetic Data - arXiv","Rapid advancements in genome sequencing have produced large-scale genomics datasets used to predict mutation pathogenicity and clinical significance with machine learning. Many such datasets introduce two major difficulties: imbalanced target variables that skew regression distributions or create class imbalance in classification, and high-cardinality, skewed predictor variables that further complicate model learning. This study evaluates how data preprocessing, feature selection, and model choice affect performance on imbalanced genetic data, using 5-fold cross-validation and comparing averaged r-squared and accuracy across method combinations.","arXiv :2402 . 14980v1 [ q-bio .QM] 22 Feb 2024  \nComparative Analysis of Data Preprocessing Methods, Feature Selection Techniques and Machine Learning Models for Improved Classi􀀌cation and Regression Performance on Imbalanced Genetic Data  \nArshmeet Kaur1* and Morteza Sarmadi2  \n1* Biology Student, Evergreen Valley College, 3095 Yerba Buena Rd, San  \nJose, 95135, CA, U.S.A.  \n2 Research Scientist, Gilead Sciences, 333 Lakeside Dr, Foster City, 94404, CA, U.S.A.  \n*Corresponding author(s). E-mail(s): [Arka7783@stu.evc.edu](Arka7783@stu.evc.edu) ; Contributing [authors:](authors: msarmadi@mit.edu)[ msarmadi@mit.edu](authors: msarmadi@mit.edu);  \nAbstract  \nRapid advancements in genome sequencing have led to the collection of vast amounts of genomics data. Researchers may be interested in using machine learning models on such data to predict the pathogenicity or clinical signi􀀌cance of a genetic mutation. However, many genetic datasets contain imbalanced target variables that pose challenges to machine learning models: observations are skewed/imbalanced in regression tasks or class-imbalanced in classi􀀌cation tasks. Genetic datasets are also often high-cardinal and contain skewed predictor variables, which poses further challenges. We aimed to investigate the e􀀋ects of data preprocessing, feature selection techniques, and model selection on the performance of models trained on these datasets. We measured performance with 5-fold cross-validation and compared averaged r-squared and accuracy metrics across di􀀋erent combinations of techniques. We found that outliers/skew in predictor or target variables did not pose a challenge to regression models. We also found that class-imbalanced target variables and skewed predictors had little to no impact on classi􀀌cation performance. Random forest was the best model to use for imbalanced regression tasks. While our study uses a genetic dataset as an example of a real-world application, our 􀀌ndings can be generalized to any similar datasets.  \n1  \n1 Introduction  \nWhen dealing speci􀀌cally with predicting the pathogenicity/clinical signi􀀌cance of a mutation, we face two problems: either imbalanced regression or class-imbalanced classi􀀌cation. Our study aims to 􀀌nd methods that lead to the best performance of imbalanced regression and class-imbalanced classi􀀌cation tasks for models trained on data with the two characteristics mentioned above. While we are using a genetic dataset, our results are applicable to any datasets that share those characteristics of high-cardinality and skewed predictors.  \nThe 􀀌rst problem, imbalanced regression, occurs when machine learning models aim to predict a continuous, skewed target variable. For example, consider CADD  PHRED, a score ranging from 0 (benign) to 1 (most deleterious) (Niroulaand Vihinen, 2019). Most missense mutations (¿70%) are mildly deleterious (Kryukovet al. , 2007a) . Additionally, medical researchers studying disease might be more likely to focus on deleterious mutations. If they obtain data from a source focusing on the clinical/disease signi􀀌cance of genetic variants, such as Clinvar, observations may be skewed toward deleterious mutations. Thus, we would expect CADD  PHRED, in the dataset used here, to be skewed heavily towards deleterious mutations.  \nThe imbalanced regression problem poses two main challenges (Ribeiro and Moniz, 2020): 1) normal distributions are a common assumption of machine learning models and earlier utility-based regression. 2) Attempts to optimize model performance often result in severe bias. Many real-world target variables must be modeled as continuous variables, as categories do not provide enough information. Additionally, the majority of the research on imbalanced target variables focuses on classi􀀌cation tasks. To deal with datasets containing skewed predictors and targets, past research has attempted to use preprocessing methods to address skewed predictor variables (Branco and Torgo, 2019) . Another poss","cbCaia5rpc8uHsOq","https://ap.wps.com/l/cbCaia5rpc8uHsOq","pdf",937695,4,1,21,"English","en",105,"# Abstract\n# Introduction\n## Imbalanced regression challenges\n## Class-imbalanced classification challenges\n## Study objectives","[{\"question\":\"What problems does the study target in genetic machine learning tasks?\",\"answer\":\"It targets imbalanced regression (skewed continuous targets) and class-imbalanced classification (skewed class labels), alongside high-cardinality and skewed predictors in genetic datasets.\"},{\"question\":\"How did the authors evaluate model performance?\",\"answer\":\"They used 5-fold cross-validation and compared averaged r-squared for regression and accuracy for classification across different preprocessing, feature selection, and model combinations.\"},{\"question\":\"What main findings were reported about preprocessing and feature selection effects?\",\"answer\":\"Outliers or skew in predictor/target variables did not challenge regression models. For classification, class-imbalanced targets and skewed predictors had little to no impact on performance.\"}]","Comparative Analysis of Data Preprocessing Methods, Feature Selection Techniques and Machine Learning Models for Improved Classification and Regression Performance on Imbalanced Genetic Data - arXiv | PDF",1785942977,53,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"comparative-analysis-of-data-preprocessing-methods-feature-selection-techniques-and-machine-learning-models-for-improved-classification-and-regression-performance-on-imbalanced-genetic-data-arxiv","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/comparative-analysis-of-data-preprocessing-methods-feature-selection-techniques-and-machine-learning-models-for-improved-classification-and-regression-performance-on-imbalanced-genetic-data-arxiv/127922/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-28","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problems does the study target in genetic machine learning tasks?","Question",{"text":76,"@type":77},"It targets imbalanced regression (skewed continuous targets) and class-imbalanced classification (skewed class labels), alongside high-cardinality and skewed predictors in genetic datasets.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How did the authors evaluate model performance?",{"text":81,"@type":77},"They used 5-fold cross-validation and compared averaged r-squared for regression and accuracy for classification across different preprocessing, feature selection, and model combinations.",{"name":83,"@type":74,"acceptedAnswer":84},"What main findings were reported about preprocessing and feature selection effects?",{"text":85,"@type":77},"Outliers or skew in predictor/target variables did not challenge regression models. For classification, class-imbalanced targets and skewed predictors had little to no impact on performance.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]