[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121758-en":3,"doc-seo-121758-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121758,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","A New Improved Prediction of Software Defects - Using Machine Learning-based Boosting Techniques with NASA Dataset","Predicting when and where software bugs will appear helps improve quality and reduce software testing costs by forecasting defects at the module level. The work addresses two key dataset challenges in software defect prediction: class imbalance caused by far fewer faulty modules than non-defective ones, and noisy features from irrelevant attributes that hinder accurate learning. The study proposes machine-learning enhancements to CatBoost and Gradient Boost classifiers, using Random Over Sampler and Mutual information-based feature selection, and validates them via 10-fold cross-validation on 11 NASA PROMISE datasets with multiple evaluation metrics.","[drsinhacs@gmail.com](drsinhacs@gmail.com)  \nAbstract—Predicting when and where bugs will appear in software may assist improve quality and save on software testing expenses. Predicting bugs in individual modules of software by utilizing machine learning methods. There are, however, two major problems with the software defect prediction dataset: Social stratification (there are many fewer faulty modules than non-defective ones), and noisy characteristics (a result of irrelevant features) that make accurate predictions difficult. The performance of the machine learning model will suffer greatly if these two issues arise. Overfitting will occur, and biassed classification findings will be the end consequence. In this research, we suggest using machine learning approaches to enhance the usefulness of the CatBoost and Gradient Boost classifiers while predicting software flaws. Both the Random Over Sampler and Mutual info classification methods address the class imbalance and feature selection issues inherent in software fault prediction. Eleven datasets from NASA's data repository, \"Promise,\" were utilised in this study. Using 10-fold cross-validation, we classified these 11 datasets and found that our suggested technique outperformed the baseline by a significant margin. The proposed methods have been evaluated based on their abilities to anticipate software defects using the most important indices available: Accuracy, Precision, Recall, F1 score, ROC values, RMSE, MSE, and MAE parameters. For all 11 datasets evaluated, the suggested methods outperform baseline classifiers by a significant margin. We tested our model to other methods of flaw identification and found that it outperformed them all. The computational detection rate of the suggested model is higher than that of conventional models, as shown by the experiments. .  \nKeywords-oftware Defect Prediction, Machine Learning, Class Imbalance, Feature Selection, NASA Promise Dataset, Catboost, Gradient Boost, Random Over Sampler, Cross Validation.  \nA New Improved Prediction of Software Defects Using Machine Learning-based Boosting Techniques  \nwith NASA Dataset  \nJayanti Goyal1, Ripu Ranjan Sinha2  \n1Research Scholar, Computer Science Department  \nRajasthan Technical University (RTU), Kota  \n[goyal.jayanti@gmail.com](goyal.jayanti@gmail.com)  \n2Professor, Computer Science, S. S. Jain Subodh P.G. College  \nRajasthan Technical University, Kota  \nI. INTRODUCTION  \nThe use of software has permeated every aspect of modern life. Software systems have had a significant impact on the economies of today's established and emerging nations, and software products are used in almost every industry and sector[1], from retail to transportation to banking to healthcare to government. Algorithms, procedures, and active modules makeup the programme. Allocation of resources and planning [2] including Time, human expertise, computer resources, tools, and infrastructure are all necessities in the design and development of a software system. As long as software plays a significant role, developers will need to think about how often bugs occur. Software failure rates are sometimes rising, even at companies with extensive development expertise. [3][4][5] .  \nWhen the outcomes of a software programme or product do not correspond to the needs of the end user, we have a software fault. The failures[6][7], unpredictability, or unexpected outcomes induced by these faults are the  \nconsequence of either source code or requirement problems. These issues negatively affect software quality and programme reliability and can lead to unnecessary expenditures of time, energy, and money. When problems occur, it takes more time and money to do maintenance. This makes early fault prediction in software a topic of interest for study. Over the course of the past two decades, academics have proposed several prediction models employing various machine learning classifiers. [8][9] . Inappropriately[10][11], Uneven data ","cbCaidu75CqNtpoz","https://ap.wps.com/l/cbCaidu75CqNtpoz","pdf",651808,1,13,"English","en",105,"# Introduction\n## Software quality, bugs, and the cost of failures\n## Software defect prediction challenge: imbalance and noisy features\n## Dataset preparation and public NASA PROMISE data\n## ML defect prediction workflow and evaluation goal","[{\"question\":\"What problem does software defect prediction aim to solve?\",\"answer\":\"It aims to predict when and where bugs will appear in software, enabling earlier interventions that improve quality and reduce testing costs.\"},{\"question\":\"Which dataset issues make software defect prediction difficult?\",\"answer\":\"Class imbalance (many more non-defective than faulty modules) and noisy characteristics from irrelevant features make accurate predictions harder and can degrade model performance.\"},{\"question\":\"How does the proposed approach improve CatBoost and Gradient Boost predictions?\",\"answer\":\"It enhances these classifiers using Random Over Sampler to address class imbalance and Mutual information-based feature selection to handle irrelevant features, then evaluates on NASA PROMISE datasets with 10-fold cross-validation.\"}]","A New Improved Prediction of Software Defects - Using Machine Learning-based Boosting Techniques with NASA Dataset | PDF",1785806679,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-new-improved-prediction-of-software-defects-using-machine-learning-based-boosting-techniques-with-nasa-dataset","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-new-improved-prediction-of-software-defects-using-machine-learning-based-boosting-techniques-with-nasa-dataset/121758/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does software defect prediction aim to solve?","Question",{"text":75,"@type":76},"It aims to predict when and where bugs will appear in software, enabling earlier interventions that improve quality and reduce testing costs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which dataset issues make software defect prediction difficult?",{"text":80,"@type":76},"Class imbalance (many more non-defective than faulty modules) and noisy characteristics from irrelevant features make accurate predictions harder and can degrade model performance.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed approach improve CatBoost and Gradient Boost predictions?",{"text":84,"@type":76},"It enhances these classifiers using Random Over Sampler to address class imbalance and Mutual information-based feature selection to handle irrelevant features, then evaluates on NASA PROMISE datasets with 10-fold cross-validation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]