[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117933-en":3,"doc-seo-117933-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117933,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","A Comparative Analysis of Machine Learning Models for Corporate Default Forecasting","This study evaluates the potential advantages of machine learning for corporate default forecasting by comparing the discriminatory power of random forest and XGBoost models against traditional statistical approaches. Using out-of-time predictions, results show superior discrimination for machine learning. Training sample size reduction lowers predictive power and narrows the performance gap, while model dimensionality changes affect statistical models only modestly. Adding predictors increases machine learning predictive power, and clustering improves discrimination across firm-size groups. Machine learning also classifies micro firms with notably higher ability and supports forecasting labour-market job risk for Portuguese non-financial micro cooperations.","A Comparative Analysis of Machine Learning Models for Corporate Default Forecasting  \nAlexander Seum  \nDissertation written under the supervision of  \nProfessor Eva Schliephake  \nDissertation submitted in partial fulfilment of requirements for the MSc in Economics, at the Universidade Católica Portuguesa, 5.4.2023.  \nA Comparative Analysis of Machine Learning Models for Corporate Default Forecasting  \nAlexander Seum  \nAbstract  \nThis study examines the potential benefits of utilizing machine learning models for default forecasting by comparing the discriminatory power of the random forest and XGBoost models with traditional statistical models. The results of the evaluation with out-of-time predictions show that the machine learning models exhibit a higher discriminatory power compared to the traditional models. The reduction in the sample size of the training dataset leads to a decrease in predictive power of the machine learning models, reducing the difference in performance between the two model types. While modifications in model dimensionality have a limited impact on the discriminatory power of the statistical models, the predictive power of machine learning models increases with the addition of further predictors. When employing a clustering approach, both traditional and machine learning models exhibit an improvement in discriminatory power in the small, medium, and large firm size clusters compared to the previous non-clustering specifications. Machine learning models exhibit a significantly higher ability to classify micro firms. The findings ofthis research indicate that the machine learning models exhibit superior discriminatory power compared to the traditional models across the different specifications. Machine learning models can be used to forecast the potential impact of corporate default of non-financial micro cooperations on the Portuguese labour market by estimating the number of jobs at risk.  \nKey words: credit risk, default forecasting, machine learning, random forest  \nAcknowledgements  \nI would like to express my sincere appreciation to Professor Eva Schliephake for serving as my supervisor during the completion of my master’s thesis. Throughout the process, her guidance, expertise, and support were invaluable to me.  \nI am grateful to Banco de Portugal for providing me with the opportunity to conduct my research and for all of the resources and support they provided me. I would like to extend a special thanks to Miguel Portela and the whole BPLIM team, who went above and beyond to help me with my work.  \nTable of Contents  \nI. Introduction .....................................................................................................6  \nII. Literature review ...........................................................................................8  \nIII. Model Building .............................................................................................9  \nIII.A. Research Objectives .............................................................................9  \nIII.B. Data Collection and Study Design .......................................................9  \nIII.C. Data Preparation................................................................................. 10  \nIII.D. Selection of Variables ........................................................................ 11  \nIII.D.i. Supervised Feature Selection 11  \nIII.D.ii. Unsupervised Feature Selection 12  \nIII.D.iii. Feature Selection Process 13  \nIII.E. Class Imbalance ................................................................................. 15  \nIV. Explanatory Data Analysis ........................................................................ 18  \nIV.A. Overview of Defaults ......................................................................... 18  \nIV.B. Firm Classification and Distribution.................................................. 18  \nV. Methods..............................................................","cbCaisRyC3wmBXb4","https://ap.wps.com/l/cbCaisRyC3wmBXb4","pdf",4565235,1,61,"English","en",105,"# I. Introduction\n# II. Literature review\n# III. Model Building\n## III.A. Research Objectives\n## III.B. Data Collection and Study Design\n## III.C. Data Preparation\n## III.D. Selection of Variables\n## III.E. Class Imbalance\n# IV. Explanatory Data Analysis\n## IV.A. Overview of Defaults\n## IV.B. Firm Classification and Distribution\n# V. Methods\n## V.A. Statistical Models\n## V.B. Machine Learning Models\n## V.C. Hyperparameter Tuning\n## V.D. Model Evaluation\n## V.E. Model Validation\n# VI. Results\n## VI.A. Type of Forecast\n## VI.B. Out-of-Time Predictions\n## VI.C. Sample Size of the Training Dataset\n## VI.D. Dimensionality\n## VI.E. Firm Size Clustering Approach\n## VI.F. Labour Market Implications of Corporate Default\n# VII. Discussion\n# VIII. Conclusion\n# References\n# Appendix","[{\"question\":\"Which machine learning models are compared for corporate default forecasting?\",\"answer\":\"The study compares random forest and XGBoost with traditional statistical models to assess discriminatory power for default forecasting.\"},{\"question\":\"How do out-of-time predictions affect the comparison between model types?\",\"answer\":\"Out-of-time evaluation shows machine learning models deliver higher discriminatory power than traditional statistical models.\"},{\"question\":\"What factors influence predictive performance in the machine learning approach?\",\"answer\":\"Reducing the training dataset size decreases predictive power for machine learning and narrows the performance difference; adding further predictors increases predictive power.\"}]","A Comparative Analysis of Machine Learning Models for Corporate Default Forecasting | PDF",1785680429,154,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-comparative-analysis-of-machine-learning-models-for-corporate-default-forecasting","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-comparative-analysis-of-machine-learning-models-for-corporate-default-forecasting/117933/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning models are compared for corporate default forecasting?","Question",{"text":75,"@type":76},"The study compares random forest and XGBoost with traditional statistical models to assess discriminatory power for default forecasting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do out-of-time predictions affect the comparison between model types?",{"text":80,"@type":76},"Out-of-time evaluation shows machine learning models deliver higher discriminatory power than traditional statistical models.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors influence predictive performance in the machine learning approach?",{"text":84,"@type":76},"Reducing the training dataset size decreases predictive power for machine learning and narrows the performance difference; adding further predictors increases predictive power.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]