[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119222-en":3,"doc-seo-119222-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119222,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Credit Scoring - A comparison between statistical and machine learning techniques for probability of default estimation","Credit risk, defined by the Basel Committee as the potential for a borrower to default on obligations, requires effective management to optimize risk-adjusted returns. This study examines a publicly available loan default dataset from the University of California, Irvine (UCI) Machine Learning Repository, comparing statistical and machine learning models. The analysis evaluates strengths and limitations in estimating default probability, emphasizing robust data preprocessing, careful model selection, and interpretability methods. Findings highlight the complex interplay of factors shaping credit risk outcomes.","Credit Scoring: A comparison between statistical and machine learning techniques for probability of default estimation  \nJoão D. Sousa Dias  \nMaster in Monetary and Financial Economics  \nSupervisor:  \nProf. Paulo Viegas de Carvalho, Invited Assistant Professor, ISCTE Business School  \nSeptember, 2024  \nDepartment of Political Economy  \nCredit Scoring: A comparison between statistical and machine learning techniques for probability of default estimation  \nJoão D. Sousa Dias  \nMaster in Monetary and Financial Economics  \nSupervisor:  \nProf. Paulo Viegas de Carvalho, Invited Assistant Professor, ISCTE Business School  \nSeptember, 2024  \nAcknowledgements  \nThis work would not have been possible without the help, support and guidance of different people. I would like to thank them for their impact on my life and subsequently, on this work.  \nFirst of all, I would like to express my gratitude to Prof. Paulo Viegas de Carvalho. His expertise, attention and advice were crucial for me, helping me surpassing many roadblocks found through all stages ofthis work.  \nTo my whole family, more particularly to my grandfather José António for being an academia figure forme. His constant motivation and support helped me becoming a learning enthusiast, trait that I always find to be useful.  \nTo Diana, mainly for her love and patience, but also for showing me another perspective on different things in life. Watching different plays and going into museums loosened my mind when this work most needed.  \nTo all my friends, that along this journey were always there for me to help me distracting and laughing, the latter being the thing that I most enjoy.  \nTo ISCTE, for receiving me as their student, helping me embracing my goals and providing me the tools to learn, this being the second thing most enjoyable in my life.  \nOnce again, from the bottom of my heart, thank you all!  \nii  \nResumo  \nRisco de crédito é definido pelo Comité de Basileia como a probabilidade de um devedor entrar emincumprimento para com as suas obrigações creditícias, sendo que é necessária uma gestão efetiva domesmo para otimizar rendibilidades ajustadas ao risco. Esta dissertação pretende ser um estudo sobre um conjunto de dados de empréstimos concedidos, publicamente disponível no repositório de Machine Learning da Universidade da California, Irvine (UCI), onde uma comparação é efetuada entre modelosestatísticos e modelos baseados em machine learning. Esta análise comparativa evidencia os vários pontos fortes e limitações respetivos a cada tipo de modelo, pelo aprofundamento das suas característicase resultados na estimação da probabilidade de incumprimento. As conclusões apontam para aimportância de um tratamento de dados robusto, da seleção do melhor modelo e na utilização de técnicas de interpretabilidade, destacando a complexidade dos vários fatores que influenciam o risco de crédito.  \niv  \nAbstract  \nCredit risk, defined by the Basel Committee as the potential for a borrower to default on obligations, necessitates effective management to optimize risk-adjusted returns. This work intends to be a study on a publicly available loan default dataset from the University of California, Irvine (UCI) Machine Learning Repository, where a comparison is conducted between statistical and machine learning models. The comparative analysis of these models highlights their strengths and limitations, offering insights into their application in credit risk assessment. The findings underscore the importance of robust data preprocessing, model selection, and interpretability techniques in predicting credit defaults, highlighting the complex interplay of various factors influencing credit risk.  \nvi  \nINDEX  \nAcknowledgements i  \nResumo iii  \nAbstract v  \nChapter 1: Introduction 13  \nChapter 2: Literature review 17  \nChapter 3: Methodology 21  \n3.1 Statistical Models 21  \n3.1.1 Logistic Regression 21  \n3.2 Machine Learning Models 22  \n3.2.1 Decision Trees 22  \n3.2.2 Random Forests 23  \n3.2.3 ","cbCailCw0zZixIUQ","https://ap.wps.com/l/cbCailCw0zZixIUQ","pdf",1009407,1,65,"English","en",105,"# Acknowledgements\n# Resumo\n# Abstract\n# Index\n## Chapter 1: Introduction\n## Chapter 2: Literature review\n## Chapter 3: Methodology\n## Chapter 4: Exploratory Data Analysis\n## Chapter 5: Results and explanations\n## Chapter 6: Conclusions and recommendations for future work\n## References\n## Annex","[{\"question\":\"What dataset is used for the credit scoring comparison?\",\"answer\":\"The study uses a publicly available loan default dataset from the University of California, Irvine (UCI) Machine Learning Repository.\"},{\"question\":\"Which model families are compared in the methodology?\",\"answer\":\"The comparison includes statistical models, such as logistic regression, and machine learning models, including decision trees, random forests, artificial neural networks, support vector machines, XGBoost, LightGBM, and AdaBoost.\"},{\"question\":\"Why are interpretability techniques emphasized in the results?\",\"answer\":\"Interpretability techniques are used to explain predictions through methods such as Shapley Additive Explanations and LIME, helping reveal the factors influencing credit risk.\"}]","Credit Scoring - A comparison between statistical and machine learning techniques for probability of default estimation | PDF",1785723158,164,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"credit-scoring-a-comparison-between-statistical-and-machine-learning-techniques-for-probability-of-default-estimation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/credit-scoring-a-comparison-between-statistical-and-machine-learning-techniques-for-probability-of-default-estimation/119222/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What dataset is used for the credit scoring comparison?","Question",{"text":75,"@type":76},"The study uses a publicly available loan default dataset from the University of California, Irvine (UCI) Machine Learning Repository.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which model families are compared in the methodology?",{"text":80,"@type":76},"The comparison includes statistical models, such as logistic regression, and machine learning models, including decision trees, random forests, artificial neural networks, support vector machines, XGBoost, LightGBM, and AdaBoost.",{"name":82,"@type":73,"acceptedAnswer":83},"Why are interpretability techniques emphasized in the results?",{"text":84,"@type":76},"Interpretability techniques are used to explain predictions through methods such as Shapley Additive Explanations and LIME, helping reveal the factors influencing credit risk.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]