[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120875-en":3,"doc-seo-120875-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120875,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Machine learning and credit risk: Empirical evidence from small-and mid-sized businesses - Research report","This paper compares two approaches for estimating credit risk for small- and mid-sized businesses (SMBs): a classic parametric method based on an ordered probit model and a nonparametric machine learning approach using a historical random forest (HRF) model calibrated on past time-series behavior. The study uses a proprietary granular quarterly dataset covering 464 Italian SMBs from 2015–2017 from a European investment bank and an international insurance company. Results show the HRF model outperforming the ordered probit model, particularly under high information asymmetry, and Shapley values are used to evaluate variable relevance in credit risk prediction.","Socio-Economic Planning Sciences 90 (2023) 101746  \n| Machine learning and credit risk: Empirical evidence from small-and mid-sized businesses |  |  |  |\n| --- | --- | --- | --- |\n| Alessandro Bitettoa,∗, Paola Cerchielloa, Stefano Filomenib, Alessandra Tandaa, Barbara Tarantino a\u003Cbr>a University of Pavia, Department of Economics and Management, Italy b University of Essex, Essex Business School, Finance Group, Colchester, UK |  |  |  |\n| A R T I C L E I N F O |  | A B S T R A C T |  |\n| JEL classification:\u003Cbr>C52 C53 D82 D83 G21\u003Cbr>G22\u003Cbr>Keywords: Credit rating SMB\u003Cbr>Historical random forest Machine learning Relationship banking Invoice lending |  | In this paper, we compare two different approaches to estimate the credit risk for small- and mid-sized businesses (SMBs), namely a classic parametric approach, by fitting an ordered probit model, and a nonparametric approach, calibrating a machine learning historical random forest (HRF) model. The models are applied to a unique and proprietary dataset comprising granular firm-level quarterly data collected from a European investment bank and an international insurance company on a sample of 464 Italian SMBs over the period 2015–2017. Results show that the HRF approach outperforms the traditional ordered probit model, highlighting how advanced estimation methodologies that use machine learning techniques can be successfully implemented to predict SMB credit risk, i.e. when facing high asymmetries of information. Moreover, by using Shapley values, we are able to assess the relevance of each variable in predicting SMB credit risk. |  |\n\n1. Introduction  \nThe determination of corporate credit ratings and credit risk is a key topic that has alimented the academic debate both theoretically and empirically and has extreme relevance for the industry and the regulatory and supervisory bodies in the financial system as it is a tool to ensure the allocational efficiency of financial markets and intermediaries [3–5].  \nThe determination of corporate ratings is particularly challenging for small- or mid-businesses (SMBs) that are mostly unlisted. SMBs indeed represent a large segment of the corporate market in several economies, especially in the European context, and they are generally subject to relevant asymmetries of information that make the estimation of accurate credit ratings more difficult.  \nIn this context, many scholars have investigated the opportunity to use alternative sources of information, also including soft information derived from intensive relationship banking [6–8], whose importance also has been acknowledged by regulators.1 Soft information, however, might not be always effective in improving lending activity by  \nbanks [9] and it cannot be easily transferred in especially complex organizations [10–15] and this advocates the need to find alternative solutions to exploit hard information to obtain to more accurate credit ratings also for informationally opaque SMBs.  \nOver time, the technological and methodological advancementsin the models determining credit risk and credit rating allowed the inclusion of sophisticated AI techniques to improve credit rating accuracy, which however might suffer from limited explainability. Except for a few studies implementing alternative methodologies [16–18], the literature has been mainly focused on the types of information a financial intermediary should use in assessing SMB credit risk. To date, limited evidence is provided on the opportunity to employ Explainable AI methodologies to estimate SMB credit ratings and few studies perform a comparison of classic parametric and machine learning (ML) approaches.  \nTo fill this gap, this paper has the objective to provide a comparison between two alternative approaches to estimate SMBs credit ratings: on one side, we use a classic parametric approach, namely an ordered  \n∗ Correspondence to: Department of Economics and Management, University of Pavia, Via San Felice 5, 27100 Pavia, Ital","cbCaiecKPknkANNK","https://ap.wps.com/l/cbCaiecKPknkANNK","pdf",1820889,1,9,"English","en",105,"# Introduction\n## Credit ratings and information asymmetry for SMBs\n## Rationale for explainable and ML-based alternatives\n# Methods and Data (overview)\n## Ordered probit as parametric baseline\n## Historical Random Forest (HRF) as nonparametric approach\n## Proprietary dataset and study scope","[{\"question\":\"Which two modeling approaches are compared to estimate SMB credit risk?\",\"answer\":\"The paper compares an ordered probit model (classic parametric approach) with a historical random forest (HRF) machine learning model (nonparametric approach).\"},{\"question\":\"What dataset and sample are used in the study?\",\"answer\":\"The analysis uses a proprietary granular quarterly dataset of 464 Italian SMBs covering 2015–2017, drawn from a European investment bank and an international insurance company and matched with Orbis data.\"},{\"question\":\"How do the results evaluate the performance of HRF versus ordered probit?\",\"answer\":\"The HRF approach outperforms the traditional ordered probit model, indicating that machine learning techniques can improve credit risk prediction under high information asymmetries.\"}]","Machine learning and credit risk: Empirical evidence from small-and mid-sized businesses - Research report | PDF",1785732445,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-and-credit-risk-empirical-evidence-from-small-and-mid-sized-businesses-research-report","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-and-credit-risk-empirical-evidence-from-small-and-mid-sized-businesses-research-report/120875/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which two modeling approaches are compared to estimate SMB credit risk?","Question",{"text":75,"@type":76},"The paper compares an ordered probit model (classic parametric approach) with a historical random forest (HRF) machine learning model (nonparametric approach).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What dataset and sample are used in the study?",{"text":80,"@type":76},"The analysis uses a proprietary granular quarterly dataset of 464 Italian SMBs covering 2015–2017, drawn from a European investment bank and an international insurance company and matched with Orbis data.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the results evaluate the performance of HRF versus ordered probit?",{"text":84,"@type":76},"The HRF approach outperforms the traditional ordered probit model, indicating that machine learning techniques can improve credit risk prediction under high information asymmetries.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]