[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125881-en":3,"doc-seo-125881-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125881,1099523885336,"Violet","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Machine Learning 工作流用于信用违约预测 - 工作流方法","Rising FinTech interest makes credit default prediction a central task for assessing borrower creditworthiness and guiding loan approval and risk management decisions. The paper presents a workflow-based approach that improves CDP through multiple pipeline steps tailored to different machine learning strengths. It begins with Weight of Evidence preprocessing for outlier removal, missing-value handling, and unified scaling, then trains model families with ensemble learning and multi-objective genetic hyperparameter optimization. The workflow targets both predictive accuracy and financial considerations to support more reliable credit risk assessment.","A machine learning work􀀃ow to address credit default prediction  \narXiv :2403 .03785v 1 [ cs .CE] 6 Mar 2024  \nRambod Rahmani 1a , Marco Parola 1b and Mario G.C.A. Cimino 1c  \n1Dept. of Information Engineering, University of Pisa, Largo L. Lazzarino 1, Pisa, Italy {r.rahmani@studenti, [marco.parola@.ing](marco.parola@.ing), [mario.cimino@](mario.cimino@}.unipi.it)[}](mario.cimino@}.unipi.it)[.unipi.it](mario.cimino@}.unipi.it)  \nKeywords: FinTech, Credit Scoring, Default Credit Prediction, Machine Learning, NSGA-II, Weight of Evidence.  \nAbstract: Due to the recent increase in interest in Financial Technology (FinTech), applications like credit default prediction (CDP) are gaining signi􀀂cant industrial and academic attention. In this regard, CDP plays a crucial role in assessing the creditworthiness of individuals and businesses, enabling lenders to make informed decisions regarding loan approvals and risk management. In this paper, we propose a work􀀃ow-based approach to improve CDP, which refers to the task of assessing the probability that a borrower will default on his or her credit obligations. The work􀀃ow consists of multiple steps, each designed to leverage the strengths of different techniques featured in machine learning pipelines and, thus best solve the CDP task. We employ a comprehensive and systematic approach starting with data preprocessing using Weight of Evidence encoding, a technique that ensures in a single-shot data scaling by removing outliers, handling missing values, and making data uniform for models working with different data types. Next, we train several families of learning models, introducing ensemble techniques to build more robust models and hyperparameter optimization via multi-objective genetic algorithms to consider both predictive accuracy and 􀀂nancial aspects. Our research aims at contributing to the FinTech industry in providing a tool to move toward more accurate and reliable credit risk assessment, bene􀀂ting both lenders and borrowers.  \n1 Introduction and background  \nIn the 􀀂nancial sector, credit scoring is a crucial task in which lenders must assess the creditworthiness of potential borrowers. In order to determine credit risk, several characteristics related to income, credit history, and other relevant aspects of the borrower must be deeply investigated.  \nTo manage 􀀂nancial risks and make critical decisions about whether to lend money to their customers, banks and other 􀀂nancial organizations must gather consumer information to identify reliable borrowers from those unable to repay debt. This results in solving a credit default prediction problem, or in other words a binary classi􀀂cation problem (Moula et al., 2017) .  \nIn order to address this challenge, over the years several statistical techniques have been embedded in a wide range of applications for the development of 􀀂nancial services in credit scoring and risk assessment (Sudjianto et al., 2010;  \nDevi and Radhika, 2018) . However, such models ofa [https://orcid.org/0009-0009-2789-5397](https://orcid.org/0009-0009-2789-5397)  \nb [https://orcid.org/0000-0003-4871-4902](https://orcid.org/0000-0003-4871-4902)  \nc [https://orcid.org/0000-0002-1031-1959](https://orcid.org/0000-0002-1031-1959)  \nten struggle to represent complex 􀀂nancial patterns because they rely on 􀀂xed functions and statistical assumptions (Luo et al., 2017) . While they have some advantages such as transparency and interpretability, their performance tends to suffer when faced with the challenges presented by the vast amounts of data and intricate relationships in credit prediction tasks.  \nOn the contrary, Deep Learning (DL) approaches have garnered signi􀀂cant attention across diverse domains, including the 􀀂nancial sector. This is due to their superior performance compared to traditional statistical and Machine Learning (ML) models (Teles et al., 2020) . In particular, DL has made great strides in several application areas, such as medical imaging (Parola et ","cbCainT7HKFNjM9G","https://ap.wps.com/l/cbCainT7HKFNjM9G","pdf",309012,6,1,9,"English","en",105,"# Introduction and background\n## Credit scoring and default prediction as classification\n## Limitations of traditional statistical models\n## Advantages of deep learning in finance\n## Preprocessing with Weight of Evidence (WoE)\n## Study goal and proposed workflow\n## Evaluation on benchmark datasets\n## Paper organization","[{\"question\":\"What is the main goal of the proposed work for credit default prediction?\",\"answer\":\"It proposes a workflow-based method to improve credit default prediction by combining preprocessing, model training, ensembling, and multi-objective hyperparameter optimization.\"},{\"question\":\"Why is Weight of Evidence (WoE) encoding used in the workflow?\",\"answer\":\"WoE supports target encoding to capture nonlinear feature-target relationships, handles missing values via separate binning, and scales numerical and categorical features into a unified continuous representation.\"},{\"question\":\"How does the workflow address financial objectives and class imbalance?\",\"answer\":\"It includes hyperparameter optimization using multi-objective genetic algorithms to consider both predictive accuracy and financial aspects, and it introduces a loss function focused on hard-to-classify examples to mitigate imbalance.\"}]","Machine Learning 工作流用于信用违约预测 - 工作流方法 | PDF",1785901815,23,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"machine-learning-workflow-for-credit-default-prediction-workflow-approach","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/machine-learning-workflow-for-credit-default-prediction-workflow-approach/125881/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What is the main goal of the proposed work for credit default prediction?","Question",{"text":77,"@type":78},"It proposes a workflow-based method to improve credit default prediction by combining preprocessing, model training, ensembling, and multi-objective hyperparameter optimization.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Why is Weight of Evidence (WoE) encoding used in the workflow?",{"text":82,"@type":78},"WoE supports target encoding to capture nonlinear feature-target relationships, handles missing values via separate binning, and scales numerical and categorical features into a unified continuous representation.",{"name":84,"@type":75,"acceptedAnswer":85},"How does the workflow address financial objectives and class imbalance?",{"text":86,"@type":78},"It includes hyperparameter optimization using multi-objective genetic algorithms to consider both predictive accuracy and financial aspects, and it introduces a loss function focused on hard-to-classify examples to mitigate imbalance.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":108,"slug":138},19,"General","general"]