[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116989-en":3,"doc-seo-116989-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116989,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Big Data for Credit Risk Analysis - Efficient Machine Learning Models Using PySpark","Big Data increasingly strengthens traditional credit scoring, especially when client data from open and new banking platforms is used for machine learning-based personal credit evaluation. Key challenges include Big Data quality and model risk. The work provides PySpark code that enables computationally efficient statistical learning and machine learning algorithms for this credit-scoring scenario, with a performance comparison of logistic regression, decision tree, random forest, neural network, and support vector machine. Results show logistic regression achieving stronger determination and fewer false negatives while remaining cheaper and more interpretable. The study also discusses practical steps, risks, and benefits of applying Big Data analytics to credit scoring.","Chapter 1  \nBig Data for Credit Risk Analysis: Efficient Machine Learning Models Using PySpark  \nThis paper and PySpark code was presented on 12 July 2018 and published by Springer in 2023. For citation, please use the following:  \nAshofteh, A. (2023)‘Big Data for Credit Risk Analysis: Efficient Machine Learning Models Using PySpark’, in Pilz, J. , Melas, V. B. , and Bathke, A. (eds) Statistical Modeling and Simulation for Experimental Design and Machine Learning Applications. Cham: Springer International Publishing, pp. 245–265. doi: 10. 1007/978-3-031-40055-1_ 14.  \nAfshin Ashofteh  \nAbstract Recently, Big Data has become an increasingly important source to support traditional credit scoring. Personal credit evaluation based on machine learning approaches focuses on the application data of clients in open banking and new banking platforms with challenges about Big Data quality and model risk. This paper represents a PySpark code for computationally efficient use of statistical learning and machine learning algorithms for the application scenario of personal credit evaluation with a performance comparison of models including logistic regression, decision tree, random forest, neural network, and support vector machine. The findings of this study reveal that the logistic regression methodology represents a more reasonable coefficient of determination and a lower false-negative rate than other models. Additionally, it is computationally less expensive and more comprehensible. Finally, the paper highlights the steps, perils, and benefits of using Big Data and machine learning algorithms in credit scoring.  \nKeywords: Credit score; Big Data; Machine Learning; Risk Management; Finance.  \n1.1 INTRODUCTION  \nRisk management with the ability to incorporate new and Big Data sources and benefit from emerging technologies such as cloud and parallel computing platforms is critically important for financial service providers, supervisory authorities, and regulators if they are to remain competitive and relevant [1] .  \nFinancial institutions’ growing interest in non-traditional data may be seen as a hypothetical occurrence, a reaction to the most recent financial crisis.  \nAfshin Ashofteh  \nNOVA Information Management School (NOVA IMS), Universidade Nova de Lisboa, Campus de Campolide, 1070-312 Lisboa, Portugal. e-mail: [aashofteh@novaims.unl.pt](aashofteh@novaims.unl.pt)  \nHowever, the financial crisis not only prompted several statutory and supervisory initiatives that require significant disclosure of data but also provided a positive atmosphere to get the advantages of new data sources such as non-traditional data sets [2, 3] .  \nThere are sources of supply and demand for this increased acceptance of non-traditional data. On the supply side, technology advancements like mobile phones [4] that have expanded storage space and computing power while cutting costs have fueled rises in the new data sources. In addition, mobile data and social data have recently been used to monitor different risks [5] . On the demand side, loan providers are becoming more interested in learning how data analysis may improve credit scoring and lower the risk of default [6] .  \nSome of the largest and most established financial institutions such as banks, insurance companies, payday lenders, peer-to-peer lending platforms, microfinance providers, leasing companies, and payment by installment companies are now taking a fresh look at their customers’ transactional data to enhance the early detection of fraud. They use innovative machine learning models that exploit novel data sources like Big Data, social data, and mobile data. Credit risk management may benefit in the long run if these advancements result in better credit choices. However, there are shorter-term hazards if early users of non-traditional data credit scoring mostly disregard the model risk and technical aspects of new methods that might affect credit scoring [7] . For instance, one crucial issue ","cbCaigoIBA9aJjPy","https://ap.wps.com/l/cbCaigoIBA9aJjPy","pdf",2533376,1,21,"English","en",105,"# 1.1 Introduction\n## Risk management importance\n## Non-traditional data and market drivers\n## Challenges of model risk and class imbalance\n# 1.2 Data Processing\n## Data processing with PySpark","[{\"question\":\"What scenario does the paper focus on in credit risk analysis?\",\"answer\":\"The paper focuses on personal credit evaluation using client application data from open banking and new banking platforms, addressing issues of Big Data quality and model risk.\"},{\"question\":\"Which machine learning models are compared, and what is the main performance finding?\",\"answer\":\"Logistic regression, decision tree, random forest, neural network, and support vector machine are compared. Logistic regression shows a more reasonable coefficient of determination and a lower false-negative rate, with lower computational cost and better interpretability.\"},{\"question\":\"Why is model risk more critical when using non-traditional Big Data for credit scoring?\",\"answer\":\"The document highlights hazards such as class imbalance from rare distress events and the difficulty of forecasting defaults in sparse, imbalanced environments, which can affect model reliability over time.\"}]","Big Data for Credit Risk Analysis - Efficient Machine Learning Models Using PySpark | PDF",1785672985,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"big-data-for-credit-risk-analysis-efficient-machine-learning-models-using-pyspark","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/big-data-for-credit-risk-analysis-efficient-machine-learning-models-using-pyspark/116989/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What scenario does the paper focus on in credit risk analysis?","Question",{"text":75,"@type":76},"The paper focuses on personal credit evaluation using client application data from open banking and new banking platforms, addressing issues of Big Data quality and model risk.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are compared, and what is the main performance finding?",{"text":80,"@type":76},"Logistic regression, decision tree, random forest, neural network, and support vector machine are compared. Logistic regression shows a more reasonable coefficient of determination and a lower false-negative rate, with lower computational cost and better interpretability.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is model risk more critical when using non-traditional Big Data for credit scoring?",{"text":84,"@type":76},"The document highlights hazards such as class imbalance from rare distress events and the difficulty of forecasting defaults in sparse, imbalanced environments, which can affect model reliability over time.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]