[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125353-en":3,"doc-seo-125353-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125353,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Overfitting in predictive process monitoring - A Monte Carlo analysis of machine learning models - Master’s thesis","Predictive process monitoring aims to estimate remaining time until a case completes, often using machine learning. Limited evidence exists on how models react to changing data conditions, particularly regarding in-sample overfitting. This master’s thesis studies five machine learning models under systematic variations in data size, process complexity, and signal quality using controlled simulations. Synthetic event logs are generated with the SynBPS framework and throughput time is predicted from first-event features, evaluated via in-sample R2 and factorial ANOVA.","Master’s thesis 2025 30 ECTS  \nFaculty of Science and Technology  \nOverfitting in predictive process monitoring: A Monte Carlo analysis of machine learning models  \nKristoffer Lien  \nABSTRACT  \nBusiness processes can vary in complexity depending on how activities are structured and how resources are allocated. Estimating the remaining time until a case is completed is a common goal in predictive process monitoring, and machine learning is often used to solve this task. Little is known about how different models behave under varying data conditions, especially in terms of in-sample overfitting. This thesis investigates how five different machine learning models respond to changes in data size, process complexity, and signal quality using controlled simulations.  \nThe study is based on synthetic event logs generated with the SynBPS framework, which allows us to fully control the structure of the process and the data it produces. The target variable is throughput time, and the predictions are made using only the features of the first event in each trace. The five types of models tested are linear regression, lasso regression, random forest, gradient boost, and multilayer perceptron. For each model, the training sample size, the availability of resources, and the signal-to-noise ratio are systematically varied under all conditions. In-sample R2 is used as a performance metric and factorial ANOVA is applied to find which factors explain most of the variance.  \nIn all models, the sample size had the biggest impact on the sample R2 , especially in low-signal settings. The random forest achieved high R2 even when most features were irrelevant, suggesting that it may overfit in noisy environments. Linear and lasso regression showed more stable behavior, with limited capacity to overfit but also lower peak performance. The multilayer perceptron and gradient boosting models were more sensitive to signal quality and prone to erratic performance under poor data conditions.  \nIn general, the thesis highlights how the in-sample overfitting is shaped by the interaction between model flexibility, signal strength, and data volume. These insights can help researchers design better benchmarks and guide practitioners in choosing appropriate models depending on the nature of the available data.  \nContents  \n1 Introduction 5  \n1.0.1 Background and objectives .................. 5  \n1.1 Theory and previous research ..................... 6  \n1.1.1 Predictive process monitoring in business process management .............................. 6  \n1.1.2 Components of throughput time ............... 7  \nEvent-level decomposition of throughput time .......... 7  \nDistribution of throughput time ................. 8  \nTrace length .......................... 9  \nActivity duration ........................ 9  \nActivity offset ......................... 11  \nProcess instability ....................... 12  \nExpectation and variance of the resource availability offset .... 12  \nExpectation and variance of the business hours offset ....... 13  \n1.1.3 Classes of input features for predictive models ........ 14  \n1.1.4 Appropriate model families .................. 16  \nLinear regression ........................ 17  \nLasso regression ........................ 19  \nGradient boosting ........................ 20  \nRandom forest ......................... 21  \nMultilayer perceptron ...................... 22  \n1.1.5 Assessing sample size requirements ............. 23  \n1.2 Assessing in-sample overfitting .................... 24  \nCauses of in-sample overfitting ................. 25  \n1.3 Research questions .......................... 26  \n2 Method 28  \n2.1 Experimental design .......................... 28  \n2.1.1 Predictive value of manipulated features ........... 29  \n2.1.2 Number of additional features without predictive value ... 30  \n2.1.3 Number of traces ....................... 31  \n2.1.4 Process complexity ...................... 31  \n2.1.5 Resource availability ...........","cbCaisbcPCP2pgKb","https://ap.wps.com/l/cbCaisbcPCP2pgKb","pdf",525408,1,62,"English","en",105,"# Introduction\n## Background and objectives\n## Theory and previous research\n## Research questions\n# Method\n## Experimental design\n## Simulation setup\n## Prediction models\n## Use of AI\n# Results\n## Random forest\n## Gradient boosting\n## Perceptron\n## Linear regression\n## Lasso regression\n# Discussion\n# Conclusions and Recommendations","[{\"question\":\"What problem does the thesis address in predictive process monitoring?\",\"answer\":\"It analyzes how different machine learning models behave when predicting remaining throughput time, focusing on in-sample overfitting under varying data conditions.\"},{\"question\":\"How are the simulations and event logs generated?\",\"answer\":\"Synthetic event logs are generated with the SynBPS framework, which allows controlled manipulation of the process structure and the data producing conditions.\"},{\"question\":\"Which factors most strongly affect in-sample performance across models?\",\"answer\":\"Across all tested models, training sample size shows the largest impact on in-sample R2, especially under low-signal settings.\"}]","Overfitting in predictive process monitoring - A Monte Carlo analysis of machine learning models - Master’s thesis | PDF",1785898358,156,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"overfitting-in-predictive-process-monitoring-a-monte-carlo-analysis-of-machine-learning-models-masters-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/overfitting-in-predictive-process-monitoring-a-monte-carlo-analysis-of-machine-learning-models-masters-thesis/125353/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address in predictive process monitoring?","Question",{"text":75,"@type":76},"It analyzes how different machine learning models behave when predicting remaining throughput time, focusing on in-sample overfitting under varying data conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the simulations and event logs generated?",{"text":80,"@type":76},"Synthetic event logs are generated with the SynBPS framework, which allows controlled manipulation of the process structure and the data producing conditions.",{"name":82,"@type":73,"acceptedAnswer":83},"Which factors most strongly affect in-sample performance across models?",{"text":84,"@type":76},"Across all tested models, training sample size shows the largest impact on in-sample R2, especially under low-signal settings.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]