[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116871-en":3,"doc-seo-116871-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116871,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","TAILORING MACHINE LEARNING FOR PROCESS MINING","Machine learning models are commonly embedded in process mining pipelines for data transformation, noise reduction, anomaly detection, classification, and prediction. However, many approaches rely on ad-hoc assumptions about data distributions that conflict with the non-parametric distributions typically found in process data. Training also ignores constraints imposed by concurrency, and data encoding—though crucial—remains underused. The paper proposes a deeper analysis to align machine learning with process mining requirements and to support a sound integration methodology.","TAILORING MACHINE LEARNING FOR PROCESS MINING  \nPaolo Ceravolo  \nComputer Science Department Università degli Studi di Milano, Italy [paolo.ceravolo@unimi.it](paolo.ceravolo@unimi.it)  \narXiv :2306 . 10341v1 [ cs .LG] 17 Jun 2023  \nSylvio Barbon Junior  \nDepartment of Engineering and Architecture University of Trieste, Italy sylvio.barbonjunior@units .it  \nErnesto Damiani  \nCenter for Cyber-Physical Systems Khalifa University, Abu Dhabi, UAE [ernesto.damiani@ku.ac.ae](ernesto.damiani@ku.ac.ae)  \nWil van der Aalst  \nProcess and Data Science  \nRWTH Aachen University  \n[wvdaalst@pads.rwth-aachen.de](wvdaalst@pads.rwth-aachen.de)  \nABSTRACT  \nMachine learning models are routinely integrated into process mining pipelines to carry out tasks like data transformation, noise reduction, anomaly detection, classification, and prediction. Often, the design of such models is based on some ad-hoc assumptions about the corresponding data distributions, which are not necessarily in accordance with the non-parametric distributions typically observed with process data. Moreover, the learning procedure they follow ignores the constraints concurrency imposes to process data. Data encoding is a key element to smooth the mismatch between these assumptions but its potential is poorly exploited. In this paper, we argue that a deeper insight into the issues raised by training machine learning models with process data is crucial to ground a sound integration of process mining and machine learning. Our analysis of such issues is aimed at laying the foundation for a methodology aimed at correctly aligning machine learning with process mining requirements and stimulating the research to elaborate in this direction.  \nKeywords Process Mining · Machine Learning  \n1 Introduction  \nProcess Mining (PM) is a consolidated discipline grounded on data mining and business process management. The exploitation of traditional PM tasks (discovery, conformance checking, and enhancement) is today a reality in many organizations [1, 2] . In the last decade, a wave of new results in artificial intelligence has triggered the interest of the PM research community in using supervised or unsupervised Machine Learning (ML) techniques for gaining insight into business processes and providing advice on how to improve their inefficiencies.  \nIn today’s practice, ML models are routinely integrated into PM data pipelines [3] to carry out tasks like data transformation, noise reduction, anomaly detection, classification, and prediction. For example, ML is playing a key role in the interface between PM and sensor platforms. Advances in sensing technologies have made it possible to deploy distributed monitoring platforms capable of detecting fine-grained events. The granularity gap between these events and the activities considered by classic PM analysis has often been bridged using ML models [4, 5] that compute virtual activity logs, a problem which is also known as log lifting [6] . ML has been proposed as a key technology to strengthen existing techniques, for example, using trace clustering to reduce the diversity that a process discovery algorithm must handle in analyzing an event log [7, 8, 9, 10], to simplify the discovered models [11, 12, 13], or to  \nTailoring Machine Learning for Process Mining  \nsupport real-time analysis on event streams [14, 15, 16] . ML is adopted to apply predictive models to the executing cases of a process. This research area, known as predictive process monitoring, exploits event log data to foresee future events, remaining time, or the outcome of cases, in support of decision making [17, 18, 19] . Root cause analysis [20] and data explainability [21] are other tasks that can be applied to event log data using ML techniques, in order to improve our understanding of a business process. ML models have also been used in addition to (or in lieu of) classic linear programming [22] to optimize business processes’ resource consumption and to provide insights","cbCaikP7N8YENNRZ","https://ap.wps.com/l/cbCaikP7N8YENNRZ","pdf",867682,1,16,"English","en",105,"# Abstract\n# 1 Introduction\n## Process Mining and ML Integration\n## Assumptions, Encoding, and Concurrency Constraints","[{\"question\":\"Why is integrating machine learning into process mining pipelines not straightforward?\",\"answer\":\"Because mapping PM tasks to ML tasks requires assumptions about training functions and hyperparameters that must reflect the properties of process-specific data, including non-parametric distributions and concurrency constraints.\"},{\"question\":\"What kinds of tasks do machine learning models perform in process mining pipelines?\",\"answer\":\"They support data transformation, noise reduction, anomaly detection, classification, and prediction, including log lifting and predictive monitoring based on event logs.\"},{\"question\":\"How does data encoding affect machine learning when working with process data?\",\"answer\":\"Encoding determines compatibility with ML algorithms and influences sample complexity, the resulting data distribution, and feature relevance, impacting tasks like concept drift detection and zero-shot learning.\"}]","TAILORING MACHINE LEARNING FOR PROCESS MINING | PDF",1785672162,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"tailoring-machine-learning-for-process-mining","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/tailoring-machine-learning-for-process-mining/116871/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is integrating machine learning into process mining pipelines not straightforward?","Question",{"text":76,"@type":77},"Because mapping PM tasks to ML tasks requires assumptions about training functions and hyperparameters that must reflect the properties of process-specific data, including non-parametric distributions and concurrency constraints.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What kinds of tasks do machine learning models perform in process mining pipelines?",{"text":81,"@type":77},"They support data transformation, noise reduction, anomaly detection, classification, and prediction, including log lifting and predictive monitoring based on event logs.",{"name":83,"@type":74,"acceptedAnswer":84},"How does data encoding affect machine learning when working with process data?",{"text":85,"@type":77},"Encoding determines compatibility with ML algorithms and influences sample complexity, the resulting data distribution, and feature relevance, impacting tasks like concept drift detection and zero-shot learning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":29,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]