[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125925-en":3,"doc-seo-125925-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125925,2336474459895,"Aria","https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916",8,"Research & Report","A novel framework for generic Spark workload characterization and similar pattern recognition using machine learning - Article","Comprehensive workload characterization is essential for understanding Spark applications and enabling downstream objectives such as performance improvement. This work proposes a novel, scalable framework for generic Spark workload characterization using consistent geometric measurements while profiling only quantitative metrics at the application task level in a non-intrusive way. It extends the framework with unsupervised machine learning, including clustering and feature selection, to recognize similar workloads without predefined labels. The approach identifies 24 representative workloads across multiple domains, achieving up to 90.9% F-Measure and up to 94.5% normalized mutual information, outperforming prior literature methods.","Journal of Parallel and Distributed Computing 189 (2024) 104881  \nContents lists available at ScienceDirect  \nJournal of Parallel and Distributed Computing  \njournal [homepage:](homepage: www.elsevier.com/locate/jpdc)[ www.elsevier.com/locate/jpdc](homepage: www.elsevier.com/locate/jpdc)  \n| A novel framework for generic Spark workload characterization and similar pattern recognition using machine learning\u003Cbr>Mariano Garralda-Barrio ∗ , Carlos Eiras-Franco, Verónica Bolón-Canedo\u003Cbr>CITIC, Universidade da Coruña, A Coruña, Spain |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O | A B S T R A C T\u003Cbr>Comprehensive workload characterization plays a pivotal role in comprehending Spark applications, as it enables the analysis of diverse aspects and behaviors. This understanding is indispensable for devising downstream tuning objectives, such as performance improvement. To address this pivotal issue, our work introduces a novel and scalable framework for generic Spark workload characterization, complemented by consistent geometric measurements. The presented approach aims to build robust workload descriptors by proﬁling only quantitative metrics at the application task-level, in a non-intrusive manner. We expand our framework for downstream workload pattern recognition by incorporating unsupervised machine learning techniques: clustering algorithmsand feature selection. These techniques signiﬁcantly improve the process of grouping similar workloads without relying on predeﬁned labels. We eﬀectively recognize 24 representative Spark workloads from diverse domains, including SQL, machine learning, web search, graph, and micro-benchmarks, available in HiBench. Our framework achieves a high accuracy F-Measure score of up to 90.9% and a Normalized Mutual Information of up to 94.5% in similar workload pattern recognition. These scores signiﬁcantly outperform the results obtained in a comparative analysis with an established workload characterization approach in the literature. |  |\n| Dataset link: [https://](https://)[ ](https://)[github.com/mgarralda/hadoop-spark-cluster/](github.com/mgarralda/hadoop-spark-cluster/)[ ](github.com/mgarralda/hadoop-spark-cluster/)[tree/main/spark-event-logs](tree/main/spark-event-logs) |  |  |\n| Keywords:\u003Cbr>Big data\u003Cbr>Workload characterization Apache spark\u003Cbr>Pattern recognition Machine learning |  |  |\n\n1. Introduction  \nWith the volume of data increasing exponentially, the era of big data has emerged as one of the most signiﬁcant trends in high-performance computing. To extract valuable insights, big data workloads demand specialized environments on large-scale computing infrastructures. To address this requirement, Apache Spark™ [1] is a widely-used framework that provides a uniﬁed multi-language engine for computing heterogeneous workloads, such as data engineering, data analytics, and machine learning applications. Comprehensive workload characterization is essential for understanding Spark applications. This, in turn, facilitates the development of downstream objectives, such as performance prediction models [11], workload prediction [25], and autotuning of big data applications [33], allowing for proactive optimization of resource allocation. However, the distributed nature of Spark infrastructure (􀀂􀀃), large datasets (􀀄􀀅), diverse application characteristics (􀀆􀀇􀀇), and numerous conﬁguration settings (􀀈􀀅) present signiﬁcant challenges in characterizing Spark workloads.  \nAs a deﬁnition, workload characterization [24] consists of a description of the workload by means of several quantitative metrics, such as at  \nmicro-architecture-level, system-level, and application-level. Nevertheless, most ongoing eﬀorts in workload characterization tend to focus on system-level properties, which include machine-speciﬁc details, rather than pure workload characterization. This would involve understanding the behavior and patterns of a workload without making any assumptions about the underlying system or hardw","cbCaibkvSJ6Bug4o","https://ap.wps.com/l/cbCaibkvSJ6Bug4o","pdf",2821493,4,1,16,"English","en",105,"# Introduction\n## Workload characterization and its importance\n## Metrics: static vs dynamic\n## Challenges in characterizing Spark workloads","[{\"question\":\"What is the main goal of the proposed framework for Spark workloads?\",\"answer\":\"The framework aims to generically characterize Spark workloads and recognize similar workload patterns to support downstream performance-related objectives.\"},{\"question\":\"How does the framework build workload descriptors without intrusive instrumentation?\",\"answer\":\"It profiles only quantitative metrics at the application task level using a non-intrusive approach and employs consistent geometric measurements to form robust descriptors.\"},{\"question\":\"Which machine learning methods are used for grouping similar workloads?\",\"answer\":\"The framework uses unsupervised machine learning, specifically clustering algorithms and feature selection, to group workloads without relying on predefined labels.\"}]","A novel framework for generic Spark workload characterization and similar pattern recognition using machine learning - Article | PDF",1785902062,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"a-novel-framework-for-generic-spark-workload-characterization-and-similar-pattern-recognition-using-machine-learning-article","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/a-novel-framework-for-generic-spark-workload-characterization-and-similar-pattern-recognition-using-machine-learning-article/125925/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main goal of the proposed framework for Spark workloads?","Question",{"text":76,"@type":77},"The framework aims to generically characterize Spark workloads and recognize similar workload patterns to support downstream performance-related objectives.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the framework build workload descriptors without intrusive instrumentation?",{"text":81,"@type":77},"It profiles only quantitative metrics at the application task level using a non-intrusive approach and employs consistent geometric measurements to form robust descriptors.",{"name":83,"@type":74,"acceptedAnswer":84},"Which machine learning methods are used for grouping similar workloads?",{"text":85,"@type":77},"The framework uses unsupervised machine learning, specifically clustering algorithms and feature selection, to group workloads without relying on predefined labels.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]