[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120482-en":3,"doc-seo-120482-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120482,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Machine learning in big data - A performance benchmarking study of Flink-ML and Spark MLlib","Machine learning in big data frameworks underpins real-time analytics, decision making, and predictive modeling. This study compares Flink-ML (Apache Flink’s ML extension) and Spark MLlib (Apache Spark’s ML library) across performance, scalability, streaming behavior, iterative computation efficiency, and integration with external deep learning tools. Flink-ML targets event-driven, low-latency workflows with streaming-based training and inference, while MLlib emphasizes batch and micro-batch streaming. Results indicate similar training times, closely matched memory use, and slightly higher accuracy for Flink-ML, with marginal throughput advantages for Spark MLlib during inference.","Submitted: 2025-02-15 | Revised: 2025-04-04 | Accepted: 2025-04-13  \nCC-BY 4.0  \nKeywords: machine learning, apache flink, apache spark, Flink-ML, MLlib  \nMessaoud MEZATI 1 , Ines AOURIA 1*  \n1 Kasdi Merbah University, Algeria, [mezati.messaoud@univ-ouargla.dz](mezati.messaoud@univ-ouargla.dz), [aouria.ines@univ-ouargla.dz](aouria.ines@univ-ouargla.dz)[ ](aouria.ines@univ-ouargla.dz)* Corresponding author: [aouria.ines@univ-ouargla.dz](aouria.ines@univ-ouargla.dz)  \nMachine learning in big data:  \nA performance benchmarking study  \nofFlink-ML and Spark MLlib  \nAbstract  \nMachine learning (ML) in big data frameworks plays a critical role in real-time analytics, decision making, and predictive modeling. Among the most prominent ML libraries for large-scale data processing are Flink-ML, the machine learning extension of Apache Flink, and MLlib, the machine learning library of Apache Spark. This paper provides a comparative analysis of these two frameworks, evaluating their performance, scalability, streaming capabilities, iterative computation efficiency, and ease of integration with external deep learning frameworks. Flink-ML is designed for real-time, event-driven ML applications and provides native support for streaming-based model training and inference. In contrast, Spark MLlib is optimized for batch processing and micro-batch streaming, making it more suitable for traditional machine learning workflows. Experimental results show that training time is nearly identical for both frameworks, with Spark MLlib requiring 4006.4 seconds and Flink-ML 4003.2 seconds, demonstrating comparable efficiency in batch training and streaming-based model updates. Accuracy results show that Flink-ML (74.9%) slightly outperforms Spark MLlib (74. 7%), suggesting that continuous learning in Flink-ML may contribute to better generalization. Inference throughput is slightly higher for Spark MLlib (8.4 images/sec) compared to Flink-ML (8.2 images/sec), suggesting that Spark's batch execution provides a slight advantage in processing efficiency. Both frameworks consume the same amount of memory (30.2%), confirming that TensorFlow's deep learning operations dominate resource consumption rather than architectural differences between Spark and Flink. These results highlight the tradeoffs between Flink-MLand Spark MLlib, and guide data scientists and engineers in selecting the appropriate framework based on specific ML workflow requirements and scalability considerations.  \n1. INTRODUCTION  \nMachine learning (ML) has become a fundamental component of big data analytics, enabling predictive modeling, real-time decision making, and large-scale data processing. As industries increasingly rely on ML Gao et al. (2024) For applications such as fraud detection, recommendation systems, and predictive maintenance, the choice of an appropriate ML framework becomes critical. Apache Spark and Apache Flink are two of the most widely adopted big data processing frameworks, each offering different capabilities for machine learning tasks (Khalid & Yousaf, 2021) .  \nApache Spark's MLlib is a mature and well-established ML library optimized for batch processing and micro-batch streaming, making it highly effective for traditional ML workloads. It provides a rich set of prebuilt algorithms and deep learning integrations that facilitate scalable ML model training (Zeydan & ManguesBafalluy, 2022) . On the other hand, Flink-ML, Apache Flink's machine learning extension, is designed for real-time, event-driven ML applications and provides native support for streaming-based model training and inference (Dritsas & Trigka, 2025) . While Spark MLlib dominates in batch ML workloads, Flink-ML's lowlatency processing capabilities position it as an alternative for dynamic and adaptive learning scenarios.  \nDespite the growing adoption of both frameworks, there is limited comparative research evaluating their suitability for different machine learning paradigms. Existing studies often ","cbCaibV93OtjJ28k","https://ap.wps.com/l/cbCaibV93OtjJ28k","pdf",749565,1,10,"English","en",105,"# Introduction\n## Purpose and contributions\n# Architectural differences between Flink-ML and MLlib\n## Overview of Apache Flink and Apache Spark architectures\n# Experimental setup and benchmarking methodology\n# Performance evaluation results\n# Discussion and practical recommendations\n# Conclusion and future research directions","[{\"question\":\"What comparison dimensions does the study use for Flink-ML and Spark MLlib?\",\"answer\":\"It compares performance, scalability, streaming capabilities, iterative computation efficiency, and integration with external deep learning frameworks.\"},{\"question\":\"How do training time and accuracy results differ between the two frameworks?\",\"answer\":\"Training time is nearly identical, with Spark MLlib at 4006.4 seconds and Flink-ML at 4003.2 seconds. Flink-ML achieves slightly higher accuracy (74.9%) than Spark MLlib (74.7%).\"},{\"question\":\"Which framework shows an advantage in inference throughput and why?\",\"answer\":\"Spark MLlib shows slightly higher inference throughput (8.4 images/sec) versus Flink-ML (8.2 images/sec). The study attributes the small edge to Spark’s batch execution efficiency.\"}]","Machine learning in big data - A performance benchmarking study of Flink-ML and Spark MLlib | PDF",1785730309,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-in-big-data-a-performance-benchmarking-study-of-flink-ml-and-spark-mllib","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-in-big-data-a-performance-benchmarking-study-of-flink-ml-and-spark-mllib/120482/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What comparison dimensions does the study use for Flink-ML and Spark MLlib?","Question",{"text":75,"@type":76},"It compares performance, scalability, streaming capabilities, iterative computation efficiency, and integration with external deep learning frameworks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do training time and accuracy results differ between the two frameworks?",{"text":80,"@type":76},"Training time is nearly identical, with Spark MLlib at 4006.4 seconds and Flink-ML at 4003.2 seconds. Flink-ML achieves slightly higher accuracy (74.9%) than Spark MLlib (74.7%).",{"name":82,"@type":73,"acceptedAnswer":83},"Which framework shows an advantage in inference throughput and why?",{"text":84,"@type":76},"Spark MLlib shows slightly higher inference throughput (8.4 images/sec) versus Flink-ML (8.2 images/sec). The study attributes the small edge to Spark’s batch execution efficiency.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]