[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123025-en":3,"doc-seo-123025-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123025,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","From Micro-benchmarks to Machine Learning - Unveiling the Efficiency and Scalability of Hadoop and Spark","With the exponential growth of data, the demand for efficient and scalable data processing has become critical. The study evaluates Hadoop and Spark, key open-source Big Data components, through comprehensive performance analysis in virtualized environments using a suite of benchmarks. Workloads range from micro-benchmarks like Sort, WordCount, and TeraSort to web-search tasks such as PageRank and machine-learning methods including Naive Bayes and K-means. Results highlight Spark’s in-memory processing advantages across scenarios, while noting Hadoop’s potential edge under resource constraints and with large inputs.","JIM International Journal of  \nInteractive Mobile Technologies  \n[Onli](Online-Journals.org)[ne-Jo](Online-Journals.org)[urnals](Online-Journals.org)[.org](Online-Journals.org)  \niJIM | eISSN: 1865-7923 | Vol. 18 No. 17 (2024) |   \n[https://doi.org/10.3991/ijim.v18i17.44555](https://doi.org/10.3991/ijim.v18i17.44555)  \nPAPER  \nFrom Micro-benchmarks to Machine Learning: Unveiling the Efficiency and Scalability of Hadoop and Spark  \nSalah Eddine Hebabaze1(􀀍), Mohamed El Ghmary2, Hamid El Bouabidi1, Sara Maftah1, Mohamed Amnai1  \n1Ibn Tofaïl University, Kenitra, Morocco  \n2Sidi Mohamed Ben Abdellah University, Fez, Morocco  \nSalaheddine.Hebabaze@ [uit.ac.ma](uit.ac.ma)  \nABSTRACT  \nWith the exponential growth of data, the demand for efficient and scalable data processing solutions has become paramount. Hadoop and Spark, pivotal components of the open-source Big Data landscape, have been put to the test in this study. We conducted a comprehensive performance analysis of Hadoop and Spark in virtualized environments, evaluating their prowess across a suite of benchmarks. The benchmarks encompassed a spectrum of workloads, from micro-benchmarks such as Sort, WordCount, and TeraSort to web search tasks such as PageRank and machine learning endeavors including Naive Bayes and K-means. The central focus was to gauge their performance, efficiency, and resource utilization. The findings ofthis study underscore the benefits of Spark’s in-memory processing, demonstrating its superiority over Hadoop in various scenarios. Spark excels in machine learning and web search applications, particularly when handling smaller inputs. Its efficient memory management and support for multiple iterations make it a strong choice. In resource-constrained environments or when dealing with large input files and limited memory, Hadoop may still hold an edge. The design and implementation of data processing solutions in virtualized environments should carefully consider the specific demands of each framework. This study not only presents a performance comparison of Hadoop and Spark across different benchmarks but also emphasizes the vital implications for designing and deploying data processing solutions in virtualized settings. It serves as a cornerstone for informed decision-making, paving the way for optimized algorithms and techniques in the dynamic landscape of big data processing.  \nKEYWORDS  \nbig data, Hadoop, Apache Spark, MapReduce, HiBench benchmark, machine learning, memory resource limitations, data workloads  \n1 INTRODUCTION  \nNowadays, traditional data management systems can’t process the data’s complexity because of its size, structure, and limited processing time [1] . Massive and  \nHebabaze, S. E., El Ghmary, M., El Bouabidi, H., Maftah, S., Amnai, M. (2024) . From Micro-benchmarks to Machine Learning: Unveiling the Efficiency and Scalability of Hadoop and Spark. InternationalJournal of Interactive Mobile Technologies (iJIM), 18(17), pp. 46–60. [https://doi.org/10.3991/ijim.v18i17](https://doi.org/10.3991/ijim.v18i17) . 44555  \nArticle submitted 2023-09-08. Revision uploaded 2024-06-24. Final acceptance 2024-06-29.  \n© 2024 by the authors of this article. Published under CC-BY.  \n46 International Journal of Interactive Mobile Technologies (iJIM) iJIM | Vol. 18 No. 17 (2024)  \nFrom Micro-benchmarks to Machine Learning: Unveiling the Efficiency and Scalability of Hadoop and Spark  \ncomplex data, known as “Big Data,” requires new approaches such as Hadoop and Spark. Hadoop is batch-oriented, while Spark is in-memory and real-time, making it faster and more versatile [2] . This paper compares Hadoop and Spark, two popular data processing frameworks, highlighting their strengths and weaknesses. It emphasizes the importance of selecting the right framework for specific tasks. The paper also contributes to the ongoing discussion among professionals and organizations about improving web-based mobile learning in educational contexts by providing insight","cbCainhgPSMnmti4","https://ap.wps.com/l/cbCainhgPSMnmti4","pdf",386071,1,15,"English","en",105,"# Introduction\n# Related Works\n# Performance Evaluation Methodology\n## Benchmarks and Workloads\n# Results and Discussion\n## Efficiency and Scalability Findings\n# Conclusion","[{\"question\":\"What benchmarks and workloads are used to compare Hadoop and Spark?\",\"answer\":\"The study uses micro-benchmarks such as Sort, WordCount, and TeraSort, along with web search tasks like PageRank. It also evaluates machine learning workloads including Naive Bayes and K-means.\"},{\"question\":\"Why does Spark generally outperform Hadoop in the reported scenarios?\",\"answer\":\"Spark’s in-memory processing is emphasized as a key advantage, supported by findings that show better efficiency and performance across multiple workloads.\"},{\"question\":\"When may Hadoop still be the better choice according to the paper?\",\"answer\":\"The paper suggests Hadoop can hold an edge in resource-constrained environments or when dealing with large input files under limited memory conditions.\"}]","From Micro-benchmarks to Machine Learning - Unveiling the Efficiency and Scalability of Hadoop and Spark | PDF",1785814231,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-micro-benchmarks-to-machine-learning-unveiling-the-efficiency-and-scalability-of-hadoop-and-spark","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-micro-benchmarks-to-machine-learning-unveiling-the-efficiency-and-scalability-of-hadoop-and-spark/123025/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What benchmarks and workloads are used to compare Hadoop and Spark?","Question",{"text":75,"@type":76},"The study uses micro-benchmarks such as Sort, WordCount, and TeraSort, along with web search tasks like PageRank. It also evaluates machine learning workloads including Naive Bayes and K-means.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does Spark generally outperform Hadoop in the reported scenarios?",{"text":80,"@type":76},"Spark’s in-memory processing is emphasized as a key advantage, supported by findings that show better efficiency and performance across multiple workloads.",{"name":82,"@type":73,"acceptedAnswer":83},"When may Hadoop still be the better choice according to the paper?",{"text":84,"@type":76},"The paper suggests Hadoop can hold an edge in resource-constrained environments or when dealing with large input files under limited memory conditions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]