[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117738-en":3,"doc-seo-117738-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},117738,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Parallelization of Machine Learning Algorithms Respectively on Single Machine and Spark","Big data platforms have made data mining more essential, yet running machine learning on a traditional single machine can be slow and inefficient for large-scale inputs. This paper studies how to parallelize several classic machine learning algorithms on two targets: a single-machine environment and the Spark big data platform. By comparing runtime and efficiency between non-parallel and parallel implementations, the work demonstrates clear, measurable improvements from parallelization on both systems, especially for Spark-based iterative workloads.","arXiv :2206 .07090v1 [ cs .DC] 8 May 2022  \nParallelization of Machine Learning Algorithms Respectively on Single Machine and Spark  \nJiajun Shen  \nNorthwesten Polytechnical University  \n[lonelyinnovator@mail.nwpu.edu.cn](lonelyinnovator@mail.nwpu.edu.cn)  \nJune 16, 2022  \nAbstract  \nWith the rapid development of big data technologies, how to dig out useful information from massive data becomes an essential problem. However, using machine learning algorithms to analyze large data may be time-consuming and inefﬁcient on the traditional single machine. To solve these problems, this paper has made some research on the parallelization of several classic machine learning algorithms respectively on the single machine and the big data platform Spark. We compare the runtime and efﬁciency of traditional machine learning algorithms with parallelized machine learning algorithms respectively on the single machine and Spark platform. The research results have shown signiﬁcant improvement in runtime and efﬁciency of parallelized machine learning algorithms.  \nI. Introduction  \ni. Big Data Platform Spark  \nWe are now in the age of data explosion, everyday there will be tons of data generated, which contains a great deal of valuable information. So how to dig out useful information from massive data becomes a signiﬁcant problem. However, the traditional algorithms on single machine processing are insufﬁcient to process too much data either in computing capacity orin efﬁciency.  \nHadoop [1], a distributed big data framework emerged to solve these problems. The core design of Hadoop framework is HDFS and MapReduce [2] . HDFS provides storage for massive data, while MapReduce provides computing for massive data. Although Hadoop is efﬁcient to speed up processing large data through parallelized computing, the Mapreduce computing model is only applicable to the ofﬂine batch processing scenarios because  \nit can not calculate data fast and in real-time. Spark [3], the new big data framework, not only retains the scalability and fault tolerance of Hadoop MapReduce but also solves the issue of MapReduce. Spark is a fast data analysis framework based on the memory storage computing, it uses resilient distributed datasets (RDDs), which can provide interactive searching in real-time, store datasets in memory to improve reads and writes, reuse the datasets during computation and optimize iterative workloads. Therefore, Spark is more suitable for data mining and machine learning algorithms which have lots of iterations. This  \npaper uses Spark platform to do some research on the parallelization of machine learning algorithms.  \nii. Machine Learning Algorithms  \nMachine learning addresses the question of how to build computers that improve automatically through experience. It is one of today's  \nmost rapidly growing technical ﬁelds, being applied to computer vision, speech recognition, natural language processing, robot control, and other applications [4] .  \nMachine learning offers a wide range of statistical algorithms for analysis, mining, and prediction. It includes various techniques such as association rule mining, decision trees, regression, support vector machines, and other data mining techniques. All these algorithms are computationally expensive which makes them the ideal cases for implementation using parallel architecture/parallel programming methods [5] . Therefore, it is signiﬁcant to design the parallel methods for these machine learning algorithms, and this paper has done some research on the parallelization of these classic machine learning algorithms.  \niii. Algorithm Parallelization  \nIn a single machine, different threads can run in parallel in different cores when using multithreading on multi-core CPUs. Besides, parallel computing can also be performed on multiple GPUs.  \nIn the distributed environment, methods for parallelization can be classiﬁed into model parallelization, data parallelization, and hybrid parallelization. Model par","cbCaibySfHeUo1FY","https://ap.wps.com/l/cbCaibySfHeUo1FY","pdf",474370,1,"English","en",105,"# Introduction\n## Big Data Platform Spark\n## Machine Learning Algorithms\n## Algorithm Parallelization\n# Literature Review\n## SVM Algorithm Parallelization\n## K-Means Algorithm Parallelization\n## Neural Network Algorithm Parallelization","[{\"question\":\"Why do traditional single-machine machine learning methods struggle with big data?\",\"answer\":\"They are insufficient in computing capacity and/or efficiency for processing large volumes of data, leading to time-consuming analysis.\"},{\"question\":\"What advantages does Spark provide for parallel machine learning workloads?\",\"answer\":\"Spark retains scalability and fault tolerance from Hadoop MapReduce while enabling fast in-memory computation using RDDs, which suits iterative algorithms and real-time interactive searching.\"},{\"question\":\"How does the paper evaluate the benefit of parallelization?\",\"answer\":\"It compares runtime and efficiency between traditional machine learning algorithms and their parallelized versions on a single machine and on Spark.\"}]","Parallelization of Machine Learning Algorithms Respectively on Single Machine and Spark | PDF",1785679286,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"parallelization-of-machine-learning-algorithms-respectively-on-single-machine-and-spark","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/parallelization-of-machine-learning-algorithms-respectively-on-single-machine-and-spark/117738/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do traditional single-machine machine learning methods struggle with big data?","Question",{"text":74,"@type":75},"They are insufficient in computing capacity and/or efficiency for processing large volumes of data, leading to time-consuming analysis.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What advantages does Spark provide for parallel machine learning workloads?",{"text":79,"@type":75},"Spark retains scalability and fault tolerance from Hadoop MapReduce while enabling fast in-memory computation using RDDs, which suits iterative algorithms and real-time interactive searching.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the paper evaluate the benefit of parallelization?",{"text":83,"@type":75},"It compares runtime and efficiency between traditional machine learning algorithms and their parallelized versions on a single machine and on Spark.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]