[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116849-en":3,"doc-seo-116849-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116849,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Small-Scale Distributed Machine Learning in R - Master of Science Dissertation","Machine learning models often need large training datasets and substantial computation, making access to sufficient computing power challenging for students. Distributed machine learning mitigates this by dispatching tasks across a network of attached computers, enabling greater compute capacity than a single machine. This dissertation presents core distributed computing concepts and summarizes common parallel and distributed approaches in R. It focuses on the doRedis backend, demonstrating setup simplicity, elastic worker attachment/detachment, and measurable benefits for parallelizing training tasks such as random forests, hyper-parameter tuning, and cross-validation.","University of Cape Town  \nSmall-Scale Distributed Machine Learning  \nin R  \nStudent: Brenden Taylor TYLBRE007  \nSupervisor: Mr Stefan S Britz Co-supervisor: Dr Etienne Pienaar  \nA dissertation for the degree of Master of Science specialising  \nin Data Science  \nat the  \nDepartment of Statistical Sciences  \nThe copyright of this thesis vests in the author. No quotation from it or information derived from it is to be published without full acknowledgement of the source. The thesis is to be used for private study or noncommercial research purposes only.  \nPublished by the University of Cape Town (UCT) in terms of the non-exclusive license granted to UCT by the author.  \nAbstract  \nMachine learning is increasing in popularity, both in applied and theoretical statistical fields. Machine learning models generally require large amounts of data to train and thus are computationally expensive, both in the absolute sense of actual compute time, and in the relative sense of the numerical complexity of the underlying calculations. Particularly for students of machine learning, appropriate computing power can be difficult to come by. Distributed machine learning, which involves sending tasks to a network of attached computers, can offer users access to significantly more computing power than otherwise by leveraging more processors than in a single computer.  \nThis research outlines the core concepts of distributed computing and provides brief outlines of the more common approaches to parallel and distributed computing in R, with reference to the specific algorithms and aspects of machine learning that are investigated. One particular parallel backend, doRedis, offers particular advantages as it is easy to set up and implement, and allows for the elastic attaching and detaching of computers from a distributed network. This paper will describe core features of the doRedis package and show, by means of applying certain aspects of the machine learning process, that it is both viable and beneficial to distribute these machine learning aspects.  \nThere is the potential for significant time savings when distributing machine learning model training. Particularly for students, the time required for setting up of a distributed network in which to use doRedis is far outweighed by the benefits. The implication that this research aims to explore, is that students will be able to leverage the many computers often available in computer labs to train more complex machine learning models in less time than they would otherwise be able to when using the built-in parallel packages that are already common in R. In fact, certain machine learning packages that already parallelise model training can be distributed to a network of computers, thereby further increasing the gains realised by parallelisation. In this way, more complex machine learning is more accessible.  \nThis research outlines the benefits that lie in the distribution of machine learning problems in an accessible, small-scale environment. This small-scale ‘proof of concept’ performs well enough to be viable for students, while also creating a bridge, and introducing the knowledge required, to deploy large-scale distribution of machine learning problems.  \nContents  \n1 Introduction 3  \n1.1 Background ..................................... 3  \n1.2 Scope and Outline of the Research ........................ 4  \n1.3 Relevance ...................................... 5  \n2 Parallel and Distributed Computing 7  \n2.1 Parallel Computing ................................. 7  \n2.2 Distributed Computing .............................. 8  \n2.3 Parallel and Distributed Computing Concepts .................. 9  \n2.3.1 Multicore and Cluster Computing .................... 9  \n2.3.2 Parallelisable Algorithms ......................... 9  \n2.4 Parallel and Distributed Computing in R .................... 10  \n2.4.1 parallel .................................. 11  \n2.4.2 future ................................... 11  \n2.4.3 Do","cbCaioMRx657YjlN","https://ap.wps.com/l/cbCaioMRx657YjlN","pdf",1526859,1,42,"English","en",105,"# Introduction\n## Background\n## Scope and Outline of the Research\n## Relevance\n# Parallel and Distributed Computing\n## Parallel Computing\n## Distributed Computing\n## Parallel and Distributed Computing Concepts\n## Parallel and Distributed Computing in R\n# Distributed Machine Learning\n## Random Forests\n## Hyper-parameter Tuning\n## Cross Validation\n# doRedis as a foreach Parallel Backend\n## Terminology\n## Functions, Specifications and Options\n## Considerations\n# Benchmark of the Time Improvement in ML Training when using doRedis\n## Three Parallel Backends\n## Random Forests and Hyper-parameter Tuning\n## Cross Validation\n## Elasticity of doRedis Demonstration\n# Conclusion\n## Future Work and Recommendations\n# Appendices\n## Hardware\n## Using doRedis","[{\"question\":\"Why does the dissertation focus on distributed machine learning for students?\",\"answer\":\"Model training can be computationally expensive and requires substantial computing power. Distributed machine learning enables students to leverage multiple computers to reduce training time and make more complex models practical.\"},{\"question\":\"What is doRedis and why is it highlighted in the research?\",\"answer\":\"doRedis is a parallel backend in R that is easy to set up and supports elastic attaching and detaching of computers from a distributed network.\"},{\"question\":\"Which parts of the machine learning workflow are evaluated using doRedis?\",\"answer\":\"The dissertation applies distributed approaches to key steps including random forests, hyper-parameter tuning, and cross-validation, and benchmarks time improvement against other parallel backends.\"}]","Small-Scale Distributed Machine Learning in R - Master of Science Dissertation | PDF",1785672062,106,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"small-scale-distributed-machine-learning-in-r-master-of-science-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/small-scale-distributed-machine-learning-in-r-master-of-science-dissertation/116849/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does the dissertation focus on distributed machine learning for students?","Question",{"text":75,"@type":76},"Model training can be computationally expensive and requires substantial computing power. Distributed machine learning enables students to leverage multiple computers to reduce training time and make more complex models practical.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is doRedis and why is it highlighted in the research?",{"text":80,"@type":76},"doRedis is a parallel backend in R that is easy to set up and supports elastic attaching and detaching of computers from a distributed network.",{"name":82,"@type":73,"acceptedAnswer":83},"Which parts of the machine learning workflow are evaluated using doRedis?",{"text":84,"@type":76},"The dissertation applies distributed approaches to key steps including random forests, hyper-parameter tuning, and cross-validation, and benchmarks time improvement against other parallel backends.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]