[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118699-en":3,"doc-seo-118699-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},118699,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",6,"Technology","Tools and Frameworks for Big Learning in Scala - Leveraging the Language for High Productivity and Performance","Implementing machine learning for large-scale data such as web graphs and social networks is hindered by prohibitively long sequential runtimes and by the difficulty of parallelization, even when using frameworks like MapReduce that mask complexity. The work presents three ongoing efforts to help researchers and practitioners quickly implement and experiment with parallel or distributed learning. It also emphasizes Scala-specific language features to build efficient and correct parallel systems more easily.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nTools and Frameworks for Big Learning in Scala: Leveraging the Language for High Productivity and Performance  \nHeather Miller  \nEPFL, Switzerland [heather.miller@epfl.ch](heather.miller@epfl.ch)  \nPhilipp Haller  \nEPFL, Switzerland and Stanford University [philipp.haller@epfl.ch](philipp.haller@epfl.ch)  \nMartin Odersky  \nEPFL, Switzerland [martin.odersky@epfl.ch](martin.odersky@epfl.ch)  \nAbstract  \nImplementing machine learning algorithms for large data, such as the Web graph and social networks, is challenging. Even though much research has focused on making sequential algorithms more scalable, their running times continue to be prohibitively long.  \nMeanwhile, parallelization remains a formidable challenge for this class of problems, despite frameworks like MapReduce which hide much of the associated complexity.  \nWe present three ongoing efforts within our team, previously presented at venues in other ﬁelds, which aim to make it easier for machine learning researchers and practitioners alike to quickly implement and experiment with their algorithms in a parallel or distributed setting. Furthermore, we hope to highlight some of the language features unique to the Scala programming language in the treatment of our frameworks, in an effort to show how these features can be used to produce efﬁcient and correct parallel systems more easily than ever before.  \n1 Introduction  \nThere is a growing need to facilitate the analysis of large data. Machine learning (ML) has provided elegant and sophisticated solutions to many complex problems on a small scale, which if ported to large scale problems, could open up new applications and avenues of research for numerous ﬁelds. Unfortunately, ML research efforts are routinely limited by the complexity and running time of sequential algorithms. Despite their success in other contexts, parallel programming paradigms like MapReduce/Hadoop are, in many ways, too limited to express the charicteristics of problems which frequently arise in ﬁelds like ML. Inspired by these limitations, we present several projects which, in different ways, aim to make it easier for ML researchers and practitioners to implement and experiment with their algorithms in a parallel or distributed setting.  \nThe goals of this paper are two-fold; we aim to showcase some of our tools, libraries, and frameworks suited to large-scale parallel and distributed learning, meanwhile highlighting language features that the Scala programming language integrates in a unique way, and which we argue make using and developing parallel systems much more easily achievable.  \nWe ﬁrst present three projects suited for large-scale ML. Two of which have been recently presented at non-ML oriented venues: Menthor [1], a framework for distributed graph-processing, and Scala's Parallel Collections [2], a library which provides parallel bulk operations on generic collection types, such as Arrays, Maps, and Sets. And ﬁnally, OptiML [3], a parallel domain-speciﬁc language (DSL) for machine learning on heterogeneous hardware platforms 1.  \n1As a member of the Scala team at EPFL, Heather has worked on Scala parallel collections, and is also one of the primary authors of the Menthor ML framework. As a member of Stanford's PPL, Philipp works on debugging tools for the Delite DSL framework, as a member of the Scala team at EPFL, he works on the Scala language and (concurrency) libraries.  \nFinally, we hope to show how each of these projects rely on a unique combination of interesting language features, which in many respects are responsible for making their implementations possible, and we note the existence of other related projects like Spark [4] and FACTORIE [5] who have also chosen Scala for its language features.  \n2 Our Tools, Libraries, and Framewor","cbCaib1K9TYVD9a7","https://ap.wps.com/l/cbCaib1K9TYVD9a7","pdf",102863,1,"English","en",105,"# Introduction\n## Goals of the paper\n# Our Tools, Libraries, and Frameworks for Big Learning\n## Menthor\n## Parallel Collections\n## OptiML","[{\"question\":\"Why is large-scale machine learning difficult to implement in practice?\",\"answer\":\"Sequential algorithms often run too slowly on large data, and parallelization remains challenging. Even MapReduce-style frameworks can be too limited for many machine learning problem characteristics.\"},{\"question\":\"What is the main contribution of the presented efforts?\",\"answer\":\"The paper presents three parallel/distributed frameworks and libraries to speed up implementation and experimentation for machine learning researchers and practitioners.\"},{\"question\":\"How does Scala’s language design relate to the proposed frameworks?\",\"answer\":\"The efforts highlight Scala language features used to make efficient and correct parallel systems easier to develop than before.\"}]","Tools and Frameworks for Big Learning in Scala - Leveraging the Language for High Productivity and Performance | PDF",1785684953,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"tools-and-frameworks-for-big-learning-in-scala-leveraging-the-language-for-high-productivity-and-performance","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/technology/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/tools-and-frameworks-for-big-learning-in-scala-leveraging-the-language-for-high-productivity-and-performance/118699/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is large-scale machine learning difficult to implement in practice?","Question",{"text":74,"@type":75},"Sequential algorithms often run too slowly on large data, and parallelization remains challenging. Even MapReduce-style frameworks can be too limited for many machine learning problem characteristics.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What is the main contribution of the presented efforts?",{"text":79,"@type":75},"The paper presents three parallel/distributed frameworks and libraries to speed up implementation and experimentation for machine learning researchers and practitioners.",{"name":81,"@type":72,"acceptedAnswer":82},"How does Scala’s language design relate to the proposed frameworks?",{"text":83,"@type":75},"The efforts highlight Scala language features used to make efficient and correct parallel systems easier to develop than before.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,112,117,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":110,"slug":111},50,"technology",{"id":113,"doc_module":4,"doc_module_name":45,"category_name":114,"show_sort_weight":115,"slug":116},7,"Healthcare",40,"healthcare",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":120,"slug":121},8,"Research & Report",30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]