[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117376-en":3,"doc-seo-117376-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117376,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Enriching the Machine Learning Workloads in BigBench - Technical Report","BigBench (TPCx-BB) provides an end-to-end, business-oriented benchmarking foundation for big data systems, yet its machine-learning coverage is limited and depends on few library implementations. This technical report enriches BigBench V2 by adding three new machine-learning workloads and expanding algorithm coverage. The workloads combine multiple algorithms and compare implementations across MLlib, SystemML, Scikit-learn, and Pandas, demonstrating the extension’s relevance for evaluating usability and performance of popular ML libraries within benchmark workflows.","arXiv :2406 . 10843v1 [ cs .LG] 16 Jun 2024  \nEnriching the Machine Learning Workloads in  \nBigBench  \nTechnical Report, May 2024  \nMatthias Polag, Todor Ivanov, and Timo Eichhorn  \nFrankfurt Big Data Lab, Goethe University Frankfurt, Frankfurt, Hessen, Germany matthias,todor,[timo@dbis.cs.uni-frankfurt.de](timo@dbis.cs.uni-frankfurt.de)  \nAbstract. In the era of Big Data and the growing support for Machine Learning, Deep Learning and Artificial Intelligence algorithms in the current software systems, there is an urgent need of standardized application benchmarks that stress test and evaluate these new technologies. Relying on the standardized BigBench (TPCx-BB) benchmark, this work enriches the improved BigBench V2 with three new workloads and expands the coverage of machine learning algorithms. Our workloads utilize multiple algorithms and compare different implementations for the same algorithm across several popular libraries like MLlib, SystemML, Scikit-learn and Pandas, demonstrating the relevance and usability of our benchmark extension.  \nKeywords: Benchmarking · Big Data · Machine Learning · BigBench  \n1 Introduction  \nIn current times Big Data is often seen as the next big thing in innovation [10] and has the potential to generate significant growth for the economy as a whole [18] . For example, the article Smart analytics: How marketing drives short-term and long-term growth [22] claims that big data analysis can increase the return of investment in retailer marketing by 15 to 20 percent.  \nToday there exist several frameworks designed to process big data like Hadoop [27] and Spark [28], yet it can be challenging to decide which system best meets their needs. To help with that decision end-to-end application benchmarks like BigBench (TPCx-BB) [8,4] were created. They evaluate different Big Data systems on a set of tasks and measure their capabilities. While BigBench and BigBench V2 [7] currently cover many common business tasks, they have only few examples of machine learning tasks as part of their workload and solely rely on the Mahout and partially MLlib libraries to implement the necessary algorithms. As suggested by Singh [25] and briefly described in the vision of ABench [15], we enrich BigBench with new machine learning algorithms and workloads, which then help us to evaluate and compare various popular machine learning libraries.  \nIn this paper we provide an extension of BigBench V2 by adding additional machine learning workloads and algorithms using different libraries. The main contributions of this work are:  \n2 M. Polag et al.  \n– Enriching BigBench V2 with more machine learning algorithms.  \n– Implementing the workloads in several different libraries (MLlib, SystemML, Scikit-learn and Pandas) .  \n– Comparison and evaluation of the libraries.  \nThe paper is structured as follows: Section 2 gives an overview of related work. Section 3 starts with a brief overview of existing Machine Learning algorithmsand libraries and introduces the BigBench Machine Learning workloads and their implementations. Section 4 describes the experimental setup, methodology and execution, and presents the results of the experimental evaluation. Finally, Section 5 summarizes the lessons learned.  \n2 Related Work  \nMachine Learning (ML) benchmarks are mainly realized as micro benchmarks. They focus on measuring the performance of specific components or tasks.  \nDeepBench [24] measures the performance of running basic neural network operations regarding the used hardware and covers essential algorithms like General Matrix Multiply (GEMM), Convolution, and Recurrent Neural Network (RNN) . The results are adequate indicators of the execution time necessary to train an entire model.  \nDawnBench [2] uses image classification and question answering workloads to evaluate deep learning systems. It proposes a new metric, using time per accuracy as an indicator. The metric reports both the end-to-end training time to achieve a state-of-the-","cbCaipwVXuSyjzno","https://ap.wps.com/l/cbCaipwVXuSyjzno","pdf",226072,1,14,"English","en",105,"# Introduction\n## Related Work\n## Enriching the Machine Learning Workloads in BigBench\n# Experimental Setup and Evaluation\n## Lessons Learned","[{\"question\":\"What gap does the report address in existing BigBench workloads?\",\"answer\":\"BigBench V2 includes only a few machine-learning tasks and relies mainly on Mahout and partially MLlib, leaving limited coverage of ML algorithms and library implementations.\"},{\"question\":\"How does the extension enrich BigBench V2?\",\"answer\":\"It adds three new machine-learning workloads and expands the range of machine-learning algorithms included in the benchmark.\"},{\"question\":\"Which machine-learning libraries are compared in the new workloads?\",\"answer\":\"The report uses and compares multiple implementations across MLlib, SystemML, Scikit-learn, and Pandas for the same or related algorithms.\"}]","Enriching the Machine Learning Workloads in BigBench - Technical Report | PDF",1785675454,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"enriching-the-machine-learning-workloads-in-bigbench-technical-report","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/enriching-the-machine-learning-workloads-in-bigbench-technical-report/117376/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What gap does the report address in existing BigBench workloads?","Question",{"text":75,"@type":76},"BigBench V2 includes only a few machine-learning tasks and relies mainly on Mahout and partially MLlib, leaving limited coverage of ML algorithms and library implementations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the extension enrich BigBench V2?",{"text":80,"@type":76},"It adds three new machine-learning workloads and expands the range of machine-learning algorithms included in the benchmark.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine-learning libraries are compared in the new workloads?",{"text":84,"@type":76},"The report uses and compares multiple implementations across MLlib, SystemML, Scikit-learn, and Pandas for the same or related algorithms.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]