[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120654-en":3,"doc-seo-120654-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120654,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Block size estimation for data partitioning in HPC applications using machine learning techniques","Extensive use of HPC infrastructures for data-intensive workloads increases the need for effective data partitioning strategies, where choosing a suitable block size is central to improving parallel execution speed and overall scalability. This paper presents a supervised machine learning methodology to estimate data block sizes for HPC applications. Evaluation on the dislib testbed and multiple algorithms, datasets, and infrastructures—including MareNostrum 4—shows the approach can accurately split datasets, enabling efficient data-parallel execution in high-performance environments.","Block size estimation for data partitioning in HPC applications using machine learning techniques  \nRiccardo Cantini􀀃 , Fabrizio Marozzo􀀃 , Alessio Orsino􀀃 , Domenico Talia􀀃 , Paolo Trunﬁo􀀃 , Rosa M. Badiay , Jorge Ejarquey , Fernando Vazquezy  \n􀀃 DIMES, University of Calabria, Rende, Italy, e-mail: frcantini, fmarozzo, aorsino, talia, trunﬁ[o](og@dimes.unical.it)[g](og@dimes.unical.it)[@dimes.unical.it](og@dimes.unical.it)[ ](og@dimes.unical.it)y Barcelona Supercomputing Center, Barcelona, Spain, e-mail: frosa.m.badia, jorge.ejarque, [fernando.vazquez](fernando.vazquezg@bsc.es)[g](fernando.vazquezg@bsc.es)[@bsc.es](fernando.vazquezg@bsc.es)  \narXiv :2211 . 10819v1 [ cs .DC] 19 Nov 2022  \nAbstract—The extensive use of HPC infrastructures and frameworks for running data-intensive applications has led to a growing interest in data partitioning techniques and strategies. In fact, ﬁnding an effective partitioning, i.e. a suitable size for data blocks, is a key strategy to speed-up parallel data-intensive applications and increase scalability. This paper describes a methodology for data block size estimation in HPC applications, which relies on supervised machine learning techniques. The implementation of the proposed methodology was evaluated using as a testbed dislib, a distributed computing library highly focused on machine learning algorithms built on top of the PyCOMPSs framework. We assessed the effectiveness of our solution through an extensive experimental evaluation considering different algorithms, datasets, and infrastructures, including the MareNostrum 4 supercomputer. The results we obtained show that the methodology is able to efﬁciently determine a suitable way to split a given dataset, thus enabling the efﬁcient execution of data-parallel applications in high performance environments.  \nIndex Terms—data partitioning, high performance computing, data-parallel applications, machine learning, big data  \nI. INTRODUCTION  \nData partitioning refers to splitting a dataset into small and ﬁxed-size units, called blocks or chunks, to enable efﬁcient data-parallel processing and storing in distributed-memory based systems. Several issues related to data partitioning must be addressed to reduce execution times and ensure good scalability of applications. For example, when a dataset is mapped on a set of nodes of a parallel/distributed computing system, two very critical problems are highlighted: (i) the choice of the destination node for a given block (i.e., the node where that block will be stored); and (ii) the selection of an appropriate block size. The ﬁrst problem has been addressed in several studies [1]–[3], in which scheduling algorithms have been proposed to minimize the movement of data at run-time. The second problem, less studied in the literature and addressed in this work, requires to take a decision before the application is running as it strongly depends on the features of the input dataset, the algorithm, and the execution environment.  \nThe block size can heavily affect the trade-off between single node efﬁciency and parallelism in data-intensive applications. Speciﬁcally, a larger size reduces parallelism (fewer blocks) but makes tasks larger. Although this can lead to an overhead reduction, it must be ensured that the block size does not exceed the memory available on the individual nodes, so as to avoid memory saturation. On the other hand, a smaller size leads to a ﬁner exploitation of parallelism, while introducing  \na larger overhead due to communication, synchronization, and task management, which can negatively impact performance.  \nTypically, block size estimation is not an easy task for programmers. In fact, they usually proceed by following a trial and error approach, only supported by simple heuristics and domain knowledge (i.e., the awareness of the behavior of the algorithm in a given distributed environment) . As a result, this tuning process is often time-consuming and resource-intensive, espec","cbCaidcDD6vEHMSq","https://ap.wps.com/l/cbCaidcDD6vEHMSq","pdf",504176,1,10,"English","en",105,"# Introduction\n## Data partitioning and the block-size challenge\n## Trade-off between single-node efficiency and parallelism\n## Proposed supervised ML methodology\n# Related Work\n# Proposed Methodology\n# Use Case and Testbed\n# Experimental Evaluation and Results\n# Conclusion","[{\"question\":\"Why is block size estimation important in HPC data partitioning?\",\"answer\":\"Block size strongly affects the balance between single-node efficiency and parallelism. It also must fit within available node memory to avoid saturation while limiting communication, synchronization, and task-management overhead.\"},{\"question\":\"How does the proposed methodology estimate an appropriate block size?\",\"answer\":\"It uses supervised machine learning with a cascade of tree-based classifiers. The model is trained on a log of past executions represented by descriptive features about algorithm, dataset, and execution environment.\"},{\"question\":\"How was the methodology evaluated and what were the results?\",\"answer\":\"Experiments were run on dislib across different algorithms, datasets, and infrastructures, including the MareNostrum 4 supercomputer. Results show the method can efficiently predict suitable block sizes to support correct dataset partitioning and improve application performance.\"}]","Block size estimation for data partitioning in HPC applications using machine learning techniques | PDF",1785731184,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"block-size-estimation-for-data-partitioning-in-hpc-applications-using-machine-learning-techniques","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/block-size-estimation-for-data-partitioning-in-hpc-applications-using-machine-learning-techniques/120654/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is block size estimation important in HPC data partitioning?","Question",{"text":75,"@type":76},"Block size strongly affects the balance between single-node efficiency and parallelism. It also must fit within available node memory to avoid saturation while limiting communication, synchronization, and task-management overhead.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed methodology estimate an appropriate block size?",{"text":80,"@type":76},"It uses supervised machine learning with a cascade of tree-based classifiers. The model is trained on a log of past executions represented by descriptive features about algorithm, dataset, and execution environment.",{"name":82,"@type":73,"acceptedAnswer":83},"How was the methodology evaluated and what were the results?",{"text":84,"@type":76},"Experiments were run on dislib across different algorithms, datasets, and infrastructures, including the MareNostrum 4 supercomputer. Results show the method can efficiently predict suitable block sizes to support correct dataset partitioning and improve application performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]