[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122529-en":3,"doc-seo-122529-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122529,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",6,"Technology","Using machine learning for intelligent shard sizing on the cloud","Cloud sharding deployments often rely on conservative approximations to choose the number of cloud instances and shard sizes, which can mismatch real demand and create overloaded deployments. Traditional reactive refinement occurs after load spikes, forcing additional work on an already stressed system. The paper introduces an application-specific machine learning approach using multiple linear regression to predict request latency and determine whether cloud capacity satisfies the service level agreement, enabling accurate shard and server sizing and reducing reactive refinement needs. Experiments on a popular database schema report highly accurate predictions and detailed validation results.","Using machine learning for intelligent shard sizing on the cloud  \nNarayanan Venkateswaran1, Anurag Shekhar2, Suvamoy Changder3 and  \nNarayan C. Debnath4  \n1Department of Computer Science and Engineering, National Institute of Technology Durgapur, India  \n2MySQL Oracle India Pvt Ltd., India  \n3Department of Computer Science and Engineering, National Institute of Technology Durgapur, India 4Department of Software Engineering, Eastern International University, Vietnam  \n\n| Received Dec 31st, 2018 |\n| --- |\n| Keyword:\u003Cbr>Machine Learning Sharding\u003Cbr>Horizontal Partitioning Cloud\u003Cbr>Server Sizing Deployment Planning Resource Allocation Data Sizing |\n\nCorresponding Author:  \nSharding implementations use conservative approximations for determining the number of cloud instances required and the size of the shards to be storedon each of them. Conservative approximations are often inaccurate and result in overloaded deployments, which need reactive refinement. Reactive refinement results in demand for additional resources from an already overloaded system and is counterproductive.  \nThis paper proposes an algorithm that eliminates the need for conservative approximations and reduces the need for reactive refinement. A multiple linear regression based machine learning algorithm is used to predict the latency of requests for a given application deployed on a cloud machine. The predicted latency helps to decide accurately and with certainty if the capacity of the cloud machine will satisfy the service level agreement for effective operation of the application. Application of the proposed methods on a popular database schema on the cloud resulted in highly accurate predictions. The results of the deployment and the tests performed to establish the accuracy have been presented in detail and are shown to establish the authenticity of the claims.  \nSuvamoy Changder,  \nDepartment of Computer Science and Engineering, National Institute of Technology Durgapur, Mahatma Gandhi Avenue, Durgapur 713209, West Bengal, INDIA  \nEmail: [suvamoy.nitdgp@gmail.com](suvamoy.nitdgp@gmail.com)  \nArticle Info ABSTRACT  \n1. Introduction  \nAny increase in the number of users of an application causes the data stored and used by the application to increase. This raises the demand for storage and processing resources. Since the cloud allows for such ondemand scaling of resources, it is a popular choice for similar applications requiring elastic backends [16] . Sharding topologies split the database into multiple shards and store one or more shards in each cloud instance [11] [18] . Resharding operations, shard splits and merges, are done on demand by leveraging the elasticity provided by the cloud instances. Several datastores with successful and popular sharding implementations have added elastic extensibility on the cloud [6][26] .  \nSharding implementations distribute the data uniformly across the set of available servers [23] [27] . If the uniform distribution causes workload in excess of server capacity, the amount of data stored is refined reactively.  \nReactive refinement is often performed after the load spike is detected on a shard. Refining shards imposes an additional workload to the system, and can prove to be inefficient.  \nThis paper proposes an application specific, machine learning based partitioning scheme as an alternative for the uniform distribution of data across servers. The methods proposed in this paper allow for accurate server sizing and reduce the need for reactive refinement in an application specific sharding solution. Unlike the existing solutions, the proposed solution attempts to create accurate partitions from the beginning by learning from empirical and existing data.  \nsection 2 talks about existing sharding solutions, their use of reactive refinement and the associated problems. section 3 on page 4 proposes the predictive sharding scheme. The proposed method is tested quantitatively in section 4 on page 7 and it is shown that the ","cbCaifAWgmh6qkGf","https://ap.wps.com/l/cbCaifAWgmh6qkGf","pdf",753285,1,16,"English","en",105,"# Introduction\n## Background and Applicability\n## Reactive refinement and its drawbacks\n## Predictive sharding scheme\n## Quantitative evaluation","[{\"question\":\"Why do conservative approximations in cloud sharding lead to overloaded deployments?\",\"answer\":\"They can be inaccurate when mapping the required number of instances and shard sizes to real demand, so the deployment may not match actual workload needs and becomes overloaded.\"},{\"question\":\"What problem does reactive refinement introduce after a load spike?\",\"answer\":\"Reactive refinement increases workload on the system already experiencing high load, and it can be inefficient while it requests more resources.\"},{\"question\":\"How does the proposed machine learning approach improve shard sizing accuracy?\",\"answer\":\"It uses multiple linear regression to predict request latency for an application on a given cloud machine, helping decide accurately whether capacity will satisfy the service level agreement and reducing the need for reactive refinement.\"}]","Using machine learning for intelligent shard sizing on the cloud | PDF",1785811108,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"using-machine-learning-for-intelligent-shard-sizing-on-the-cloud","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/using-machine-learning-for-intelligent-shard-sizing-on-the-cloud/122529/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do conservative approximations in cloud sharding lead to overloaded deployments?","Question",{"text":75,"@type":76},"They can be inaccurate when mapping the required number of instances and shard sizes to real demand, so the deployment may not match actual workload needs and becomes overloaded.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does reactive refinement introduce after a load spike?",{"text":80,"@type":76},"Reactive refinement increases workload on the system already experiencing high load, and it can be inefficient while it requests more resources.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed machine learning approach improve shard sizing accuracy?",{"text":84,"@type":76},"It uses multiple linear regression to predict request latency for an application on a given cloud machine, helping decide accurately whether capacity will satisfy the service level agreement and reducing the need for reactive refinement.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,117,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":29,"slug":116},7,"Healthcare","healthcare",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":119,"show_sort_weight":120,"slug":121},8,"Research & Report",30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]