[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119236-en":3,"doc-seo-119236-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119236,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Scalable and Efficient Machine Learning as a Service","Machine learning progress is driving rapid adoption across image recognition, text prediction, translation, and autonomous driving, increasing demand for Machine-Learning-as-a-Service (MLaaS). MLaaS supports model design, training, and inference serving with automated, scalable, and efficient execution. This dissertation introduces three MLaaS methods: SimiGrad for fine-grained adaptive batching in large-scale training, DistQuant for distributed quantization of partitioned weights to cut communication cost without accuracy loss, and RRL for SLO-aware low-latency serving using region-based reinforcement learning scheduling. Experiments on major benchmarks demonstrate improved performance, efficiency, and reduced latency and SLO violations.","University of Nevada, Reno  \nScalable and Efficient Machine Learning as a Service  \nA dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Computer Science and Engineering  \nby  \nHeyang Qin  \nDr. Feng Yan, Dissertation Advisor Dr. Lei Yang, Dissertation Co-advisor  \nDecember 2022  \n© by Heyang Qin 2022 All Rights Reserved  \nThe Graduate School  \nWe recommend that the dissertation prepared under our supervision by  \nHeyang Qin  \nentitled  \nScalable and Efficient Machine Learning as a Service  \nbe accepted in partial fulfillment of the requirements for the degree of  \nDOCTOR OF PHILOSOPHY  \nFeng Yan, Ph.D., Advisor  \nLei Yang, Ph.D., Co-advisor  \nDongfang Zhao, Ph.D., Committee Member  \nEmily Hand, Ph.D., Committee Member  \nPengbo Chu, Ph.D., Graduate School Representative  \nMarkus Kemmelmeier, Ph. D., Dean, Graduate School  \nDecember 2022  \ni  \nAbstract  \nDriven by the sustained advances of machine learning and its application to multiple domains ranging from image recognition, text prediction to translation and autonomous driving, the past few years have witnessed a surging demand for Machine-Learning-asa-Service (MLaaS) . MLaaS is an emerging computing paradigm that facilitates machine learning model design, model training, inference serving and provides optimized executions of machine learning tasks in an automated, scalable, and efficient manner. This dissertation proposes three novel approaches for MLaaS, namely SimiGrad, DistQuant, and RRL, to improve the scale and efficiency of MLaaS training and inference, respectively.  \nFor MLaaS training, we propose SimiGrad, a fine-grained adaptive batching approach for large scale training using gradient similarity measurement. Large scale training requires massive parallelism to finish the training within a reasonable amount of time. To support massive parallelism, large batch training is the key enabler but often at the cost of generalization performance. We propose a fully automated and lightweight adaptive batching methodology to enable fine-grained batch size adaption (e.g., at a mini-batch level) that can achieve state-of-the-art performance with record breaking batch sizes. The core component of our method is a lightweight yet efficient representation of the critical gradient noise information. We open-source the proposed methodology and extensive evaluationson popular benchmarks (e.g., CIFAR10, ImageNet, and BERT-Large) demonstrate that the proposed methodology outperforms state-of-the-art methodologies using adaptive batching approaches or hand-tuned static strategies in both performance and batch size. Particularly, we achieve a new state-of-the-art batch size of 78K in BERT-Large pretraining with a SQuAD score of 90.69 compared to 90.58 reported in previous state-of-the-art with 59K batch size.  \nii  \nAnother key challenge for MLaaS training is the communication cost which limits how much the training can scale. Quantization is a popular method for reducing communication cost yet it imposes non-trivial encoding and decoding overheads and may lead to degraded model performance. Our key observation is that model weights are partitioned and cachedin GPU memory in common distributed training methods such as model and pipeline parallelism. If quantization can be performed on the partitioned weights in parallel while cached in GPU memory, the quantization speed can be significantly improved and we can further reduce the communication overhead for weights gathering. To this end, we propose DistQuant, a distributed quantization scheme for compressing partitioned weights during distributed training. DistQuant preserves model performance by canceling out the noise introduced by quantization and is transparent to training pipelines. We both theoretically and empirically show that DistQuant can achieve much higher precision than state-of-the-art quantization approaches. Evaluation on large-scale models including BERT and GPT2 in","cbCaivEScZg0QHDy","https://ap.wps.com/l/cbCaivEScZg0QHDy","pdf",1884179,1,132,"English","en",105,"# Abstract\n## MLaaS overview\n## SimiGrad: adaptive batching for training\n## DistQuant: distributed quantization for communication reduction\n## RRL: reinforcement-learning serving scheduler for low latency\n## Acknowledgements","[{\"question\":\"What is Machine-Learning-as-a-Service (MLaaS) and why is it needed?\",\"answer\":\"MLaaS is an emerging computing paradigm that automates and scales machine learning model design, training, and inference serving. Growing application demand motivates MLaaS for efficient execution at scale.\"},{\"question\":\"How does SimiGrad improve MLaaS training efficiency?\",\"answer\":\"SimiGrad enables fine-grained adaptive batching using gradient similarity measurement, allowing large-scale training with improved generalization. It uses a lightweight representation of critical gradient noise and outperforms existing adaptive or static batching approaches on major benchmarks.\"},{\"question\":\"How do DistQuant and RRL address communication cost and serving latency?\",\"answer\":\"DistQuant compresses partitioned weights during distributed training, preserving model performance by managing quantization noise and reducing communication overhead. RRL schedules ML model serving with region-based reinforcement learning to select near-optimal parallelism configurations, improving latency and reducing SLO violations.\"}]","Scalable and Efficient Machine Learning as a Service | PDF",1785723233,333,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"scalable-and-efficient-machine-learning-as-a-service","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/scalable-and-efficient-machine-learning-as-a-service/119236/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Machine-Learning-as-a-Service (MLaaS) and why is it needed?","Question",{"text":76,"@type":77},"MLaaS is an emerging computing paradigm that automates and scales machine learning model design, training, and inference serving. Growing application demand motivates MLaaS for efficient execution at scale.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SimiGrad improve MLaaS training efficiency?",{"text":81,"@type":77},"SimiGrad enables fine-grained adaptive batching using gradient similarity measurement, allowing large-scale training with improved generalization. It uses a lightweight representation of critical gradient noise and outperforms existing adaptive or static batching approaches on major benchmarks.",{"name":83,"@type":74,"acceptedAnswer":84},"How do DistQuant and RRL address communication cost and serving latency?",{"text":85,"@type":77},"DistQuant compresses partitioned weights during distributed training, preserving model performance by managing quantization noise and reducing communication overhead. RRL schedules ML model serving with region-based reinforcement learning to select near-optimal parallelism configurations, improving latency and reducing SLO violations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]